🤏 Smol-Data Collection Tried and tested mixes for strong pretraining. Inspired by https://huggingface.co/blog/codelion/optimal-dataset-mixing • 14 items • Updated Mar 2 • 18
UEmbed: Unified Sparse and Dense Multimodal Embeddings Paper • 2608.02583 • Published 12 days ago • 50
view post Post 3338 Two big projects are open sourced soon. Get ready... See translation 11 replies · 🔥 9 9 + Reply
Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System Paper • 2606.18112 • Published Jun 18 • 29
The Verification Horizon: No Silver Bullet for Coding Agent Rewards Paper • 2606.26300 • Published Jun 24 • 53
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation Paper • 2606.26907 • Published Jun 25 • 52
Native Active Perception as Reasoning for Omni-Modal Understanding Paper • 2606.19341 • Published Jun 17 • 19
Rethinking the Role of Efficient Attention in Hybrid Architectures Paper • 2606.15378 • Published Jun 13 • 20
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Paper • 2606.17030 • Published Jun 15 • 47
OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under optimal Squared error quantization Paper • 2605.21226 • Published May 20 • 9
Measuring Maximum Activations in Open Large Language Models Paper • 2605.15572 • Published May 15 • 18
view article Article Training-Free Reasoning at 88.89% on GPQA Diamond: How Darwin Family Hit Frontier Scores Without a Single Gradient Step FINAL-Bench • May 15 • 18
view article Article Vocabulary-Augmented Prompting for Sango — Production African Language AI Without a Parallel Corpus MEYNG • May 13 • 2