TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models Paper • 2610.07767 • Published 4 days ago • 87
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents Paper • 2609.39102 • Published 10 days ago • 498
LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches Paper • 2610.06647 • Published 5 days ago • 92
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence Paper • 2609.17488 • Published 25 days ago • 559
Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model Paper • 2609.18323 • Published 24 days ago • 119