Beyond Solver Verdicts: Generative Reward Models for Autoformalization Paper • 2609.11085 • Published 4 days ago • 27
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 6 days ago • 411
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published 14 days ago • 383
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 28 days ago • 340
EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published 25 days ago • 275
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Paper • 2608.11924 • Published Aug 12 • 291
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published Jul 29 • 140
DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization Paper • 2605.31455 • Published May 29 • 6