onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction Paper • 2609.24983 • Published 2 days ago • 27
MatrAIx: Simulating the World with 8.3 Billion Persona Agents Paper • 2608.04205 • Published Aug 4 • 54
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment Paper • 2607.07820 • Published Jul 8 • 95
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 165
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 52
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration Paper • 2605.20025 • Published May 19 • 90
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling Paper • 2605.13301 • Published May 13 • 166
Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents Paper • 2605.10832 • Published May 11 • 22