RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents Paper • 2609.22000 • Published 7 days ago • 78
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 7 days ago • 131
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 8 days ago • 178
SenseNova-U1.5: Towards Native Unified Visual Intelligence Paper • 2609.11929 • Published 15 days ago • 274
OpenThinker-Agent2 Collection OpenThinker-Agent2: agentic SFT/RL datasets and 8B/32B models (cold-start SFT, RL, and the OpenThinkerAgent-32B release). • 11 items • Updated Jun 11 • 11
Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding Paper • 2606.21906 • Published Jun 20 • 27
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 25 days ago • 63
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training Paper • 2609.07108 • Published 18 days ago • 36
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 22 days ago • 101
SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback Paper • 2608.13120 • Published Aug 13 • 32
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks Paper • 2608.23035 • Published Aug 24 • 42
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 144
FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory Paper • 2608.04530 • Published Aug 5 • 14
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Paper • 2607.07508 • Published Jul 8 • 33