Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective Paper • 2610.03185 • Published 7 days ago • 27
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published Aug 31 • 63
MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering Paper • 2601.22859 • Published Jan 30 • 18 • 5
MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering Paper • 2601.22859 • Published Jan 30 • 18
SpecBundle Collection A collection of production-grade draft models for speculative decoding • 30 items • Updated Aug 4 • 20
ProxyAttn: Guided Sparse Attention via Representative Heads Paper • 2509.24745 • Published Sep 29, 2025 • 1