Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published 19 days ago • 92
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 5 days ago • 104
MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup Paper • 2609.15126 • Published 9 days ago • 7
Calibrating Teacher--Student Discrepancy for On-Policy Distillation Paper • 2609.21619 • Published 5 days ago • 7
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling Paper • 2609.19499 • Published 7 days ago • 30
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL Paper • 2609.20715 • Published 6 days ago • 33
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 6 days ago • 106
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 6 days ago • 104
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 6 days ago • 159
What Does Privileged Information Add to On-Policy Self-Distillation? Paper • 2609.20612 • Published 6 days ago • 31
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models Paper • 2609.19671 • Published 6 days ago • 41
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 6 days ago • 46
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 16 days ago • 370
ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals Paper • 2609.16816 • Published 8 days ago • 11
Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation Paper • 2609.13770 • Published 11 days ago • 9
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence Paper • 2609.17488 • Published 8 days ago • 591
SpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization Paper • 2609.14320 • Published 10 days ago • 33
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 7 days ago • 75
Register Tokens for Bounded-State Reasoning in Diffusion Language Models Paper • 2609.16372 • Published 9 days ago • 7