MInTRL: Off-policy Intervention can boost On-policy RL Paper • 2609.12419 • Published 7 days ago • 10
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training Paper • 2609.15051 • Published 4 days ago • 11
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction Paper • 2609.13285 • Published 10 days ago • 68
MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control Paper • 2609.06251 • Published 13 days ago • 3
DataFlex-RL: An Evaluation Platform for RLVR Data Policies Paper • 2609.06107 • Published 13 days ago • 125
Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work Paper • 2609.11977 • Published 14 days ago • 126
DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat Paper • 2609.11155 • Published 8 days ago • 25
CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation Paper • 2609.06931 • Published 11 days ago • 26
Negative Self-Distillation: Learning to Reason by Avoiding Flaws Paper • 2609.11699 • Published 8 days ago • 37
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics Paper • 2609.10712 • Published 9 days ago • 43
SenseNova-U1.5: Towards Native Unified Visual Intelligence Paper • 2609.11929 • Published 8 days ago • 262
Studying Image Tokenizers as Visual Languages in Unified Multimodal Models Paper • 2609.09143 • Published 10 days ago • 31
Revisiting Complete Reasoning Traces for Post-Training Paper • 2609.07103 • Published 11 days ago • 23
Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR Paper • 2609.08650 • Published 10 days ago • 12
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data Paper • 2609.05405 • Published 14 days ago • 42
It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning Paper • 2609.00638 • Published 17 days ago • 75
Cliff: Learning Process Rewards from the First Mistake Paper • 2609.02817 • Published 16 days ago • 21
Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks Paper • 2609.08404 • Published 10 days ago • 25
One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation Paper • 2608.25936 • Published 23 days ago • 17