Towards Full Pipeline FP8 Reinforcement Learning for LLMs Paper • 2609.22870 • Published 6 days ago • 15
ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models Paper • 2609.13231 • Published 23 days ago • 13
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents Paper • 2609.23377 • Published 5 days ago • 48
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 7 days ago • 131
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published 21 days ago • 114
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 7 days ago • 32
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence Paper • 2609.20886 • Published 9 days ago • 28
From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention Paper • 2609.21788 • Published 7 days ago • 13
Region-Level Policy Optimization for Fine-grained MLLM Perception Paper • 2609.19745 • Published 8 days ago • 41
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL Paper • 2609.20715 • Published 8 days ago • 42
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 8 days ago • 56
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 9 days ago • 79
ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments Paper • 2609.19134 • Published 9 days ago • 100
ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models Paper • 2609.18487 • Published 9 days ago • 47
SenseNova-U1.5: Towards Native Unified Visual Intelligence Paper • 2609.11929 • Published 15 days ago • 274
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 11 days ago • 214
Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models Paper • 2609.13680 • Published 13 days ago • 13