Agentic RAG Evaluation: Budget Allocation Across Questions, Trajectories, and Reads Paper • 2610.05034 • Published 5 days ago • 18
Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL Paper • 2609.37200 • Published 10 days ago • 138
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training Paper • 2609.07108 • Published Sep 7 • 36
WorldReward: Reward Modeling for Camera-Conditioned World Models Paper • 2609.03952 • Published Sep 3 • 28