RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling Paper • 2609.22947 • Published 10 days ago • 39
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL Paper • 2609.20715 • Published 12 days ago • 44
ImIR: Image-Instruction Tuning for All-in-One Image Restoration Paper • 2609.25267 • Published 8 days ago • 14
RULER: Instance-aware Rubric Rewards for SVG Generation Paper • 2609.25270 • Published 8 days ago • 101
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 11 days ago • 136
ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning Paper • 2609.22323 • Published 13 days ago • 10