A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization Paper • 2610.12183 • Published 2 days ago • 4
A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization Paper • 2610.12183 • Published 2 days ago • 4
In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks Paper • 2609.38173 • Published 11 days ago • 41
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 24 days ago • 84
Training Diffusion Language Models for Black-Box Optimization Paper • 2603.17919 • Published May 29 • 15
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails Paper • 2607.05910 • Published Jul 7 • 31
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments Paper • 2607.02440 • Published Jul 2 • 49
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research Paper • 2606.07591 • Published May 28 • 106
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation Paper • 2605.31264 • Published May 29 • 132
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling Paper • 2605.13301 • Published May 13 • 166
UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience Paper • 2603.24533 • Published Mar 25 • 46
Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning Paper • 2512.06533 • Published Dec 6, 2025 • 9 • 2
Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning Paper • 2512.06533 • Published Dec 6, 2025 • 9
Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning Paper • 2512.06533 • Published Dec 6, 2025 • 9
P1: Mastering Physics Olympiads with Reinforcement Learning Paper • 2511.13612 • Published Nov 17, 2025 • 135