CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers? Paper • 2610.07557 • Published 2 days ago • 52
EvoOntology: A Self-Evolving Ontology Layer for Data Agents Paper • 2609.15779 • Published 24 days ago • 159
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Paper • 2608.20492 • Published Aug 20 • 87