VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction Paper • 2610.06293 • Published 4 days ago • 53
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published Sep 1 • 569
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Paper • 2606.11042 • Published Jun 9 • 221
OmniEdu: Open Foundation Models for Learning and Teaching Paper • 2609.23088 • Published 20 days ago • 239
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published Aug 26 • 337
PixelHacker: Image Inpainting with Structural and Semantic Consistency Paper • 2504.20438 • Published Apr 29, 2025 • 51
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 25 days ago • 215