Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States Paper • 2610.01415 • Published 10 days ago • 91
UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement Paper • 2609.38721 • Published 11 days ago • 273
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published Sep 8 • 321