VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control Paper • 2609.19554 • Published 12 days ago • 42
HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses Paper • 2609.15938 • Published 15 days ago • 31
Scaling Automatic Research Agents via World Models Paper • 2608.12564 • Published about 1 month ago • 482
DarwinX: Evolving Agent Harnesses Through Natural Selection Paper • 2608.07545 • Published Jul 31 • 116
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 265
MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models Paper • 2607.27637 • Published Aug 1 • 7
OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Paper • 2608.03812 • Published Aug 4 • 29
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published Jul 26 • 96