Running on Zero MCP 4 Valen Visual Decisions 👁 4 Visual question answering with candidate probabilities
Paint-Anything: Unified Any-Color Control for Image Generation and Editing Paper • 2609.20816 • Published 9 days ago • 54
OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation Paper • 2609.22069 • Published 8 days ago • 35
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 8 days ago • 132
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published 22 days ago • 116
RNGBench Collection [EMNLP2026] Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games • 2 items • Updated 11 days ago • 2
WorldReward: Reward Modeling for Camera-Conditioned World Models Paper • 2609.03952 • Published 23 days ago • 27
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models Paper • 2608.25518 • Published Aug 26 • 59
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution Paper • 2608.25593 • Published Aug 26 • 69
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows Paper • 2608.19741 • Published Aug 20 • 12