Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing Paper • 2606.01393 • Published May 31
MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos Paper • 2607.00491 • Published Jul 1
Faithful, Enriched, and Precise: Benchmarking Natural-Science Illustration Generation by T2I models Paper • 2606.05949 • Published Jun 5
SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models Paper • 2608.04244 • Published Aug 4 • 4
CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents Paper • 2606.22883 • Published Jun 22 • 38
CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents Paper • 2606.22883 • Published Jun 22 • 38
TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation Paper • 2606.02320 • Published Jun 1 • 13
Sample-Efficient Post-Training for LEGO Spatial-Physics Reasoning Paper • 2606.07602 • Published May 29 • 6
WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models Paper • 2604.18224 • Published Apr 20 • 21
OProver: A Unified Framework for Agentic Formal Theorem Proving Paper • 2605.17283 • Published May 17 • 32
ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding Paper • 2603.27064 • Published Mar 28 • 27
Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality Paper • 2604.04418 • Published Apr 6 • 1
Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality Paper • 2604.04418 • Published Apr 6 • 1
ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding Paper • 2603.27064 • Published Mar 28 • 27