DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? Paper • 2608.10366 • Published 3 days ago • 7
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published 8 days ago • 46
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published 14 days ago • 96
GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning Paper • 2608.02585 • Published 11 days ago • 24
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 14 days ago • 38