Submitted by Yuzhe Gu 33 AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification Intern Large Models 3
Submitted by Shengyuan Ding 51 Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games Intern Large Models 41 5
Submitted by Yuzhe Gu 25 ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Intern Large Models 12 2
Submitted by Yifei Li 31 OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs Intern Large Models 50 2
2 What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents Intern Large Models 71
Submitted by Shuangrui Ding 46 WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation Intern Large Models 511 3
Submitted by Yining Li 13 TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration Intern Large Models 2