How to Loop MoE: Flatten the Experts, Untie the Attention Paper • 2609.35751 • Published 9 days ago • 2
Beyond Solver Verdicts: Generative Reward Models for Autoformalization Paper • 2609.11085 • Published 27 days ago • 35
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean Paper • 2609.09264 • Published 29 days ago • 9
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Paper • 2608.04001 • Published Aug 4 • 1
Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers Paper • 2601.07036 • Published Jan 11
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean Paper • 2609.09264 • Published 29 days ago • 9
BIOCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models Paper • 2510.20095 • Published Oct 23, 2025 • 2
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Paper • 2506.09082 • Published May 3 • 1