Beyond Solver Verdicts: Generative Reward Models for Autoformalization Paper • 2609.11085 • Published 5 days ago • 31
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 7 days ago • 414
Demystifying Agent Skills: Why They Work-Until They Don't Paper • 2608.14036 • Published Aug 14 • 170
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 29 days ago • 340
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Paper • 2608.11924 • Published Aug 12 • 291
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published Jul 30 • 312
Benchmarking Sensor Robustness in Plasma Diagnostic Models: A Systematic Evaluation on TokaMark Paper • 2607.11915 • Published Jul 5 • 4
NeuroCogMap Reveals Cognitive Organization of Large Language Models Paper • 2607.00397 • Published Jul 1 • 13
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 254
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection Paper • 2605.30189 • Published May 28 • 9
Geo-Align: Video Generation Alignment via Metric Geometry Reward Paper • 2605.23903 • Published May 22 • 10
Look Before You Leap: Autonomous Exploration for LLM Agents Paper • 2605.16143 • Published May 15 • 10
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation Paper • 2605.13724 • Published May 13 • 105
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation Paper • 2604.28196 • Published Apr 30 • 76
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing Paper • 2604.04911 • Published Apr 6 • 36
Adam's Law: Textual Frequency Law on Large Language Models Paper • 2604.02176 • Published Apr 2 • 511