Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring Paper • 2605.30834 • Published May 29 • 11
Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents Paper • 2606.26080 • Published Jun 24 • 12
Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents Paper • 2606.26080 • Published Jun 24 • 12
DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Paper • 2606.05758 • Published Jun 4 • 5
Capturing LLM Capabilities via Evidence-Calibrated Query Clustering Paper • 2605.17110 • Published May 16 • 2
Capturing LLM Capabilities via Evidence-Calibrated Query Clustering Paper • 2605.17110 • Published May 16 • 2
Breakeven complexity: A new perspective on neural partial differential equation solvers Paper • 2605.15399 • Published May 14
From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing Paper • 2605.15181 • Published May 14 • 12
Exploration and Exploitation Errors Are Measurable for Language Model Agents Paper • 2604.13151 • Published Apr 14 • 25
Exploration and Exploitation Errors Are Measurable for Language Model Agents Paper • 2604.13151 • Published Apr 14 • 25
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks Paper • 2603.24755 • Published Mar 25 • 30
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States Paper • 2603.19987 • Published Mar 20 • 9
How to Move Your Dragon: Text-to-Motion Synthesis for Large-Vocabulary Objects Paper • 2503.04257 • Published Mar 6, 2025 • 2
T1: Tool-integrated Self-verification for Test-time Compute Scaling in Small Language Models Paper • 2504.04718 • Published Apr 7, 2025 • 43
Distilling LLM Agent into Small Models with Retrieval and Code Tools Paper • 2505.17612 • Published May 23, 2025 • 82