MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training Paper • 2606.30406 • Published about 1 month ago • 17
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published 21 days ago • 76
Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Grounding Paper • 2606.18101 • Published Jun 16 • 15
Online Skill Learning for Web Agents via State-Grounded Dynamic Retrieval Paper • 2606.04391 • Published Jun 3 • 11
TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL Paper • 2606.01599 • Published Jun 1 • 17
Less is Enough: Synthesizing Diverse Data in Feature Space of LLMs Paper • 2602.10388 • Published Feb 11 • 246
MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information Paper • 2510.03632 • Published Oct 4, 2025 • 42