EVISKILL: Grounding Skill Evolution in Replayable Evidence Paper • 2610.05030 • Published 6 days ago • 50
Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training Paper • 2610.07510 • Published 5 days ago • 12
TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models Paper • 2610.07767 • Published 4 days ago • 87
Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym Paper • 2609.37267 • Published 11 days ago • 40
LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures Paper • 2610.04292 • Published 7 days ago • 38
What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Document Paper • 2610.04128 • Published 8 days ago • 12
VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks Paper • 2610.00972 • Published 9 days ago • 59
Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It Paper • 2610.03195 • Published 8 days ago • 46
HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents Paper • 2610.03574 • Published 8 days ago • 60
ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization Paper • 2610.00906 • Published 9 days ago • 83
X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization Paper • 2609.32993 • Published 14 days ago • 78
SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation Paper • 2609.37539 • Published 11 days ago • 12
EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making Paper • 2609.38334 • Published 9 days ago • 80
AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks Paper • 2609.38288 • Published 11 days ago • 136
ExpVoyager: Direct Experience Navigation for Dynamic Agent Skill Synthesis Paper • 2609.32630 • Published 14 days ago • 19
TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces Paper • 2609.33295 • Published 13 days ago • 72
Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents Paper • 2609.29892 • Published 16 days ago • 34
MaLiang-Harness: A Programmable Path to Image and Video Generation Paper • 2609.34309 • Published 12 days ago • 394
IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis Paper • 2609.29444 • Published 16 days ago • 21