Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control Paper • 2608.12123 • Published 2 days ago • 2 • 3
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents Paper • 2608.08389 • Published 5 days ago • 12 • 3
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace Paper • 2608.08621 • Published 5 days ago • 20 • 5
Omega-S: A Functional Resilience Index for LLM Fine-Tuning Paper • 2608.03887 • Published 10 days ago • 7 • 4
Stealing Reasoning Traces from Proprietary LLM APIs Paper • 2608.09867 • Published 4 days ago • 97 • 3
Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression Paper • 2608.04569 • Published 9 days ago • 12 • 3
Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss Paper • 2608.03796 • Published 10 days ago • 14 • 4
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces Paper • 2608.03451 • Published 10 days ago • 33 • 4
Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay Paper • 2608.05784 • Published 8 days ago • 28 • 6
Invisible Shortcuts: Why Vision Encoders Know Your Camera Paper • 2608.05424 • Published 9 days ago • 17 • 3
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents Paper • 2608.05013 • Published 10 days ago • 35 • 3
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? Paper • 2608.03874 • Published 10 days ago • 14 • 4
Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants Paper • 2607.26611 • Published 16 days ago • 32 • 3
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space Paper • 2607.25675 • Published 17 days ago • 67 • 3
A New Role for Relevance: Guiding Corpus Interaction in Agentic Search Paper • 2607.24223 • Published 18 days ago • 95 • 6
Codifying the Judge: Scalable Evaluation via Program Distillation Paper • 2607.22561 • Published May 29 • 8 • 3
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems Paper • 2607.21503 • Published 22 days ago • 28 • 4
OpenForgeRL: Train Harness-native Agents in Any Environment Paper • 2607.21557 • Published 22 days ago • 10 • 4