OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling Paper • 2608.05141 • Published 8 days ago • 1
ADIAS: Automated Design of Interactive Agentic Systems Paper • 2608.06410 • Published 10 days ago • 2
Scrouting: Cost-Aware Routing of Coding Agents by Scouting the Repository First Paper • 2608.04804 • Published 8 days ago • 1
Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization Paper • 2608.01078 • Published 11 days ago • 1
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms Paper • 2608.04074 • Published 9 days ago • 2
Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories Paper • 2608.02276 • Published 10 days ago • 4
Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments Paper • 2608.02670 • Published 11 days ago • 2
InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval Paper • 2608.04761 • Published 7 days ago • 1
NOVA: A Verification-Aware Agent Harness for Architecture Evolution in Industrial Recommender Systems Paper • 2606.27243 • Published 16 days ago • 1
Antares: Foundation Models for Agentic Vulnerability Localization Paper • 2608.02407 • Published 10 days ago • 2
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Paper • 2607.23722 • Published 18 days ago • 1
Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment Paper • 2608.02786 • Published 10 days ago • 1
DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation Paper • 2608.10636 • Published 2 days ago • 5
UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models Paper • 2608.08627 • Published 4 days ago • 6
OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills Paper • 2607.20121 • Published 22 days ago • 1
Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions Paper • 2607.12406 • Published about 1 month ago • 1
Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study Paper • 2607.02436 • Published Jul 2 • 1
Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents Paper • 2608.03327 • Published 7 days ago • 3
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles Paper • 2607.27670 • Published 9 days ago • 3