WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents Paper • 2609.36887 • Published 8 days ago • 19
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 6 days ago • 269
HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents Paper • 2610.03574 • Published 5 days ago • 55
DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation Paper • 2610.00360 • Published 7 days ago • 6
Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation Paper • 2610.00348 • Published 8 days ago • 13
Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs Paper • 2609.32259 • Published 8 days ago • 95
Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts Paper • 2610.00314 • Published 8 days ago • 101