The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean Paper • 2609.09218 • Published Sep 6 • 2
The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean Paper • 2609.09218 • Published Sep 6 • 2
Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations Paper • 2609.30867 • Published 12 days ago • 6
Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations Paper • 2609.30867 • Published 12 days ago • 6
Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations Paper • 2609.30867 • Published 12 days ago • 6
Sleeping RL ComtradeBench: An OpenEnv Benchmark for Reliable LLM Tool-Use Under Adversarial API Conditions 📊 Benchmark LLM agents on robust data‑fetching tool use
Sleeping RL ComtradeBench: An OpenEnv Benchmark for Reliable LLM Tool-Use Under Adversarial API Conditions 📊 Benchmark LLM agents on robust data‑fetching tool use