Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
Abstract
Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Across 26 economics papers, ARGUS abstains on about 40% of paper-dimension assessments for lack of retrievable evidence. In a five-paper pilot with labels reconciled by two annotators, it assigns a higher risk level than the labels on 25 of the 33 assessments it completes. A rule fixed before the labels arrived removes most of this in-sample; weighted agreement stays low. ARGUS provides evidence-linked risk reports that localize potential weaknesses for expert review, without adjudicating causal claims. Code and data: https://github.com/yonghongzhang-io/ARGUS
Community
ARGUS audits the evidence that difference-in-differences studies report for their identification assumptions: it maps each paper onto an 11-dimension assumption–implication–evidence rubric, flags risk per dimension, and abstains when the evidence cannot be retrieved. It does not judge whether an estimate is correct. In a five-paper human pilot it is systematically more severe than the reconciled labels (25 of 33 answered cells), which we report as a limitation. Data: https://huggingface.co/datasets/yonghongzhang/ARGUS (ClimateNLP @ EMNLP 2026)
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Claim-Gated Source-Risk Auditing for Generative Search (2026)
- Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers (2026)
- Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding (2026)
- Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay (2026)
- A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation (2026)
- PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress (2026)
- ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper