Stable Within, Unidentified Across: Semantic Identification of Benchmark Effects and Rankings
Abstract
Evaluation conclusions depend on evaluator-controlled semantics: legal references, scoreability, and aggregation. We call an artifact-defined endpoint Q(s) evaluation-semantically identified when it is invariant over a declared family. A frozen 217-row analysis appears stable within its restricted contract family. In TraceElephant, F_asym yields precise task-disjoint adoption effects from -40.76 to -43.03 percentage points, whereas both jointly frozen contrast-comparable treatments in F_cc yield exactly zero with [0, 0] intervals. Their union F_audit therefore does not identify T, even though F_cc identifies T = 0. A factorial audit attributes 100 percent of the disagreement to one-sided terminal deletion, with zero nonterminal contribution. Rankings remain specification-dependent across the symmetric treatments. This is a localized non-implication witness and reusable family-indexed audit, not a prevalence estimate, exhaustive semantic taxonomy, or independent replication.
Get this paper in your agent:
hf papers read 2608.19269 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper