Papers
arxiv:2608.19269

Stable Within, Unidentified Across: Semantic Identification of Benchmark Effects and Rankings

Published on Aug 18
Authors:

Abstract

Evaluation conclusions depend on evaluator-controlled semantics: legal references, scoreability, and aggregation. We call an artifact-defined endpoint Q(s) evaluation-semantically identified when it is invariant over a declared family. A frozen 217-row analysis appears stable within its restricted contract family. In TraceElephant, F_asym yields precise task-disjoint adoption effects from -40.76 to -43.03 percentage points, whereas both jointly frozen contrast-comparable treatments in F_cc yield exactly zero with [0, 0] intervals. Their union F_audit therefore does not identify T, even though F_cc identifies T = 0. A factorial audit attributes 100 percent of the disagreement to one-sided terminal deletion, with zero nonterminal contribution. Rankings remain specification-dependent across the symmetric treatments. This is a localized non-implication witness and reusable family-indexed audit, not a prevalence estimate, exhaustive semantic taxonomy, or independent replication.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.19269
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.19269 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2608.19269 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.19269 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.