Erratum
Known defects in this archive, with what is affected and what to trust. Recorded rather than quietly patched, because two of them cannot be corrected without re-running the paid inference stage.
1. run_manifest.json embeds the wrong ledger (open)
Affected: runs/full/provenance/run_manifest.json, key ledger.
The full tier was produced by a real inference run and then re-invoked against
the warm logit cache. The re-invocation correctly declined to overwrite
ledger_inference.json, but still stamped its own near-zero ledger into the run
manifest, so the two disagree:
| items_processed | throughput | USD | peak VRAM | |
|---|---|---|---|---|
ledger_inference.json (authoritative) |
903,558 | 2,808 views/s | $0.025 | 12.1 GB |
run_manifest.json → ledger (wrong) |
0 | 0 views/s | ~$0.000004 | 0.57 GB |
Trust ledger_inference.json. Every figure and the reported compute table
are derived from it, not from the manifest copy.
The cause is fixed in source: run_inference now reads the preserved ledger
back and returns it, so the manifest, the printed summary and
ledger_inference.json cannot disagree in future runs. The stale manifest in
this archive is not regenerated because only the GPU stage writes it.
2. git_dirty: true is not meaningful here (open)
Affected: run_manifest.json, keys git_commit and git_dirty.
_git_dirty() calls git status --porcelain and reports true on any output,
including untracked files. The documented workflow creates artifacts/split.json
before inference and .gitignore does not exclude it, so git_dirty is true on
a clean checkout as well. The flag therefore does not distinguish "modified
tracked source" from "generated an expected artifact", and should not be read as
evidence that the producing source was modified.
git_commit records the upstream base commit, not a commit containing the
corrections applied at run time.
3. Figures and summaries were stale (fixed)
Affected: runs/*/figures/*, runs/*/results/results_summary.md.
runs/full/figures/fig4_cost_throughput.png previously showed 1.65e+07 views/s
and near-zero spend, taken from the ledger described in item 1 at a moment when
it was transiently wrong. The summaries additionally declared H1 from accuracy
alone and described the fitted model as three scalars rather than four.
All are regenerated in this revision. Figure 4 now shows 0.0894 GPU-h, $0.025, 2,808 views/s and 12.1 GB, and the summaries carry a per-metric H1 table and the correct four-scalar count.
4. H1 does not hold strictly (finding, not a defect)
Accuracy, NLL, Brier and AURC all worsen monotonically with severity. ECE does not: it dips at severities 3 to 4 and again from clean to severity 1. A paired object-grouped bootstrap over base object IDs places both reversals within noise of zero, so the direction of H1 is supported while strict monotonicity is not. The summaries state this; earlier revisions did not.
5. Bootstrap depth deviates from the pre-registration (open)
configs/full.yaml runs 200 bootstrap replicates; the pre-registration
specified 1,000. The reduction was made because the CPU bootstrap dominated
runtime. spec_hash does not cover evaluation settings, so it is unchanged and
does not flag this. Deeper intervals can be recomputed from the cached logits
with no GPU.
Intervals on delta_*_vs_clean rows in results.csv are NaN: degradation
deltas are reported as point estimates only, without paired intervals.