pointcal-c / ERRATUM.md
grKnight's picture
Add erratum for the manifest ledger and git_dirty defects
3563719 verified
|
Raw
History Blame Contribute Delete
3.55 kB

Erratum

Known defects in this archive, with what is affected and what to trust. Recorded rather than quietly patched, because two of them cannot be corrected without re-running the paid inference stage.

1. run_manifest.json embeds the wrong ledger (open)

Affected: runs/full/provenance/run_manifest.json, key ledger.

The full tier was produced by a real inference run and then re-invoked against the warm logit cache. The re-invocation correctly declined to overwrite ledger_inference.json, but still stamped its own near-zero ledger into the run manifest, so the two disagree:

items_processed throughput USD peak VRAM
ledger_inference.json (authoritative) 903,558 2,808 views/s $0.025 12.1 GB
run_manifest.jsonledger (wrong) 0 0 views/s ~$0.000004 0.57 GB

Trust ledger_inference.json. Every figure and the reported compute table are derived from it, not from the manifest copy.

The cause is fixed in source: run_inference now reads the preserved ledger back and returns it, so the manifest, the printed summary and ledger_inference.json cannot disagree in future runs. The stale manifest in this archive is not regenerated because only the GPU stage writes it.

2. git_dirty: true is not meaningful here (open)

Affected: run_manifest.json, keys git_commit and git_dirty.

_git_dirty() calls git status --porcelain and reports true on any output, including untracked files. The documented workflow creates artifacts/split.json before inference and .gitignore does not exclude it, so git_dirty is true on a clean checkout as well. The flag therefore does not distinguish "modified tracked source" from "generated an expected artifact", and should not be read as evidence that the producing source was modified.

git_commit records the upstream base commit, not a commit containing the corrections applied at run time.

3. Figures and summaries were stale (fixed)

Affected: runs/*/figures/*, runs/*/results/results_summary.md.

runs/full/figures/fig4_cost_throughput.png previously showed 1.65e+07 views/s and near-zero spend, taken from the ledger described in item 1 at a moment when it was transiently wrong. The summaries additionally declared H1 from accuracy alone and described the fitted model as three scalars rather than four.

All are regenerated in this revision. Figure 4 now shows 0.0894 GPU-h, $0.025, 2,808 views/s and 12.1 GB, and the summaries carry a per-metric H1 table and the correct four-scalar count.

4. H1 does not hold strictly (finding, not a defect)

Accuracy, NLL, Brier and AURC all worsen monotonically with severity. ECE does not: it dips at severities 3 to 4 and again from clean to severity 1. A paired object-grouped bootstrap over base object IDs places both reversals within noise of zero, so the direction of H1 is supported while strict monotonicity is not. The summaries state this; earlier revisions did not.

5. Bootstrap depth deviates from the pre-registration (open)

configs/full.yaml runs 200 bootstrap replicates; the pre-registration specified 1,000. The reduction was made because the CPU bootstrap dominated runtime. spec_hash does not cover evaluation settings, so it is unchanged and does not flag this. Deeper intervals can be recomputed from the cached logits with no GPU.

Intervals on delta_*_vs_clean rows in results.csv are NaN: degradation deltas are reported as point estimates only, without paired intervals.