thundercode commited on
Commit
6df4c6e
·
verified ·
1 Parent(s): 6b7a6df

release: add docs/LIMITATIONS.md

Browse files
Files changed (1) hide show
  1. docs/LIMITATIONS.md +699 -62
docs/LIMITATIONS.md CHANGED
@@ -4,82 +4,708 @@ An honest, exhaustive catalogue of everything SatQuery AI does **not** do, does
4
  know. Negative results and open items are listed here rather than omitted, because a limitation that
5
  is not written down is a limitation that will be discovered by someone else at the worst moment.
6
 
7
- **Status tags:** `OPEN` · `NOT RUN` · `REJECTED` · `DEFERRED` · `BY DESIGN`.
 
 
 
 
 
 
 
 
 
 
 
8
 
9
  ---
10
 
11
  ## 1. Model quality
12
 
13
- | # | Limitation | Detail |
14
- |---|---|---|
15
- | 1 | **Grounding IoU is low in absolute terms** | mean best IoU **0.2838** (canonical) / **0.2566** (matched6). The trained head clearly beats the zero-shot baseline (0.0972), but 0.28 is not "solved". |
16
- | 2 | **Grounding is protocol-sensitive** | two protocols × two decode variants give very different numbers: head_threshold 0.2838, head_argmax **0.1215**, zero-shot 0.0972. An absolute value is meaningless without its protocol. |
17
- | 3 | **Optical-SAR accuracy is carried by common classes** | accuracy **0.931** but macro-F1 **0.434161**. 5 of 19 classes are absent in the scored split and contribute 0.0 to macro-F1 by construction. Never quote accuracy alone. |
18
- | 4 | **Change-VQA is weak on rare classes** | accuracy 0.697626 / macro-F1 0.378373 (test) and 0.651469 / 0.372309 (test2). The wide accuracy–macroF1 gap is the signature of class imbalance. |
19
- | 5 | **VQA is weak-but-related** | the live VQA path answers broadly related content (e.g. "Grassland") rather than a crisp class. |
20
- | 6 | **Optical-SAR returns a bare class index** | the live service returns `class_18`, not a human-readable CLC label. |
21
- | 7 | **Calibration made things worse** | ECE **0.013755 → 0.014929** (`ece_improvement` −0.001174). Retained only because it is part of the frozen config — **not** because it helped. |
22
- | 8 | **The VLM adapter is not accepted** | metrics usable (exact_match 0.963, F1 0.96432, +49.5 pp) but the artifact is **ACCEPTANCE-REJECTED**; the deployed caption/VQA path uses the **unadapted** model. |
23
- | 9 | **The router number is not a test result** | 0.965116 is **validation**, **ungated**, **n = 86**, corpus-limited. The test split was **NOT RUN**. |
24
- | 10 | **The router's routing is not perfect** | known residuals below (§2). |
25
 
26
- ## 2. Router residuals (known misroutes)
 
 
 
27
 
28
- | # | Query | Behaviour | Note |
 
 
 
 
 
 
 
29
  |---|---|---|---|
30
- | 11 | *"What is the new runway?"* | reads `change`, not `vqa` | the word "new" triggers a change reading |
31
- | 12 | *"How much built-up area was added?"* | reads `vqa` (under-trigger) | a change-style quantifier the router does not catch |
32
- | 13 | *"What changed between the earlier and later image?"* with **one** asset | console **reads** `change`, **dispatches** `change_vqa` | **intentional** — reading is asset-count-blind, dispatch is asset-count-aware — but visually surprising. Documented in [`architecture/04-router.md`](architecture/04-router.md). |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
33
 
34
  ## 3. Evaluation gaps
35
 
36
- | # | Gap | State |
37
- |---|---|---|
38
- | 14 | **No system-level end-to-end benchmark** | **NOT RUN — none exists.** No end-to-end accuracy is claimed anywhere. |
39
- | 15 | **Router test split** | **NOT RUN** |
40
- | 16 | **Benchmark adapters** | **NOT RUN** |
41
- | 17 | **End-to-end latency benchmark** | **NOT RUN** (per-specialist latency is recorded only incidentally) |
42
- | 18 | **Cross-dataset generalisation** | **NOT RUN** — each specialist is evaluated only on its own training-family split |
43
- | 19 | **Human evaluation** | **NOT RUN** |
44
- | 20 | **Robustness / adversarial evaluation** | **NOT RUN** |
45
- | 21 | **BigEarthNet label semantics** | the local subset is **100 % single-label** vs the official 1–11 multi-label scheme, so its metrics are **not comparable** to published numbers |
46
- | 22 | **Reliability diagram** | the shipped diagram is the **pre-scaling** curve (labelled as such); the calibrated curve is not plotted |
47
- | 23 | **Statistical significance for most metrics** | only the grounding resolution decision (448 vs 224) has a paired test with a confidence interval. Other per-task numbers are point estimates. |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
48
 
49
  ## 4. Operational limitations
50
 
51
- | # | Limitation | State |
52
- |---|---|---|
53
- | 24 | **B-07: transient tunnel gaps** | a request can hang or return 504. Patch prepared, **NOT deployed**. **OPEN** |
54
- | 25 | **B-07 root shape** | in `auto` mode a tunnel timeout falls through to the forward path, burning `wake_timeout_s` (120 s) on a `302`; worst case ≈ **249 s** (150 + 120). Measured. |
55
- | 26 | **Cold start is tens of seconds** | Render free tier sleeps; the Codespace may be stopped. Documented, not hidden. |
56
- | 27 | **B-02: `codespace_name` trailing newline** | cosmetic; the wake path strips it. **OPEN (cosmetic)** |
57
- | 28 | **No database, auth, or queue** | stateless gateway **by design** — no persistence of runs or users. |
58
- | 29 | **Deployment repos are private** | their links 404 for an outside audience. **BY DESIGN** |
59
- | 30 | **`deploy/` in the monorepo is stale/untracked** | not the deployed source. Trap. |
60
- | 31 | **No APM, distributed tracing, or cost accounting** | observability is limited to the health payload and per-run traces. |
61
- | 32 | **Single-region, no HA** | one Render service, one Codespace. |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
62
 
63
  ## 5. Packaging and licensing
64
 
65
- | # | Limitation | State |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
66
  |---|---|---|
67
- | 33 | **No `LICENSE` file** | none exists in the source repository. **OPEN** — a licence must be selected before public release of the code. |
68
- | 34 | **Backbones are not redistributed** | fetched from the Hub at run time; their licences are their own. |
69
- | 35 | **Artifact bloat** | `artifacts/` is ~3.7 GB, mostly reproducible caches, a duplicate 774 MB ZIP, a duplicate 241 MB probe, two 231 MB feature caches, a superseded 276 MB head, and ~235 MB of selection manifests — not released weights. |
70
- | 36 | **The released weights require their backbones** | the six artifacts are small modules; a consumer must also fetch the pinned backbones. |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
71
 
72
  ## 6. Documentation caveats
73
 
74
- | # | Caveat | State |
75
- |---|---|---|
76
- | 37 | **`docs/FINAL_DELIVERY_REPORT.md` §6 is stale** | it still lists the bundled EO change pair as DEGRADED and B-01 as BLOCKED; both were resolved on 2026-09-25. |
77
- | 38 | **The monorepo `README.md` was materially stale** | it described a hermetic frontend, an in-progress Render/Codespace, and a `/v1/*` contract. Superseded by this release's README. |
78
- | 39 | **`hf/` docs were stale** | they asserted the project owns no weights and has no HF credentials. Both were false at release time. |
79
- | 40 | **Anatomy plate image variant** | points at the 720×720 variant of an image the recorded run analysed at 730×730. Cosmetic. |
80
- | 41 | **The original master plan describes a superseded deployment** | it specifies a Gradio GUI + HF Space + ZeroGPU + Railway. The shipped system is a static frontend + Render + Codespace tunnel, serving JSON. |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
81
 
82
- ## 7. Explicit non-claims
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
83
 
84
  These are things a reader might reasonably assume, which this project does **not** claim:
85
 
@@ -87,22 +713,33 @@ These are things a reader might reasonably assume, which this project does **not
87
  - **No claim of production readiness for model quality.** The deployment runs; the models carry the
88
  limitations above.
89
  - **No claim that the trained heads generalise** beyond their training-family test splits.
90
- - **No claim that calibration improves confidence** — it made ECE worse.
91
- - **No claim that the VLM adapter is accepted** for production use.
92
  - **No claim of an end-to-end accuracy number** — none exists.
93
- - **No claim that the router is correct on all phrasings** — residuals exist.
 
 
94
  - **No claim of robustness** to adversarial, corrupted, or out-of-distribution inputs.
95
  - **No claim of geolocation accuracy** — grounding boxes are image-relative, not geodetic.
96
  - **No claim that the system is a safety-, legal-, or life-critical tool.**
 
 
 
 
97
 
98
- ## 8. Where the evidence lives
99
 
100
  | Topic | Evidence |
101
  |---|---|
102
  | All measured metrics | `artifacts/**/*.json`, verified by `release/tools/verify_readme_metrics.py` |
103
  | Metric honesty rules | [`BENCHMARKS.md`](BENCHMARKS.md), [`EVALUATION.md`](EVALUATION.md) |
104
  | The grounding resolution rejection | `docs/PHASE7_RESOLUTION_DECISION.md` |
105
- | The VLM rejection | `artifacts/vlm/phase6_closure.json` (`why_acceptance_rejected`) |
106
- | B-07 / B-02 | [`DEPLOYMENT.md`](DEPLOYMENT.md) §8.1, §5 |
107
- | Router residuals | [`architecture/04-router.md`](architecture/04-router.md) |
108
- | Environment traps | [`REPRODUCIBILITY.md`](REPRODUCIBILITY.md) |
 
 
 
 
 
 
4
  know. Negative results and open items are listed here rather than omitted, because a limitation that
5
  is not written down is a limitation that will be discovered by someone else at the worst moment.
6
 
7
+ **Status tags:** `OPEN` · `NOT RUN` · `REJECTED` · `DEFERRED` · `BY DESIGN` · `MEASURED`.
8
+
9
+ **How to read this document.** Limitations are numbered `L-01 …` and grouped by area. Each entry
10
+ states the limitation, the measured or observed detail behind it, and the file that records it. Where
11
+ a value is a status, it is stated exactly as the project's own records state it — a `REJECTED` is never
12
+ softened to "usable", an `OPEN` ruling is never presented as settled, and a validation number is never
13
+ promoted to a test result.
14
+
15
+ **Companion documents.** [`BENCHMARKS.md`](BENCHMARKS.md) and [`EVALUATION.md`](EVALUATION.md) hold
16
+ the metric honesty rules; [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) holds the findings behind several
17
+ entries here; [`DEPLOYMENT.md`](DEPLOYMENT.md) §8/§13 holds the operational blockers B-07 and B-02;
18
+ [`architecture/04-router.md`](architecture/04-router.md) documents the router residuals.
19
 
20
  ---
21
 
22
  ## 1. Model quality
23
 
24
+ ### L-01 — Grounding IoU is low in absolute terms (`MEASURED`)
 
 
 
 
 
 
 
 
 
 
 
25
 
26
+ The grounding head's mean best IoU is **0.2838** (canonical protocol) and **0.2566** (matched6
27
+ protocol) on VRSBench eval, n = 16,159. The trained head clearly beats the zero-shot baseline
28
+ (**0.0972**), but 0.28 is not "solved". Recall@0.5 is only **0.2198** (canonical) / **0.1938**
29
+ (matched6).
30
 
31
+ **Evidence:** `artifacts/grounding/remoteclip_grounding_v001/eval_result_canonical.json`,
32
+ `eval_result_matched6.json`; `docs/PHASE7_RESOLUTION_DECISION.md`; `docs/FINAL_DELIVERY_TODO.md` §1.6.
33
+
34
+ ### L-02 — Grounding is protocol-sensitive; an absolute value is meaningless without its protocol (`MEASURED`)
35
+
36
+ The same head reports very different numbers under different decode variants:
37
+
38
+ | Variant | Protocol | mean best IoU | recall@0.5 |
39
  |---|---|---|---|
40
+ | `head_threshold` | canonical | 0.2838 | 0.2198 |
41
+ | `head_threshold` | matched6 | 0.2566 | 0.1938 |
42
+ | `head_argmax` | canonical | 0.1215 | 0.0795 |
43
+ | `zero_shot_matched` | canonical | 0.0972 | 0.0234 |
44
+
45
+ So a single grounding number quoted alone is misleading: the head/threshold versus head/argmax split
46
+ changes IoU by more than a factor of two. **Never quote one without the other.**
47
+
48
+ **Evidence:** `artifacts/grounding/remoteclip_grounding_v001/eval_result_canonical.json` and
49
+ `eval_result_matched6.json`; `release/DOCS_STYLE_GUIDE.md` §3.
50
+
51
+ ### L-03 — The grounding head did not beat the zero-shot baseline on the validation curve (`MEASURED`)
52
+
53
+ The Benchmark page carries this as an honest label: the grounding head did not beat the baseline on
54
+ validation (`docs/FINAL_DELIVERY_TODO.md` §4 P6-T01). The head's advantage over zero-shot is a
55
+ **test-split** result (0.2838 vs 0.0972), not a validation result.
56
+
57
+ **Evidence:** `docs/FINAL_DELIVERY_TODO.md` §4 P6-T01; `docs/PHASE7_RESOLUTION_DECISION.md` ("the
58
+ Phase 8 head must beat 0.0972 to justify itself").
59
+
60
+ ### L-04 — Optical-SAR accuracy is carried by common classes; macro-F1 is low (`MEASURED`, ruling `OPEN`)
61
+
62
+ The fusion head scores accuracy **0.931** but macro-F1 **0.434161** on a held-out test split of
63
+ n = 4,000 over 19 classes. **5 of the 19 classes are absent in the scored split** (`classes_absent:
64
+ [1, 11, 14, 15, 16]`) and contribute **0.0** to macro-F1 by construction
65
+ (`macro_f1_denominator: "all 19 classes (absent classes contribute 0.0)"`). The wide accuracy–macro-F1
66
+ gap is the signature of class imbalance. **Never quote accuracy without macro-F1.**
67
+
68
+ **Evidence:** `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`;
69
+ `docs/FINAL_DELIVERY_TODO.md` §1.6.
70
+
71
+ ### L-05 — The optical-SAR ruling is `OPEN` (`OPEN`)
72
+
73
+ `pre_registered_115_metric.json` states plainly: *"This tool reports ONE head's held-out accuracy and
74
+ macro-F1. It selects no head, ranks nothing and compares no arms. Whether this constitutes a Phase 12
75
+ pass is the owner's ruling."* Phase 12 is **INCOMPLETE**; the pre-registered metric's gate criterion is
76
+ owner-gated (`docs/PHASE12_CURRENCY_CORRECTION.md` §0, §3; `docs/STEP7_BACKEND_CHAIN_REPORT.md` §17).
77
+
78
+ **Evidence:** `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`
79
+ (`is_deciding_statistic: false`); `docs/PHASE12_CURRENCY_CORRECTION.md`.
80
+
81
+ ### L-06 — Change-VQA is weak on rare classes, on two test sets (`MEASURED`, ruling `OPEN`)
82
+
83
+ Change-VQA scores **accuracy 0.697626 / macro-F1 0.378373** on `Test` (n = 39,686) and
84
+ **0.651469 / 0.372309** on `Test2` (n = 31,036). The wide accuracy–macroF1 gap is the signature of
85
+ class imbalance. The ruling is **`OPEN`** (`docs/PHASE19_FINAL_HARDENING.md` §4.5: "the R-02
86
+ macro-F1/accuracy gap is unruled … not this work order's to rule"). Top-3 accuracy is 0.964698. The
87
+ head's metadata records `confidence_method: "uncalibrated"` and `test_splits_used: false` on the
88
+ training record itself.
89
+
90
+ **Evidence:** `artifacts/change_vqa/run/PROMOTION.json` (`test_accuracy`, `test_macro_f1`,
91
+ `test2_accuracy`); `artifacts/change_vqa/run/model_metadata.json`; `docs/PHASE19_FINAL_HARDENING.md`
92
+ §4.5.
93
+
94
+ ### L-07 — VQA is weak-but-related (`MEASURED`)
95
+
96
+ The live VQA path answers broadly related content rather than a crisp class. Measured live: the case
97
+ A1 query answered **"Grassland"** for a scene where a more specific answer was expected
98
+ (`LIVE_VALIDATION_POSTFIX.md`, "Model-quality note"; `release/CURRENT_RELEASE_STATE.md` §6). This is a
99
+ model-quality limitation, not a deployment fault — the pipeline dispatches and returns a real answer.
100
+
101
+ **Evidence:** `.workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md`.
102
+
103
+ ### L-08 — Optical-SAR returns a bare class index, not a human label (`MEASURED`)
104
+
105
+ The live service returns `class_18` rather than a human-readable CLC label. The live answer reads
106
+ `[optical_sar] Fused optical-SAR prediction: class_18 (margin 1.000; optical channels 4/12, SAR
107
+ channels 2/2)` — the modality accounting confirms the right channels reached the fusion head, but the
108
+ answer is not interpretable without a label map.
109
+
110
+ **Evidence:** `release/CURRENT_RELEASE_STATE.md` §6;
111
+ `.workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md`.
112
+
113
+ ### L-09 — Calibration made ECE worse (`MEASURED`)
114
+
115
+ Temperature scaling moved ECE from **0.013755 → 0.014929** (`ece_improvement: −0.001174` — **worse**)
116
+ while improving NLL marginally (0.689741 → 0.689631). It is **retained only because it is part of the
117
+ frozen config** — **not** because it helped. The fitted temperature is `0.9772731820958189`, fitted on
118
+ the Val split with n = 16,441, `effective: true`, `hit_bound: false`. The path is live
119
+ (`core/controller.py` calls `load_calibration(config)` and hands the artifact to `EvidenceEngine`), so
120
+ a deployed result carries a calibrated value — but the calibration is **not an improvement** and must
121
+ never be described as making confidence "more accurate".
122
+
123
+ **Evidence:** `artifacts/calibration_v001.json` (`metrics.ece_before`, `metrics.ece_after`,
124
+ `fit_diagnostics.temperature`); `docs/STEP7_BACKEND_CHAIN_REPORT.md` §7;
125
+ `docs/PHASE19_FINAL_HARDENING.md` §9 (calibration-success correction).
126
+
127
+ ### L-10 — The VLM adapter is `ACCEPTANCE-REJECTED` (`REJECTED`)
128
+
129
+ The Phase-6 LoRA adapter's metrics are **usable** — test exact-match **0.963**, F1 **0.96432**,
130
+ aggregate test delta **+49.5 pp** — and the artifact is `USABLE_VERIFIED`. But its acceptance status is
131
+ **`REJECTED`**: v001 rejected it on val, and the independent-test rule v002 rejected it on the test
132
+ split (one class, `Mixed forest`, lost 4 questions at z = 2.1335). The closure record keeps the two
133
+ questions separate: *"'Verified' answers: is this artifact the one we trained, and does it work?
134
+ 'Accepted' answers: did it clear the bar predeclared before we looked?"* The deployed caption/VQA path
135
+ uses the **unadapted** model by default; the adapter is attached only when `SATQUERY_VLM_ADAPTER` is
136
+ set. **USABLE ≠ ACCEPTED.**
137
+
138
+ **Evidence:** `artifacts/vlm/phase6_closure.json` (`production_adapter.acceptance_status`,
139
+ `why_acceptance_rejected`, `what_closure_does_not_claim`);
140
+ `docs/PHASE6_RUN1_REJECTION_DIAGNOSIS.md`.
141
+
142
+ ### L-11 — The VLM acceptance-rule history is a post-hoc rule change (`MEASURED`, disclosed)
143
+
144
+ v001 was pre-registered (`declared_before_training: true`) and rejected the run on a per-class guardrail
145
+ whose 1.0 pp threshold sits **below the measurement resolution** of the data (0.20 questions at n = 20;
146
+ per-class SE 2.6–10.6 pp). v002 was declared **after** run 1 (`declared_before_training: false`) and is
147
+ documented as a **relaxation** of V2, justified measurement-theoretically, not by the outcome. The
148
+ record keeps v001 and its `REJECTED` verdict verbatim and states that v002 "is not a numerically
149
+ stricter bar" than the contract's ~3 SE figure. A reader must treat the acceptance verdict as resting
150
+ on 4 questions in one class of 33 — the "unfloored minimum-size exposure" the closure reports but does
151
+ not resolve.
152
+
153
+ **Evidence:** `docs/PHASE6_RUN1_REJECTION_DIAGNOSIS.md` §4, §7.1, §7.4;
154
+ `artifacts/vlm/phase6_closure.json` (`thresholds_used.amendment`, `residual_risk`).
155
+
156
+ ### L-12 — The router number is a validation result, ungated, n = 86 (`MEASURED`, `TEST NOT RUN`)
157
+
158
+ Router task accuracy is **0.965116** — measured on **validation**, **ungated**, **n = 86**,
159
+ **corpus-limited** (`corpus_total: 576`, `corpus_groups: 54`; `corpus_limited: true`). The **test
160
+ split was NOT RUN**. The number is indicative only and must never be quoted as a test result.
161
+
162
+ **Evidence:** `artifacts/router/router_adapter_v001/metadata.json` (`val_task_accuracy`);
163
+ `artifacts/router/threshold_sweep_val.json`; `docs/FINAL_DELIVERY_TODO.md` §1.6.
164
+
165
+ ### L-13 — The router's routing is not perfect (known residuals) (`MEASURED`)
166
+
167
+ See §2. The router is a lexical/embedding classifier over a small corpus; residual misroutes exist and
168
+ are documented rather than hidden.
169
+
170
+ ---
171
+
172
+ ## 2. Router residuals (known misroutes)
173
+
174
+ ### L-14 — *"What is the new runway?"* reads `change`, not `vqa` (`OPEN`)
175
+
176
+ The `new`-as-change heuristic fires on non-`where` questions. This is "strictly better than pre-fix,
177
+ where `new` was unconditionally temporal. A lexical router cannot cleanly separate 'the new X' from
178
+ 'what's new'."
179
+
180
+ **Evidence:** `.workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md` ("Known residuals");
181
+ `release/CURRENT_RELEASE_STATE.md` §6.
182
+
183
+ ### L-15 — *"How much built-up area was added?"* reads `vqa` (under-trigger) (`OPEN`)
184
+
185
+ `built` was dropped from the temporal set during the B-08 fix and `area` no longer matches inside
186
+ `areas`, so a change-style quantifier is not caught and the query under-triggers to `vqa`.
187
+
188
+ **Evidence:** `.workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md` ("Known residuals");
189
+ `docs/FINAL_DELIVERY_TODO.md` §5 B-08.
190
+
191
+ ### L-16 — The `interpret()` / `chooseTask()` asymmetry is visually surprising (`RESOLVED`, intentional)
192
+
193
+ For *"What changed between the earlier and later image?"* with **one** asset attached, the console
194
+ **reads** `change` while dispatch correctly falls back to **`change_vqa`**. This is **intentional** —
195
+ the reading is asset-count-blind (it describes the question's intent), while dispatch is
196
+ asset-count-aware (it respects what can actually be computed with the assets present) — but a reader
197
+ who sees the reading panel and the answer disagree may mistake it for a defect.
198
+
199
+ **Evidence:** [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) §7;
200
+ `.workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md` ("Discriminator note").
201
+
202
+ ---
203
 
204
  ## 3. Evaluation gaps
205
 
206
+ ### L-17 — No system-level end-to-end benchmark exists (`NOT RUN`)
207
+
208
+ **No end-to-end accuracy is claimed anywhere.** `STATUS.md` states there is no system-level E2E
209
+ benchmark, and the Benchmark page is required to label it `NOT RUN`.
210
+
211
+ **Evidence:** `docs/FINAL_DELIVERY_TODO.md` §1.7 item 7, §5 B-04; `docs/FINAL_DELIVERY_REPORT.md` §6.
212
+
213
+ ### L-18 — The router test split was not run (`NOT RUN`)
214
+
215
+ See L-12. The test split exists but was never scored.
216
+
217
+ **Evidence:** `docs/FINAL_DELIVERY_TODO.md` §1.6; `docs/FINAL_DELIVERY_REPORT.md` §6.
218
+
219
+ ### L-19 — The benchmark adapters were not run (`NOT RUN`)
220
+
221
+ No benchmark adapter is registered; the registry is empty by construction
222
+ (`docs/PHASE19_FINAL_HARDENING.md` §9, "Benchmark pass — Not claimed").
223
+
224
+ **Evidence:** `docs/PHASE19_FINAL_HARDENING.md` §9; `docs/FINAL_DELIVERY_REPORT.md` §6.
225
+
226
+ ### L-20 — No end-to-end latency benchmark (`NOT RUN`)
227
+
228
+ Per-specialist latency is recorded only incidentally (e.g. grounding encoder latency 2.205 ms/image at
229
+ 224 in the canonical eval, 20.0 ms/image on a T4 per `docs/PHASE7_RESOLUTION_DECISION.md`). There is no
230
+ benchmark of the deployed request path across the four tiers.
231
+
232
+ **Evidence:** `artifacts/grounding/remoteclip_grounding_v001/eval_result_canonical.json`;
233
+ `docs/PHASE7_RESOLUTION_DECISION.md`; [`PERFORMANCE.md`](PERFORMANCE.md).
234
+
235
+ ### L-21 — No cross-dataset generalisation (`NOT RUN`)
236
+
237
+ Each specialist is evaluated only on its own training-family test split (LEVIR-CD for change,
238
+ VRSBench for grounding, a BigEarthNet-family split for fusion, CDVQA for change-VQA). Nothing measures
239
+ transfer to a different distribution.
240
+
241
+ **Evidence:** `docs/PHASE7_RESOLUTION_DECISION.md` ("Anything about hidden ISRO/SAC imagery … a
242
+ different distribution entirely"); the per-task evaluation sections of [`EVALUATION.md`](EVALUATION.md).
243
+
244
+ ### L-22 — No human evaluation (`NOT RUN`)
245
+
246
+ No human study of answer quality, usefulness, or failure modes was performed.
247
+
248
+ **Evidence:** `docs/FINAL_DELIVERY_REPORT.md` §6 (unverified items); this document is the only
249
+ catalogue.
250
+
251
+ ### L-23 — No robustness or adversarial evaluation (`NOT RUN`)
252
+
253
+ No evaluation of behaviour under adversarial, corrupted, or out-of-distribution inputs. The
254
+ change specialist carries invalid-data and registration-quality logic
255
+ (`docs/ARCHITECTURE_FREEZE.md` §23 false-change handling), but no robustness *evaluation* exists.
256
+
257
+ **Evidence:** `docs/FINAL_DELIVERY_REPORT.md` §6.
258
+
259
+ ### L-24 — BigEarthNet label semantics make the local metrics non-comparable (`MEASURED`)
260
+
261
+ The local BigEarthNet subset is **100 % single-label** against the official **1–11 multi-label**
262
+ scheme, so metrics computed on it are **not comparable** to published multi-label numbers.
263
+
264
+ **Evidence:** [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) §9; `docs/PHASE12_LABEL_POLICY_DECISION.md`.
265
+
266
+ ### L-25 — The reliability diagram shipped is the pre-scaling curve (`MEASURED`)
267
+
268
+ The shipped reliability curve is the **pre-scaling** curve (labelled as such); the calibrated curve is
269
+ not plotted. The caption now reads "Measured" and the bins are transcribed from
270
+ `artifacts/calibration_v001.json`, but the remaining section-03 PR curves are still labelled
271
+ illustrative.
272
+
273
+ **Evidence:** `docs/FINAL_DELIVERY_TODO.md` §4 P6-T01 note; `artifacts/calibration_v001.json`.
274
+
275
+ ### L-26 — Statistical significance exists for only one decision (`MEASURED`)
276
+
277
+ Only the grounding resolution decision (448 vs 224) has a paired test with a confidence interval
278
+ (§[`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) §3). Every other per-task number is a point estimate with no
279
+ significance test.
280
+
281
+ **Evidence:** `docs/PHASE7_RESOLUTION_DECISION.md` ("Paired analysis — independent confirmation").
282
+
283
+ ### L-27 — The optical-SAR registry state is `degraded`, and the adapter never advertises `loaded` (`MEASURED`)
284
+
285
+ The registry resolves `optical_sar` to **`degraded`**, not `available`
286
+ (`docs/STEP7_BACKEND_CHAIN_REPORT.md` §17). More generally, the capability adapter derives contract
287
+ state from **artifact presence** and therefore **never emits `loaded` or `evicted`** — "a model is
288
+ resident" is unknowable without loading one, which the metadata path must not do
289
+ (`docs/DEPLOYMENT_ARCHITECTURE.md` §3.3.1). A consumer cannot learn from `/v1/capabilities` whether a
290
+ model is actually resident.
291
+
292
+ **Evidence:** `docs/STEP7_BACKEND_CHAIN_REPORT.md` §17; `docs/DEPLOYMENT_ARCHITECTURE.md` §3.3.1.
293
+
294
+ ### L-28 — The trace's registry block is a snapshot of the *previous* request (`MEASURED`)
295
+
296
+ `SpecialistRegistry.describe()` is assigned into the trace **immediately after planning** and **before
297
+ execution**, so on a cold process the client-visible `built` map is necessarily `{}` — including for
298
+ the request that is about to build the specialist. The same query returns two different traces
299
+ depending on how many requests the process has already served. The same trace is also internally
300
+ inconsistent about which clock it uses (`selected_models` is assigned inside execution, so it *does*
301
+ reflect the current request while `built` does not). **Severity: low — observability only.**
302
+
303
+ **Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §5.8 (finding F-19).
304
+
305
+ ### L-29 — The uploaded-asset path surface has historically leaked server-side paths (`RESOLVED`, recorded)
306
+
307
+ A series of findings (F-13 … F-16) recorded that `trace.inputs`, `trace.steps[PARSE].detail`, and
308
+ `evidence[].artifact_ref` / `result.change_map` published server-side filesystem paths to an
309
+ unauthenticated client. All were fixed (paths reduced to basenames; refs set to `null` with an explicit
310
+ non-retrievable warning; no fabricated `artifact://` URIs). Two consequences remain worth recording:
311
+ the contract's §2.4 example once showed a fabricated `artifact://` URI the service cannot emit (now
312
+ corrected), and the ruling introduced a **signalling** change (F-16c): on a deployment that configures
313
+ `change.artifact_dir`, a no-change run now reports `degraded: true` where it previously reported
314
+ `degraded: false`.
315
+
316
+ **Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §5.3–§5.6; [`SECURITY.md`](SECURITY.md).
317
+
318
+ ---
319
 
320
  ## 4. Operational limitations
321
 
322
+ ### L-30 — B-07: transient tunnel gaps (`OPEN`)
323
+
324
+ A request can hang or return `504` when the tunnel agent is briefly absent. The patch is prepared,
325
+ **NOT deployed**.
326
+
327
+ **Evidence:** [`DEPLOYMENT.md`](DEPLOYMENT.md) §8.1; `docs/FINAL_DELIVERY_TODO.md` §5 B-07;
328
+ `release/CURRENT_RELEASE_STATE.md` §6.
329
+
330
+ ### L-31 — B-07 root shape: the `auto`-mode fallthrough wastes the wake budget (`MEASURED`)
331
+
332
+ In `auto` transport mode a tunnel timeout **falls through** to the forward path, which then burns
333
+ `wake_timeout_s` (120 s) on a `302` → a worst case of ≈ **249 s** (150 + 120). Measured.
334
+
335
+ **Evidence:** `release/CURRENT_RELEASE_STATE.md` §6; [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) §6;
336
+ [`DEPLOYMENT.md`](DEPLOYMENT.md) §8.1.
337
+
338
+ ### L-32 — Cold start is tens of seconds (`MEASURED`)
339
+
340
+ Render's free tier sleeps when idle, and the Codespace may be stopped (idle timeout 30 min). The first
341
+ request after idle waits for a wake. Documented, not hidden.
342
+
343
+ **Evidence:** [`DEPLOYMENT.md`](DEPLOYMENT.md) §8; `docs/DEPLOYMENT_TOPOLOGY.md` §2.
344
+
345
+ ### L-33 — B-02: `codespace_name` trailing newline (`OPEN`, cosmetic)
346
+
347
+ The `/api/health` payload reports the raw `codespace_name` with a trailing `\n`. Cosmetic; the wake
348
+ path strips it.
349
+
350
+ **Evidence:** [`DEPLOYMENT.md`](DEPLOYMENT.md) §8.2; `docs/FINAL_DELIVERY_TODO.md` §4 P2-T03.
351
+
352
+ ### L-34 — No database, auth, or queue (`BY DESIGN`)
353
+
354
+ The gateway is stateless by design. There is no persistence of runs or users, no auth layer, and no
355
+ request queue (plan §73/§74; `docs/DEPLOYMENT_ARCHITECTURE.md` §6).
356
+
357
+ **Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §2.2, §6.
358
+
359
+ ### L-35 — The deployment repositories are private (`BY DESIGN`)
360
+
361
+ The three deploy repositories return `404` for an outside audience.
362
+
363
+ **Evidence:** [`DEPLOYMENT.md`](DEPLOYMENT.md) §11.4; `docs/FINAL_DELIVERY_TODO.md` §4 P9-T01.
364
+
365
+ ### L-36 — `deploy/` in the monorepo is stale and untracked (`OPEN` trap)
366
+
367
+ Edits there do not deploy; it is not the deployed source.
368
+
369
+ **Evidence:** [`DEPLOYMENT.md`](DEPLOYMENT.md) §1.1; `docs/FINAL_DELIVERY_TODO.md` §1.1.
370
+
371
+ ### L-37 — The rate limiter is a fairness control, not a protection control (`BY DESIGN`)
372
+
373
+ The per-IP limiter keys on the first hop of the client-supplied `X-Forwarded-For`; a caller that varies
374
+ the header is never throttled (measured 0/8 throttled with a fresh value per request, versus 5/8
375
+ without). It is explicitly **not** a security boundary.
376
+
377
+ **Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §5.2; `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.1.2.
378
+
379
+ ### L-38 — No APM, distributed tracing, or cost accounting (`BY DESIGN`)
380
+
381
+ Observability is limited to the health payload and per-run traces. There is no APM, no distributed
382
+ tracing across the four tiers, and no cost accounting.
383
+
384
+ **Evidence:** [`architecture/10-observability-and-ops.md`](architecture/10-observability-and-ops.md);
385
+ [`OPERATIONS.md`](OPERATIONS.md).
386
+
387
+ ### L-39 — Single-region, no HA (`BY DESIGN`)
388
+
389
+ One Render service, one Codespace. No redundancy, no failover, no multi-region deployment.
390
+
391
+ **Evidence:** [`DEPLOYMENT.md`](DEPLOYMENT.md) §13.
392
+
393
+ ### L-40 — A saturated asset store is indistinguishable from a misconfigured one (`OPEN`)
394
+
395
+ `POST /v1/assets` answers `503` both when the store is unconfigured and when it is full; the response
396
+ cannot tell them apart, and the counter that would have separated them was **removed** rather than given
397
+ a consumer. Treat the `503` on this route as ambiguous.
398
+
399
+ **Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §5.1 (finding F-11).
400
+
401
+ ### L-41 — `change_vqa.artifact_dir` is a configured-but-inert key (`OPEN`, low severity)
402
+
403
+ `change_vqa.artifact_dir` is advertised in the same `optional_config_keys` table as the two keys that
404
+ work, but `specialists/change/vqa_specialist.py` **never reads it** (the attribute occurs exactly once,
405
+ as an assignment; the module contains no file-writing code). An operator who sets it receives no
406
+ artifacts and **no warning**.
407
+
408
+ **Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §5.6 (finding F-17).
409
+
410
+ ### L-42 — The Anatomy page's plate image variant is cosmetic-wrong (`OPEN`, cosmetic)
411
+
412
+ The Anatomy plate points at the **720×720** variant of an image the recorded run analysed at
413
+ **730×730**. The content is identical and the canvas scales it, but the page's "the ACTUAL analysed
414
+ image" wording is very slightly loose.
415
+
416
+ **Evidence:** `.workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md` ("Known residuals");
417
+ `release/CURRENT_RELEASE_STATE.md` §6.
418
+
419
+ ### L-43 — The `/api/capabilities` deployment block carried stale metadata (`RESOLVED`)
420
+
421
+ The capabilities `deployment` block once claimed `huggingface-spaces`/`zerogpu`. It was recorded as a
422
+ stale-metadata item (`docs/FINAL_DELIVERY_TODO.md` §1.7 item 6) and is reconciled in the shipped
423
+ contract; the frozen `configs/deploy.yaml` still describes the superseded target (§5 below).
424
+
425
+ **Evidence:** `docs/FINAL_DELIVERY_TODO.md` §1.7 item 6; [`DEPLOYMENT.md`](DEPLOYMENT.md) §11.
426
+
427
+ ---
428
 
429
  ## 5. Packaging and licensing
430
 
431
+ ### L-44 — There is no `LICENSE` file (`OPEN`)
432
+
433
+ **No `LICENSE` file exists** in the source repository. The README says "add a license file before
434
+ public release". A licence must be selected before public release of the code. This is a **release
435
+ blocker for the code**, not a model defect.
436
+
437
+ **Evidence:** `release/CURRENT_RELEASE_STATE.md` §6 ("**No `LICENSE` file exists** in the monorepo;
438
+ README says 'add a license file before public release'"); `release/DOCS_STYLE_GUIDE.md` §3.
439
+
440
+ ### L-45 — Backbones are not redistributed (`BY DESIGN`)
441
+
442
+ The six trained artifacts are small modules; the backbones (SmolVLM, RemoteCLIP, MiniLM, CROMA,
443
+ STANet-style encoder) are fetched from their sources at run time, and their licences are their own.
444
+ Nothing here re-licenses a backbone.
445
+
446
+ **Evidence:** [`MODELS.md`](MODELS.md); `app/serving.py` (`resolve_checkpoint_path`, "offline first");
447
+ `docs/ARCHITECTURE_FREEZE.md` §2.
448
+
449
+ ### L-46 — The released weights require their backbones (`MEASURED`)
450
+
451
+ A consumer of the six artifacts must also fetch the pinned backbones at the exact revisions; an
452
+ artifact alone is not runnable. The VLM adapter, for example, requires
453
+ `HuggingFaceTB/SmolVLM-500M-Instruct` (revision `a7da5b986cb5`) plus the tokenizer/processor files
454
+ listed in `phase6_closure.json`.
455
+
456
+ **Evidence:** `artifacts/vlm/phase6_closure.json` (`production_adapter.files_required_to_serve_standalone`);
457
+ `release/CURRENT_RELEASE_STATE.md` §3.
458
+
459
+ ### L-47 — Artifact-tree bloat (`MEASURED`)
460
+
461
+ `artifacts/` totals roughly **3.7 GB**, almost all of it caches, duplicates and features rather than
462
+ released weights:
463
+
464
+ | Path | Size | Classification |
465
  |---|---|---|
466
+ | `artifacts/grounding/remoteclip_grounding_v001/` | 1.4 GB | ARCHIVED (evidence) |
467
+ | `artifacts/grounding/remoteclip_grounding_v001.zip` | 774 MB | **DUPLICATE** of the directory above |
468
+ | `artifacts/optical_sar/fusion_head_v001/` | 276 MB | superseded by `fusion_head_production_v001` |
469
+ | `artifacts/optical_sar/fusion_features/` | 231 MB | reproducible cache |
470
+ | `artifacts/optical_sar/fusion_features_armB/` | 231 MB | reproducible cache |
471
+ | `artifacts/change/levir_change_cpu_probe_v001/` | 241 MB | **DUPLICATE** probe of `levir_change_v001` |
472
+ | `artifacts/change/levir_change_v001/` | 241 MB | ARCHIVED |
473
+ | `artifacts/phase12_selection/*.jsonl` | ~235 MB | data-selection manifests (seeds) |
474
+
475
+ **Evidence:** `release/CURRENT_RELEASE_STATE.md` §3.
476
+
477
+ ### L-48 — The VLM adapter is not committed and lives under `.scratch/` (`MEASURED`)
478
+
479
+ The adapter is **not** committed (`.gitignore` excludes `artifacts/`, `checkpoints/` and
480
+ `*.safetensors`) and its canonical path is under `.scratch/`, which a future cleanup could remove. The
481
+ closure record states that moving it to a non-scratch location is "a reasonable follow-up, not a
482
+ closure requirement", and that moving it would make the recorded evidence stale.
483
+
484
+ **Evidence:** `artifacts/vlm/phase6_closure.json` (`production_adapter.reconstruction.why_not_moved`).
485
+
486
+ ### L-49 — The frozen `configs/deploy.yaml` describes a target that does not exist (`OPEN` paperwork)
487
+
488
+ `configs/deploy.yaml` still declares `platform: huggingface-spaces`, `sdk: gradio`, `zerogpu: true`,
489
+ and the `gpu_duration_*` values. It is inert (`registry: false`), never loaded by `core/config.py`, and
490
+ cannot be edited without either failing `scripts/validate_deploy_config.py` or moving `Config.hash`. It
491
+ is deliberately left undisturbed, but a reader who finds it will reasonably think the project targets a
492
+ Gradio ZeroGPU Space.
493
+
494
+ **Evidence:** `configs/deploy.yaml`; `docs/DEPLOYMENT_DECISION.md` §4; [`DEPLOYMENT.md`](DEPLOYMENT.md) §11.
495
+
496
+ ---
497
 
498
  ## 6. Documentation caveats
499
 
500
+ ### L-50 — `docs/FINAL_DELIVERY_REPORT.md` §6 is stale (`SUPERSEDED`)
501
+
502
+ It still lists the bundled EO change pair as **DEGRADED** (726² vs 736² → shape error) and **B-01** as
503
+ **BLOCKED**. Both were resolved on 2026-09-25: the EO pair is now same-shape (both 720×720, new `-720`
504
+ URLs) → **RESOLVED**; the Hugging Face link is live on all 11 pages → **B-01 CLOSED**.
505
+
506
+ **Evidence:** `release/CURRENT_RELEASE_STATE.md` §6 ("Documentation that this sprint supersedes");
507
+ `docs/FINAL_DELIVERY_REPORT.md` §6.
508
+
509
+ ### L-51 — The monorepo `README.md` was materially stale (`SUPERSEDED`)
510
+
511
+ It described a hermetic frontend, an in-progress Render/Codespace, a `/v1/*` contract, omitted the
512
+ tunnel, and pointed at the stale `deploy/`. Superseded by this release's README.
513
+
514
+ **Evidence:** `release/CURRENT_RELEASE_STATE.md` §6.
515
+
516
+ ### L-52 — The `hf/` docs were stale (`SUPERSEDED`)
517
+
518
+ `hf/SETUP.md` and `hf/README.md` asserted the project owns no weights and has no HF credentials. Both
519
+ were false at release time.
520
+
521
+ **Evidence:** `release/CURRENT_RELEASE_STATE.md` §6.
522
+
523
+ ### L-53 — Stale-negative documentation is a structural hazard (`MEASURED`)
524
+
525
+ Three documents written after the Phase-12 A/B experiment still described it as **un-run**, even though
526
+ it had completed 34 hours earlier. The lesson recorded: "a stale negative claim is more dangerous than
527
+ a stale positive one … 'This has never been executed' is contradicted by *nothing* — no test fails, no
528
+ hash moves, no invariant breaks." The existing conformance tests check that documented things *exist*;
529
+ **no test can check that a documented absence is still absent**.
530
+
531
+ **Evidence:** `docs/PHASE12_CURRENCY_CORRECTION.md` §0, §6.
532
+
533
+ ### L-54 — The original master plan describes a superseded deployment (`SUPERSEDED`)
534
+
535
+ The master plan specifies a Gradio GUI + HF Space + ZeroGPU + Railway, and a single-image workflow set.
536
+ The shipped system is a static frontend + Render + Codespace tunnel, serving JSON, with the change-VQA
537
+ and optical-SAR capabilities added later. The plan is a design document, not a description of the
538
+ shipped system.
539
+
540
+ **Evidence:** `docs/MASTER_ARCHITECTURE_PLAN.md`; `docs/DEPLOYMENT_DECISION.md` §4;
541
+ [`DEPLOYMENT.md`](DEPLOYMENT.md) §11.
542
 
543
+ ### L-55 — `mask_ref` was wrongly listed as a documented path surface (`CORRECTED`)
544
+
545
+ An audit note listed `Region.mask_ref` as a deliberate documented path surface; that was wrong —
546
+ `core/schemas.py` is a bare `str | None = None` with no description, and `grep -r mask_ref docs/` finds
547
+ nothing. Recorded so the correction is not lost.
548
+
549
+ **Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §5.5 (note on `mask_ref`).
550
+
551
+ ### L-56 — Documentation-validated-by-execution found nine factual errors (`MEASURED`)
552
+
553
+ Validating every JSON example against the real Pydantic models and every runbook claim against the
554
+ repository caught nine factual errors proof-reading had missed — including `CoordinateSystem`
555
+ documented as `normalized`/`geographic` when the real values are `normalized_0_1`/`geo`, and `Box`
556
+ documented with nested geometry when the real model is flat. The checks are now permanent tests (48
557
+ assertions). This is a caveat about how much confidence a *document* can carry: prose is not validated
558
+ by these tests.
559
+
560
+ **Evidence:** `docs/PHASE19_FINAL_HARDENING.md` §3.6; `docs/STEP8_FINAL_CONFORMANCE_AUDIT.md` §13.
561
+
562
+ ---
563
+
564
+ ## 7. Interface, contract, and data-surface limitations
565
+
566
+ ### L-57 — The contract promises forward-compatible reading; the schema forbids it (`OPEN`, contract contradiction)
567
+
568
+ `docs/API_CONTRACT.md` §1.1 says a consumer **"MUST tolerate unknown fields on read (forward
569
+ compatibility)"**, and §1 says additive changes do not bump the version — which only works if unknown
570
+ fields are ignorable. But `ResultEnvelope`, `SpecialistResult` and `ExecutionTrace` are all
571
+ `extra="forbid"`, so `model_validate` **rejects** a body carrying an unknown key. Both halves are
572
+ load-bearing and cannot both hold. The contract's own authority clause makes the **code** authoritative,
573
+ so the document's bullet is the inaccurate half — but the document was **not** corrected (a doc edit is
574
+ a reviewable change), so a reader of the contract is still told the wrong thing.
575
+
576
+ **Evidence:** `docs/STEP7_BACKEND_CHAIN_REPORT.md` §15 (finding C-2);
577
+ [`architecture/08-api-contract.md`](architecture/08-api-contract.md).
578
+
579
+ ### L-58 — The capability adapter can never report `loaded` or `evicted` (`MEASURED`)
580
+
581
+ Because the metadata path must not construct a model, `app/deployment.py` derives contract state from
582
+ artifact **presence** and reconstructs the registry's word from that state — the inverse direction. One
583
+ observable consequence: **`loaded` and `degraded` are never emitted**, since "a model is resident" is
584
+ unknowable without loading one, and `evicted` is a runtime model-cache fact no static inspection can
585
+ observe. A consumer therefore cannot learn from `/v1/capabilities` whether a model is actually
586
+ resident.
587
+
588
+ **Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §3.3.1;
589
+ `docs/STEP8_FINAL_CONFORMANCE_AUDIT.md` §3 (finding C-4).
590
+
591
+ ### L-59 — Two subsystems use overlapping words with different meanings (`MEASURED`)
592
+
593
+ To the registry, `unavailable` means "no builder could be constructed"; to the contract it means
594
+ "present but broken, and **therefore a defect**". Transcribing one into the other without noticing
595
+ "would turn a deployment gap into a reported defect, or vice versa". The registry's `available` is not a
596
+ contract state at all, and `degraded` is *both* a whole-service status and a capability state — a reader
597
+ cannot tell from the word alone which is meant.
598
+
599
+ **Evidence:** `docs/STEP7_BACKEND_CHAIN_REPORT.md` §10 (findings H-2, M-2).
600
+
601
+ ### L-60 — The capability source of truth was historically two producers (`RESOLVED`, recorded)
602
+
603
+ Before the owner ruling, `AnalysisController.health()` enumerated **6** capabilities from the registry
604
+ while the served `describe_deployment()` enumerated **2** from two `Path.exists()` calls. The registry
605
+ is now authoritative and `describe_deployment()` delegates to the single adapter — but the historical
606
+ divergence is recorded because "any table that exists in two places will drift, and this one already
607
+ had".
608
+
609
+ **Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §3.3.1;
610
+ `docs/STEP7_BACKEND_CHAIN_REPORT.md` §10, §16 (finding M-1).
611
+
612
+ ### L-61 — `mask_ref` is an undocumented optional field (`OPEN`, low)
613
+
614
+ `core/schemas.py` declares `mask_ref` as a bare `str | None = None` with no description, and
615
+ `grep -r mask_ref docs/` finds nothing. An earlier audit note wrongly listed it as a deliberate
616
+ documented path surface; that was corrected. It remains undocumented.
617
+
618
+ **Evidence:** `docs/DEPLOYMENT_ARCHITECTURE.md` §5.5 (note on `mask_ref`).
619
+
620
+ ### L-62 — VRSBench box coordinates are normalised to 0–100 (`MEASURED`)
621
+
622
+ VRSBench's repository notes that provided evaluation box coordinates are normalized to **0–100**, so the
623
+ evaluator adapter must explicitly convert between that convention and the project's internal **0–1**
624
+ representation rather than quietly treating the numbers as pixels. A consumer that reads a grounding box
625
+ without the coordinate-system field will misinterpret it.
626
+
627
+ **Evidence:** `docs/MASTER_ARCHITECTURE_PLAN.md` §13 (grounding head / coordinate convention);
628
+ `docs/STEP7_BACKEND_CHAIN_REPORT.md` §4 (`CoordinateSystem` is exactly `{normalized_0_1, pixel, geo}`).
629
+
630
+ ### L-63 — The gateway's rate-limit and size-limit values are implementation choices, not plan facts (`OPEN`)
631
+
632
+ The plan specifies none of them. The defaults in `GatewayConfig` are choices the implementation made,
633
+ and the maintainer is asked to confirm them, because they bound one client's share of the daily budget.
634
+ No rate-limit value is specified by the plan.
635
+
636
+ **Evidence:** `docs/PHASE19_FINAL_HARDENING.md` §4.4; `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.1.
637
+
638
+ ### L-64 — `POST /v1/assets` was not in the original plan (`RESOLVED`, historically undefined)
639
+
640
+ The plan fixes three endpoints, yet `AnalysisRequest.assets` is a list of *handles*, which requires a
641
+ fourth. `docs/API_CONTRACT.md` §2.5 records the gap; the endpoint is now implemented, but the
642
+ fourth-endpoint decision was historically open and the gateway once answered `501` for it.
643
+
644
+ **Evidence:** `docs/PHASE19_FINAL_HARDENING.md` §4.2;
645
+ `docs/STEP7_BACKEND_CHAIN_REPORT.md` §16.
646
+
647
+ ---
648
+
649
+ ## 8. Dataset and corpus limitations
650
+
651
+ ### L-65 — The change model is evaluated only on LEVIR-CD-256 (`MEASURED`)
652
+
653
+ The change head's headline (pooled IoU 0.8122 / macro IoU 0.8457 / pooled F1 0.8964, n = 2048,
654
+ threshold 0.5) is measured on **LEVIR-CD-256** test. Nothing measures transfer to another change
655
+ dataset or to a different sensor pair.
656
+
657
+ **Evidence:** `artifacts/change/eval_test/eval_result.json`; [`EVALUATION.md`](EVALUATION.md).
658
+
659
+ ### L-66 — The change evaluation records a large pixel imbalance (`MEASURED`)
660
+
661
+ The change test split is heavily imbalanced: pooled counts record `tp = 5,978,997`, `fp = 523,658`,
662
+ `fn = 858,407`, `tn = 126,856,666` over `n_pixels = 134,217,728`, with a mean change fraction of
663
+ **0.0509** and only **935 of 2,048** images containing change. A pooled IoU on a 5 % positive pixel rate
664
+ is not the same statistic as a balanced one, and the macro/pooled split (macro IoU 0.718 vs pooled IoU
665
+ 0.8122) reflects that.
666
+
667
+ **Evidence:** `artifacts/change/eval_test/eval_result.json` (`metrics.pooled`, `metrics.macro`,
668
+ `mean_change_fraction`, `n_images_with_change`).
669
+
670
+ ### L-67 — The router corpus is small and corpus-limited (`MEASURED`)
671
+
672
+ The router's training/evaluation corpus is **576** queries in **54** groups (`corpus_total: 576`,
673
+ `corpus_groups: 54`), with a by-task distribution of caption 91, change 115, grounding 128, optical_sar
674
+ 50, unsupported 105, vqa 87. The validation split is **n = 86**. A 0.965116 validation accuracy over 86
675
+ examples, drawn from a 576-query corpus, is **indicative only**; it is not a benchmark result.
676
+
677
+ **Evidence:** `artifacts/router/router_adapter_v001/metadata.json` (`corpus`),
678
+ `artifacts/router/threshold_sweep_val.json`.
679
+
680
+ ### L-68 — The change-VQA test and test2 splits share scenes (`MEASURED`)
681
+
682
+ The change-VQA dataset's integrity block lists `allowed_shared_pairs: [["Test", "Test2"]]` — the two
683
+ test splits legitimately share scenes, so they are **not independent draws** of the same population. A
684
+ number from one is not a confirmation of the other.
685
+
686
+ **Evidence:** `artifacts/change_vqa/run/run_record.json` (`dataset.integrity`).
687
+
688
+ ### L-69 — The optical-SAR metric is on a 19-class space with five absent classes (`MEASURED`)
689
+
690
+ The fusion metric is defined over a **19-class** label space; the scored held-out split contains only 14
691
+ present classes (`classes_present: [0,2,3,4,5,6,7,8,9,10,12,13,17,18]`). The absent five contribute 0.0
692
+ to macro-F1 by construction, which is why accuracy (0.931) and macro-F1 (0.434161) diverge so sharply.
693
+ The metric is a single head's held-out result; it selects no head and compares no arms.
694
+
695
+ **Evidence:** `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`.
696
+
697
+ ### L-70 — The grounding evaluation is a single dataset and a single expression family (`MEASURED`)
698
+
699
+ Grounding is measured only on VRSBench referring expressions (n = 16,159). Nothing measures grounding on
700
+ a different expression style or a different imagery distribution. The `matched6` protocol (top_k = 6) is
701
+ the only protocol variation recorded.
702
+
703
+ **Evidence:** `artifacts/grounding/remoteclip_grounding_v001/eval_result_canonical.json`,
704
+ `eval_result_matched6.json`.
705
+
706
+ ---
707
+
708
+ ## 9. Explicit non-claims
709
 
710
  These are things a reader might reasonably assume, which this project does **not** claim:
711
 
 
713
  - **No claim of production readiness for model quality.** The deployment runs; the models carry the
714
  limitations above.
715
  - **No claim that the trained heads generalise** beyond their training-family test splits.
716
+ - **No claim that calibration improves confidence** — it made ECE worse (0.013755 → 0.014929).
717
+ - **No claim that the VLM adapter is accepted** for production use — it is `ACCEPTANCE-REJECTED`.
718
  - **No claim of an end-to-end accuracy number** — none exists.
719
+ - **No claim that the router is correct on all phrasings** — residuals exist (§2).
720
+ - **No claim that the router number is a test result** — it is validation, ungated, n = 86.
721
+ - **No claim that the optical-SAR or change-VQA rulings are settled** — both are `OPEN`.
722
  - **No claim of robustness** to adversarial, corrupted, or out-of-distribution inputs.
723
  - **No claim of geolocation accuracy** — grounding boxes are image-relative, not geodetic.
724
  - **No claim that the system is a safety-, legal-, or life-critical tool.**
725
+ - **No claim that the code is licensed** — no `LICENSE` file exists.
726
+ - **No claim that the release repositories are public** — three of four are private by design.
727
+
728
+ ---
729
 
730
+ ## 10. Where the evidence lives
731
 
732
  | Topic | Evidence |
733
  |---|---|
734
  | All measured metrics | `artifacts/**/*.json`, verified by `release/tools/verify_readme_metrics.py` |
735
  | Metric honesty rules | [`BENCHMARKS.md`](BENCHMARKS.md), [`EVALUATION.md`](EVALUATION.md) |
736
  | The grounding resolution rejection | `docs/PHASE7_RESOLUTION_DECISION.md` |
737
+ | The VLM rejection | `artifacts/vlm/phase6_closure.json` (`why_acceptance_rejected`), `docs/PHASE6_RUN1_REJECTION_DIAGNOSIS.md` |
738
+ | The optical-SAR ruling | `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json` |
739
+ | The change-VQA ruling | `artifacts/change_vqa/run/PROMOTION.json`, `docs/PHASE19_FINAL_HARDENING.md` §4.5 |
740
+ | The calibration result | `artifacts/calibration_v001.json`, `docs/STEP7_BACKEND_CHAIN_REPORT.md` §7 |
741
+ | B-07 / B-02 | [`DEPLOYMENT.md`](DEPLOYMENT.md) §8.1, §8.2; `docs/FINAL_DELIVERY_TODO.md` §5 |
742
+ | Router residuals | [`architecture/04-router.md`](architecture/04-router.md), `LIVE_VALIDATION_POSTFIX.md` |
743
+ | Environment traps | [`REPRODUCIBILITY.md`](REPRODUCIBILITY.md), [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) §8 |
744
+ | Stale documentation | `release/CURRENT_RELEASE_STATE.md` §6, `docs/PHASE12_CURRENCY_CORRECTION.md` |
745
+ | The stale-claim class | `docs/PHASE12_CURRENCY_CORRECTION.md` §6 |