thundercode commited on
Commit
5a57728
Β·
verified Β·
1 Parent(s): d09af94

release: add docs/DATA_PIPELINE.md

Browse files
Files changed (1) hide show
  1. docs/DATA_PIPELINE.md +1155 -0
docs/DATA_PIPELINE.md ADDED
@@ -0,0 +1,1155 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Data pipeline β€” deep reference
2
+
3
+ **Status tags used on every substantive claim:** `IMPLEMENTED` Β· `VERIFIED` Β· `MEASURED` Β· `ATTEMPTED`
4
+ Β· `NOT RUN` Β· `BLOCKED` Β· `DEFERRED` Β· `REJECTED` Β· `OPEN` Β· `RESOLVED` Β· `CLOSED`.
5
+
6
+ Every statement in this document was produced by reading a file in `C:/Users/anish/satquery-ai/`. Where a
7
+ fact is not established from the available evidence, this document writes
8
+ `UNKNOWN β€” not established from the available evidence`. Where a stage is *declared* in configuration but
9
+ read by no code, that is stated explicitly rather than presented as a working step. Where a cache spec
10
+ hash is cited, it was **recomputed** from the code that produces it, not copied from a report.
11
+
12
+ This chapter is the reference for the **end-to-end data path**: from an uploaded asset to the tensor a
13
+ specialist consumes, including every decode, every normalisation, every cache, and every intermediate
14
+ artefact. It is written to be sufficient to reconstruct the pipeline from the document alone.
15
+
16
+ **Sibling documents:** [`GEOSPATIAL.md`](GEOSPATIAL.md) (the raster contract, CRS, coordinate systems, the
17
+ sensor adapter, the tiling policy and the 224 decision), [`architecture/03-request-lifecycle.md`](architecture/03-request-lifecycle.md)
18
+ (the nine controller states), [`architecture/05-specialists.md`](architecture/05-specialists.md) (the six
19
+ specialists), [`architecture/07-configuration-freeze.md`](architecture/07-configuration-freeze.md) (the
20
+ frozen config identity), [`DATASETS.md`](DATASETS.md) (the corpora), [`REPRODUCIBILITY.md`](REPRODUCIBILITY.md)
21
+ (reproduction recipes), [`TRAINING.md`](TRAINING.md) (how each corpus is consumed).
22
+
23
+ ---
24
+
25
+ ## Table of contents
26
+
27
+ 0. [How to read this document](#0-how-to-read-this-document)
28
+ 1. [The end-to-end path](#1-the-end-to-end-path)
29
+ 2. [Validation and decoding](#2-validation-and-decoding)
30
+ 3. [Modality inference, and where it is consumed](#3-modality-inference-and-where-it-is-consumed)
31
+ 4. [The normalisation stages](#4-the-normalisation-stages)
32
+ 5. [Tiling and tile selection](#5-tiling-and-tile-selection)
33
+ 6. [The change-VQA data path](#6-the-change-vqa-data-path)
34
+ 7. [The optical-SAR data path](#7-the-optical-sar-data-path)
35
+ 8. [The router data path](#8-the-router-data-path)
36
+ 9. [The grounding data path](#9-the-grounding-data-path)
37
+ 10. [The VQA / caption data path](#10-the-vqa--caption-data-path)
38
+ 11. [Determinism and reproducibility](#11-determinism-and-reproducibility)
39
+ 12. [Where intermediate artefacts live](#12-where-intermediate-artefacts-live)
40
+ 13. [What is NOT part of the pipeline](#13-what-is-not-part-of-the-pipeline)
41
+ 14. [What is NOT RUN / OPEN / BLOCKED for this topic](#14-what-is-not-run--open--blocked-for-this-topic)
42
+ 15. [Where the evidence lives](#15-where-the-evidence-lives)
43
+
44
+ ---
45
+
46
+ ## 0. How to read this document
47
+
48
+ The pipeline has **two distinct halves**, and almost every confusion about it comes from conflating them:
49
+
50
+ 1. **The serving path** β€” an uploaded asset becomes a specialist result, request by request. This is what
51
+ runs in production. It is *stateless* with respect to data: it reads a raster, prepares it, and calls
52
+ a model. No cache is consulted except the router's, and no intermediate artefact is required for a
53
+ result to be produced.
54
+ 2. **The preparation / training path** β€” a corpus on disk becomes a *cache* of frozen features, and a head
55
+ is trained over that cache. This runs offline, once per corpus revision. It is where the
56
+ `mean Β± 2Β·std` radiometric stretch actually executes, where the `change_feat_v1` feature vector is
57
+ computed, and where the arm (A/B) is baked in.
58
+
59
+ The caches are the bridge: the training path produces them, and the **serving path does not read them**
60
+ (except the router's corpus cache, which is a training-time artefact, and the change-VQA text cache, which
61
+ is likewise training-time). A reader who assumes the serving path is cache-driven will misread the system;
62
+ the caches exist so that a *frozen* encoder is not re-run across an entire corpus on every training
63
+ experiment.
64
+
65
+ Three further distinctions are load-bearing and are stated at each point below:
66
+
67
+ - **Declared versus executed.** `configs/base.yaml` declares the tiling policy, the optical percentile
68
+ stretch and the SAR dB clip. All three are read by no code on the serving path.
69
+ - **Deterministic versus seeded.** Some stages are pure functions of their input (the display stretch, the
70
+ quality gate, the radiometric stretch). Others are stochastic and take a seeded RNG (channel dropout).
71
+ The distinction determines what "reproducible" means for each.
72
+ - **Cache-hit versus cache-miss.** Every cache in the project is keyed by a **spec hash** or a
73
+ **fingerprint**, and a mismatch is a **miss** or a **refusal**, never a silent stale read.
74
+
75
+ ---
76
+
77
+ ## 1. The end-to-end path
78
+
79
+ ### 1.1 The nine controller states
80
+
81
+ The controller runs a fixed nine-state machine, declared in `configs/base.yaml` under `agent.states` and
82
+ mirrored by the `ControllerState` enum in `core/schemas.py:79`:
83
+
84
+ ```
85
+ RECEIVE β†’ PARSE β†’ VALIDATE β†’ PLAN β†’ PREPROCESS β†’ EXECUTE β†’ AGGREGATE β†’ VERIFY β†’ RESPOND
86
+ ```
87
+
88
+ The division of labour is stated once and holds throughout: *the router understands, the policy decides,
89
+ the specialists compute, the VLM explains, and the evidence proves.* The three states that matter for the
90
+ data pipeline are VALIDATE, PREPROCESS and EXECUTE.
91
+
92
+ ### 1.2 The path, drawn
93
+
94
+ ```
95
+ POST /v1/assets (gateway; five-type allowlist, two-layer size cap)
96
+ β”‚
97
+ └─ asset_id (opaque, 128-bit; never a path)
98
+ β”‚
99
+ β–Ό
100
+ POST /v1/analyze { assets: [...], query }
101
+ β”‚
102
+ β”œβ”€ RECEIVE ── resolve handles to server-side paths
103
+ β”‚
104
+ β”œβ”€ PARSE ── router: query text β†’ Intent (task, modality, temporal, spatial, language)
105
+ β”‚
106
+ β”œβ”€ VALIDATE ── core/controller.py::_inspect_asset
107
+ β”‚ └─ preprocessing.raster.inspect_raster(path, max_pixels=..., ...)
108
+ β”‚ header-only: dims β†’ bands β†’ dtype β†’ CRS β†’ transform β†’ bounds
109
+ β”‚ β†’ nodata β†’ modality β†’ AssetMetadata
110
+ β”‚
111
+ β”œβ”€ PLAN ── policy selects a workflow (which specialists, in what order)
112
+ β”‚
113
+ β”œβ”€ PREPROCESS ── per-specialist preparation (this chapter's core)
114
+ β”‚ β”œβ”€ decode: read_bands β†’ (C, H, W); load_image_array β†’ (H, W, 3) uint8
115
+ β”‚ β”œβ”€ sensor adapter: band map + zero-fill + availability mask
116
+ β”‚ β”œβ”€ normalisation: display stretch / encoder-input stretch
117
+ β”‚ └─ resize: to the model's canonical resolution
118
+ β”‚
119
+ β”œβ”€ EXECUTE ── specialist forward pass β†’ SpecialistResult
120
+ β”‚
121
+ β”œβ”€ AGGREGATE ── combine specialist results
122
+ β”œβ”€ VERIFY ── consistency checks
123
+ └─ RESPOND ── ResultEnvelope { run_id, result, trace }
124
+ ```
125
+
126
+ `architecture/03-request-lifecycle.md` records a fact worth repeating here: the **PREPROCESS state is
127
+ never recorded** as a distinct `TraceStep` in the execution trace; the work happens inside each
128
+ specialist's `execute` rather than as a controller-level step. So the trace shows VALIDATE and EXECUTE but
129
+ not the preparation between them, and the preparation's parameters reach the trace through the
130
+ specialist's **evidence payloads** (the availability masks, the radiometry report, the decode metadata)
131
+ rather than through a PREPROCESS step.
132
+
133
+ ### 1.3 What each state reads and writes
134
+
135
+ | State | Reads | Writes | Cache consulted |
136
+ |---|---|---|---|
137
+ | RECEIVE | the uploaded asset handle | a server-side path | none |
138
+ | PARSE | the query text | an `Intent` | router corpus cache (training only) |
139
+ | VALIDATE | raster **header** only | `AssetMetadata` | none |
140
+ | PLAN | the `Intent`, the `AssetMetadata` list | a workflow plan | none |
141
+ | PREPROCESS | raster **pixels** | model-ready arrays | none |
142
+ | EXECUTE | model-ready arrays | a `SpecialistResult` | none (serving) |
143
+ | AGGREGATE / VERIFY / RESPOND | specialist results | the envelope + trace | none |
144
+
145
+ The single most important cell in that table is VALIDATE's "raster header only": the inspection that
146
+ decides whether an asset is usable **does not decode it**. `inspect_raster`'s docstring states this β€”
147
+ *"Validate and describe a raster without loading its pixel data"* β€” and it is what keeps a 25-megapixel
148
+ rejection cheap.
149
+
150
+ ---
151
+
152
+ ## 2. Validation and decoding
153
+
154
+ ### 2.1 Header-only validation β€” `inspect_raster`
155
+
156
+ Covered in full in [`GEOSPATIAL.md`](GEOSPATIAL.md) Β§1.3. The pipeline-relevant points:
157
+
158
+ - It reads dimensions, band count, dtypes, CRS, transform, bounds, nodata, resolution, driver and a tiled
159
+ flag, and constructs a `GeoMetadata` plus an `AssetMetadata`.
160
+ - The **pixel budget** is enforced here: `width * height > max_pixels` raises `OversizedImageError`, which
161
+ is recoverable β€” the caller may downscale. `max_pixels` comes from `image.max_pixels` (25,000,000).
162
+ - The **modality** is assigned here by `infer_modality` (Β§3).
163
+ - The **SHA-256** is computed here only when `compute_hash=True` (streaming, 1 MiB chunks) β€” it is needed
164
+ for manifests and dedup detection, and it is the field the change specialist uses to detect a repeated
165
+ acquisition.
166
+
167
+ ### 2.2 Band decoding β€” `read_bands`
168
+
169
+ `preprocessing/raster.py::read_bands(path, *, indexes=None, out_dtype=None)` returns
170
+ `(bands, H, W)` plus a profile that retains `crs`, `transform` and `bounds`. The profile is the source of
171
+ georeferencing for every downstream write (change maps, masks) so *"downstream code never loses
172
+ georeferencing."* It accepts an optional `indexes` list and an optional `out_dtype` cast, and wraps every
173
+ failure as `RasterReadError`.
174
+
175
+ Consumers:
176
+
177
+ - `preprocessing/imagery.py::load_image_array` β€” for the VQA, caption and grounding paths.
178
+ - `specialists/optical_sar/specialist.py::execute` β€” reads the optical and SAR arrays raw, then hands them
179
+ to the sensor adapter.
180
+ - `specialists/change/specialist.py::_write_artifacts` β€” reads T1's profile so the change map is written
181
+ with T1's georeferencing.
182
+
183
+ ### 2.3 Display decoding β€” `load_image_array`
184
+
185
+ `preprocessing/imagery.py::load_image_array` turns a raster into a displayable `(H, W, 3)` uint8 array:
186
+
187
+ 1. `read_bands` β†’ `(bands, H, W)`.
188
+ 2. Transpose to `(H, W, bands)`; take the first three bands as RGB; repeat a single band three times; or
189
+ stack a 2-D array three times.
190
+ 3. Per-**scene** 2/98 percentile stretch over finite values, clip to `[0, 1]`, scale to uint8.
191
+
192
+ Its docstring is explicit about what it does **not** do: *"It does not resample, crop, or reproject. Those
193
+ change the pixel grid, and the grounding specialist converts normalized boxes to pixel coordinates using
194
+ the ORIGINAL raster's dimensions β€” a silent resize here would put every box in the wrong place."* And
195
+ about determinism: *"The percentile stretch is deterministic: the same file always yields the same array,
196
+ so a grounding box and a VQA answer describe identical pixels."*
197
+
198
+ `load_pil_image` wraps the same array in a `PIL.Image`, and notes that georeferencing is not lost because
199
+ the caller's `AssetMetadata` already carries it.
200
+
201
+ ### 2.4 The deterministic quality gate
202
+
203
+ `preprocessing/quality.py` is *"a deterministic input-quality gate"* that sits **upstream of the model**.
204
+ It exists because of a measured failure: a loaded SmolVLM-500M-Instruct, given uniform random noise and a
205
+ prompt that explicitly said *"answer only from what is visible"*, produced *"a fluent, specific, entirely
206
+ fabricated scene description."* The gate's rationale is stated directly: *"A 500M-parameter VLM will
207
+ describe *something* for any input it is given, and asking it to self-assess reliably is asking it to do
208
+ the thing it just failed at."*
209
+
210
+ The discriminating signal is **lag-1 spatial autocorrelation**, not variance. The measured separation:
211
+
212
+ | Input | Autocorrelation |
213
+ |---|---|
214
+ | Real remote-sensing imagery | 0.6 – 0.99 |
215
+ | Uniform random noise | ~0.00 |
216
+ | A constant (blank) image | undefined; variance ~0 |
217
+
218
+ Variance alone cannot separate a flat desert scene from an all-zero tile (both have near-zero variance);
219
+ autocorrelation is high for both uniform-but-real imagery and textured imagery, and collapses only for
220
+ noise.
221
+
222
+ The constants:
223
+
224
+ | Constant | Value | Meaning |
225
+ |---|---|---|
226
+ | `MIN_AUTOCORRELATION` | `0.10` | Below this, the array is not spatially coherent |
227
+ | `FLAT_STD_EPSILON` | `1e-6` | Below this normalised std, the array is constant |
228
+ | `NOISE_ENTROPY_BITS` | `7.8` | Shannon entropy above which the histogram is noise-like |
229
+ | `BLOCKING_VERDICTS` | `{NOISE, INVALID_VALUES}` | Verdicts that must block a VLM call |
230
+
231
+ `assess_image_quality` returns an `ImageQuality` with one of five verdicts:
232
+
233
+ | Verdict | Meaning | Effect |
234
+ |---|---|---|
235
+ | `STRUCTURED` | Spatial structure consistent with real imagery | Safe to analyse |
236
+ | `FLAT` | Effectively constant | Usable, but `is_degraded` β€” lower confidence |
237
+ | `NOISE` | Uncorrelated, near-uniform | **Blocks the VLM call** |
238
+ | `TOO_SMALL` | Fewer than 16 px on a side | Usable, flagged |
239
+ | `INVALID_VALUES` | NaN or infinite values present | **Blocks the VLM call** |
240
+
241
+ Two behaviours are worth quoting. First, the non-finite rule is deliberately **tolerance-free**: *"a fixed
242
+ fraction is size-dependent, so one NaN in a 64x64 array (0.99976) would pass while one NaN in a 10x10
243
+ array (0.99) would fail. The same defect must not be tolerated or rejected depending on image
244
+ dimensions."* Any non-finite value is disqualifying. Second, the entropy check is described as a
245
+ **redundant second signal**: *"MEASURED, and the margin here is thin: structured ramp+texture gives 7.581,
246
+ uniform noise gives 7.988. That is a 0.22-bit gap against a 7.8 threshold. The autocorrelation check is
247
+ what actually carries this gate (0.952 vs -0.008, a 0.96 margin against a 0.10 threshold)."*
248
+
249
+ The gate is a pure function β€” *"Nothing here is learned, sampled, or probabilistic. Same input, same
250
+ verdict."* A known limitation is recorded: a float GeoTIFF whose nodata sentinel is NaN lands in
251
+ `INVALID_VALUES` and is refused, and the comment records that *"Filling nodata from the raster profile
252
+ belongs in the tiling work"* β€” i.e. the fix is deferred to a stage that does not exist (Β§5).
253
+
254
+ ### 2.5 Error types
255
+
256
+ Every failure in the validation and decoding path is a typed `SatQueryError` subclass. The relevant ones
257
+ for this pipeline:
258
+
259
+ | Error | Raised by | Meaning |
260
+ |---|---|---|
261
+ | `RasterReadError` | `inspect_raster`, `read_bands`, `write_raster` | The file is not a readable raster |
262
+ | `OversizedImageError` | `inspect_raster` | Exceeds the pixel budget; **recoverable** |
263
+ | `UnsupportedBandsError` | `inspect_raster`, the sensor adapter | Band count is zero, or bands cannot be placed |
264
+ | `MissingCRSError` | `parse_crs` | A CRS string is malformed |
265
+ | `CoordinateError` | `geospatial/transform.py` | A conversion lacks the context it needs |
266
+ | `SpecialistError` | specialists | A forward pass or preparation failed |
267
+
268
+ The design rule from `preprocessing/raster.py` applies throughout: *"Never raise a bare exception. Every
269
+ failure is a typed SatQueryError."*
270
+
271
+ ---
272
+
273
+ ## 3. Modality inference, and where it is consumed
274
+
275
+ `infer_modality(band_count, explicit=None)` (`preprocessing/raster.py:58`) is covered in full in
276
+ [`GEOSPATIAL.md`](GEOSPATIAL.md) Β§6. Its pipeline role is as the **first consumer** of the band count that
277
+ VALIDATE read:
278
+
279
+ - `inspect_raster` calls it and stores the result on `AssetMetadata.modality`.
280
+ - The optical-SAR specialist's `_assess_pair` reads the **declared** modality first and falls back to
281
+ `infer_modality(asset.geo.band_count)` only when it is `UNKNOWN`, appending a warning that names the
282
+ heuristic.
283
+ - The change specialist does **not** infer modality: it requires exactly two assets and compares their
284
+ CRS and registration, without asserting a modality for either.
285
+
286
+ The heuristic's limits are stated in its own comment: *"These are heuristics, not ground truth β€” the
287
+ sensor adapter is authoritative when a sensor descriptor is supplied."* A band count is not a sensor
288
+ declaration, and the adapter is the only component that can map a band to a canonical channel.
289
+
290
+ ---
291
+
292
+ ## 4. The normalisation stages
293
+
294
+ There are **three** distinct normalisation transforms in this codebase, plus one **declared but
295
+ unimplemented** stage. Conflating them is the most common misreading, so they are laid out side by side.
296
+
297
+ ### 4.1 The display stretch β€” per-scene 2/98
298
+
299
+ **Where:** `preprocessing/imagery.py::load_image_array` (and its grayscale cousin in
300
+ `specialists/change/postprocess.py::to_grayscale_float`).
301
+
302
+ **What:** one `(lo, hi)` computed over **all finite values across all bands** of one scene, then
303
+ `(arr - lo) / (hi - lo)`, clipped to `[0, 1]`.
304
+
305
+ **Output:** uint8 `(H, W, 3)` for display (or float32 grayscale for registration).
306
+
307
+ **Purpose:** presentation β€” so a VQA answer and a grounding box describe identical pixels.
308
+
309
+ **Determinism:** pure; same file β†’ same array.
310
+
311
+ **Not** the CROMA transform. `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§1.1 names three disqualifying
312
+ differences: per-scene rather than per-channel; uint8 3-band rather than float 12-channel; presentation
313
+ rather than radiometry. Reusing it for CROMA is explicitly **forbidden** by that ruling (Β§3.3).
314
+
315
+ ### 4.2 The encoder-input stretch β€” per-channel `mean Β± 2Β·std`
316
+
317
+ **Where:** `specialists/optical_sar/radiometry.py::normalise_for_croma`.
318
+
319
+ **What:** for each channel `c` of each sample, `lower = mean_c - 2Β·std_c`, `upper = mean_c + 2Β·std_c`,
320
+ then `(x - lower) / (upper - lower)`, optionally through a uint8 round-trip (`Γ—255`, clip, quantise,
321
+ `/255`). Computed **per sample** (not per batch) with `ddof=1`.
322
+
323
+ **Output:** float32 of the same shape, in `[0, 1]`.
324
+
325
+ **Purpose:** the transform the CROMA authors' own README instructs users to apply before the frozen
326
+ encoder.
327
+
328
+ **Determinism:** *"pure and deterministic: no sampling, no learned statistics, no RNG, and no dependence
329
+ on batch composition. Same input -> same output."*
330
+
331
+ **Where it runs:** the **extraction** path (`training/fusion/extract.py`), not the serving path (Β§7.5).
332
+ The extraction module applies it *"because the arm's `use_8_bit` has to be applied at exactly one
333
+ place"*, and it refuses an encoder that would stretch a second time.
334
+
335
+ **The zero-channel rule:** an all-zero channel (the sensor adapter's unavailability convention) has
336
+ `std = 0` and would divide by zero. Such channels are **skipped and left at exactly zero**, with status
337
+ `unavailable`; a present-but-constant channel is skipped with status `degenerate`. This is the C-1
338
+ discipline applied to radiometry β€” *"the same rule that forbids inventing a band forbids inventing a
339
+ dynamic range for a band that does not exist."*
340
+
341
+ ### 4.3 The declared-but-unimplemented conditioning stage
342
+
343
+ **Where declared:** `configs/base.yaml`:
344
+
345
+ ```yaml
346
+ optical:
347
+ normalization: percentile
348
+ lower_percentile: 2
349
+ upper_percentile: 98
350
+ sar:
351
+ representation: db
352
+ clip_min_db: -30
353
+ clip_max_db: 5
354
+ ```
355
+
356
+ **Status:** `DECLARED (not read)`. No code applies a percentile/dB conditioning stage. The radiometry
357
+ module's docstring states the consequence: *"the two `optical.*` and three `sar.*` config keys are still
358
+ read by no code. That is recorded, not fixed."* `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§6 item 3
359
+ adds: *"A config key that no code reads is worse than a missing key: it reads as a satisfied
360
+ requirement."*
361
+
362
+ ### 4.4 The side-by-side
363
+
364
+ | Transform | Granularity | Output | Purpose | Implemented | Runs on serving path |
365
+ |---|---|---|---|---|---|
366
+ | Display 2/98 | per-scene, all bands | uint8 (H,W,3) | Presentation | βœ… | βœ… |
367
+ | Encoder-input `mean Β± 2Β·std` | per-channel, per-sample | float32 `[0,1]` | Frozen-encoder interface | βœ… | ❌ (extraction only) |
368
+ | Change grayscale 2/98 | per-scene, grayscale | float32 `[0,1]` | Registration correlation | βœ… | βœ… |
369
+ | Percentile / dB conditioning | per-scene, per config | unbounded rescale | `base.yaml`'s intent | ❌ | ❌ |
370
+
371
+ ---
372
+
373
+ ## 5. Tiling and tile selection
374
+
375
+ Covered in full in [`GEOSPATIAL.md`](GEOSPATIAL.md) Β§10. The pipeline-relevant summary:
376
+
377
+ - `image.max_pixels: 25000000` **is enforced**, at `inspect_raster` time, raising `OversizedImageError`.
378
+ - `image.tile_size: 512`, `image.tile_overlap: 128`, `image.max_tiles: 64`, `image.top_k_tiles: 4` are
379
+ **declared and read only by** `core/config.py`'s single relationship check
380
+ (`top_k_tiles <= max_tiles`). No serving-path code slices a raster into tiles or selects a top-K.
381
+ - `change.tile_size: 256`, `change.tile_overlap: 0` are likewise declared and not consumed as a tiling
382
+ step; 256 is the change model's training resolution.
383
+
384
+ **Consequence for the pipeline:** a raster within the pixel budget is handed to a specialist **whole**.
385
+ The only size-reduction steps that actually run are per-specialist resizes to a model's canonical
386
+ resolution (e.g. `resize_to_canonical` to 120Γ—120 for CROMA, and the change feature extractor's bilinear
387
+ resize to 256Γ—256). Neither is tiling: both resample the whole image rather than selecting a region.
388
+
389
+ The VLM path's `processor_longest_edge: 512` is a **control against** the processor's default 2048, not a
390
+ tiling step. `core/config.py` enforces `processor_longest_edge <= image.tile_size`, with the measured
391
+ rationale recorded: the default 2048 upscales a 512 tile 4Γ— and splits it into ~17 sub-images, *"the real
392
+ figure is ~17x"* against the plan's estimated 4Γ—.
393
+
394
+ ---
395
+
396
+ ## 6. The change-VQA data path
397
+
398
+ The change-VQA path (`Task.CHANGE_VQA`, R-02) is the most elaborate in the project: it has a frozen
399
+ detector, two frozen feature extractors, a trained head, and four cache artefacts. It is the clearest
400
+ illustration of the serving/preparation split.
401
+
402
+ ### 6.1 The dataset and the preprocessing version tag
403
+
404
+ `training/change_vqa/dataset.py` defines the identity constants:
405
+
406
+ ```python
407
+ DATASET_ID = "cdvqa"
408
+ PREPROCESSING_VERSION = "change_vqa_preproc_v1"
409
+ SPLIT_MAP: dict[str, str] = {
410
+ "Train": "train",
411
+ "Val": "val",
412
+ "Test": "test",
413
+ "Test2": "test",
414
+ }
415
+ ```
416
+
417
+ `PREPROCESSING_VERSION` is *"bumped when the record shape or the target definition changes. Recorded in
418
+ every evaluation so two numbers computed under different definitions can never be silently compared."*
419
+ It is the **preprocessing version tag** for this path: it names the *record and target* definition, and is
420
+ distinct from the *feature spec* hashes below, which name the *tensor* definition.
421
+
422
+ `SPLIT_MAP` is a leakage control, not a convenience: `Test2` maps to `test` because Test2 is *"a SECOND
423
+ QUESTION SET over the SAME 968 scenes as Test, not an independent sample; it must never be pooled."*
424
+ `docs/PHASE10_CDVQA_DATA_STATUS.md` Β§4 measured this: Test and Test2 share **100% of their images** (968 /
425
+ 968), while Train/Val/Test are cleanly disjoint.
426
+
427
+ `LABEL_PALETTE` decodes the six change classes (NVG surface, trees, low vegetation, water, buildings,
428
+ playgrounds) from a fixed colour palette; white `(255,255,255)` is a **shared background**, not a class.
429
+
430
+ `SceneTargets` derives supervision from `label1`/`label2` with four definitions, all fractions of total
431
+ pixels: `class_mag[c]` (magnitude), `class_delta[c]` (signed), `class_ratio[c]` (per-class ratio) and
432
+ `total_changed`. The `total_changed` definition was **validated against gold answers**: over 40 real
433
+ `change_ratio` rows, `count(label1 != label2) / total` reproduced the annotated bin **34/40 = 85%** of the
434
+ time, versus **1/40** for the competing definition. The residual 15% is recorded as *"annotator/map
435
+ disagreement, which is the noise floor of the target"* rather than smoothed away.
436
+
437
+ ### 6.2 The change feature extractor
438
+
439
+ `training/change_vqa/features.py` defines the frozen change feature. The module explains why the features
440
+ are frozen: the plan budgets CDVQA at ≀ 3 h on T4Γ—2, and training end-to-end through STANet would
441
+ re-learn a representation the project already has. Freezing it drops the trainable parameter count from
442
+ ~15.8 M to ~1.5 M.
443
+
444
+ `CHANGE_FEATURE_DIM = 1045`, and it is **three things**:
445
+
446
+ | Part | Dim | Composition | Answers |
447
+ |---|---|---|---|
448
+ | `level_pooled` | 1024 | 4 levels Γ— 128 ch Γ— {mean, max} | *"what does the change look like locally?"* |
449
+ | `change_stats` | 5 | mean, std, frac>0.3, frac>0.5, frac>0.7 of the change map | *"how much of the scene changed?"* |
450
+ | `change_grid` | 16 | adaptive 4Γ—4 pool of the same map | *"WHERE did it change?"* |
451
+
452
+ The dimensions are **derived, never typed** β€” *"a literal here would drift silently from the pooling code
453
+ below and produce a shape error three files away"*:
454
+
455
+ ```python
456
+ N_LEVELS = 4
457
+ LEVEL_CHANNELS = 128
458
+ LEVEL_POOLED_DIM = N_LEVELS * LEVEL_CHANNELS * 2 # 1024
459
+ CHANGE_STATS_DIM = 2 + len(CHANGE_STAT_THRESHOLDS) # 5
460
+ CHANGE_GRID_DIM = CHANGE_GRID * CHANGE_GRID # 16
461
+ CHANGE_FEATURE_DIM = LEVEL_POOLED_DIM + CHANGE_STATS_DIM + CHANGE_GRID_DIM # 1045
462
+ ```
463
+
464
+ The spatial parts are load-bearing: *"Without `change_grid` the head would see a scene-global average and
465
+ could not distinguish 'buildings changed in one corner' from 'buildings changed everywhere', which is
466
+ precisely the difference between `largest_change` and `smallest_change`."*
467
+
468
+ The default working resolution is `DEFAULT_IMAGE_SIZE = 256` β€” *"the change model's OWN training
469
+ resolution (`configs/base.yaml: change.tile_size: 256`; LEVIR-CD patches are 256x256). Feeding the native
470
+ 512 would push the frozen encoder outside the distribution it was trained on."* 512 is offered as an
471
+ explicit, recorded alternative.
472
+
473
+ `ChangeFeatureExtractor.extract(t1, t2)` re-runs the detector's own submodules to capture the intermediate
474
+ fused levels that `STANetStyleChangeDetector.forward` does not return. That is *"a second implementation
475
+ of one forward pass, so it is a drift hazard β€” and it is closed by a test that asserts the probability map
476
+ produced here is **bit-identical** to `STANetStyleChangeDetector.forward(...).probabilities`."*
477
+
478
+ ### 6.3 The text feature extractor
479
+
480
+ `TextFeatureExtractor` wraps `sentence-transformers/all-MiniLM-L6-v2` β€” the same model the router uses, so
481
+ *"the text side of the reasoning head adds no new model to the deployment."* It produces `TEXT_FEATURE_DIM
482
+ = 384` normalised vectors and raises `FeatureExtractionError` if the encoder's output width disagrees with
483
+ the declared constant.
484
+
485
+ ### 6.4 The cache spec hashes
486
+
487
+ Two independent hash functions key the two caches. Both are `sha256(...)[:16]`.
488
+
489
+ `feature_spec_hash(image_size, *, extractor)` hashes the feature definition:
490
+ `{spec, image_size, extractor, levels, level_channels, thresholds, grid, text_encoder}`, sorted keys.
491
+
492
+ `text_spec_hash(*, encoder, revision=None)` hashes the question-feature definition:
493
+ `{spec: "change_vqa_text_v1", encoder, revision, dim, normalised}`, sorted keys. The separate hash is
494
+ deliberate: *"One combined hash would force a full re-extraction of both whenever either changed, and β€”
495
+ worse β€” would let a stale text cache pass a check it should fail."*
496
+
497
+ **Recomputed values** (this document computed them from the functions above, with the constants as read
498
+ from the source):
499
+
500
+ | Input | Spec hash |
501
+ |---|---|
502
+ | `feature_spec_hash(256, extractor="stanet:trained")` | `c801326f85a185f8` |
503
+ | `feature_spec_hash(256, extractor="stanet:untrained")` | `714efac5b0e6a6d1` |
504
+ | `text_spec_hash(encoder="sentence-transformers/all-MiniLM-L6-v2", revision="1110a243fdf4")` | `d2801ea1a314354a` |
505
+ | `text_spec_hash(encoder="sentence-transformers/all-MiniLM-L6-v2")` (unpinned) | `f87f5226c4c4a69b` |
506
+
507
+ So the change-feature cache spec for the trained detector at 256 px is **`c801326f85a185f8`**, and the
508
+ question-feature cache spec with the pinned MiniLM revision (`configs/base.yaml: router.revision:
509
+ 1110a243fdf4`) is **`d2801ea1a314354a`**. The spec hash **encodes trained-vs-untrained**, which is how the
510
+ serving specialist detects a detector-state mismatch (Β§6.6).
511
+
512
+ ### 6.5 Cache layout, resume, and the artefacts
513
+
514
+ `scripts/prepare_change_vqa.py` is the producer. Its default output directory is
515
+ `DEFAULT_OUT = REPO_ROOT / "artifacts" / "change_vqa"`, and it writes:
516
+
517
+ | Artefact | Shape | Content |
518
+ |---|---|---|
519
+ | `<out>/scene_manifest.jsonl` | one record per `(native split, scene)` | identity, integrity, split, question counts |
520
+ | `<out>/scene_targets.jsonl` | one record per scene | label-derived `SceneTargets` |
521
+ | `<out>/change_features.npz` | `(n_scenes, 1045)` + spec hash + extractor config | frozen STANet features, keyed by `scene_key` |
522
+ | `<out>/text_features_<split>.npz` | `(n_questions, 384)` + spec hash | question features, keyed by `question_id` |
523
+
524
+ The manifest is **scene-level, deliberately**: *"153,130 question rows would produce a ~50 MB JSONL that
525
+ duplicates text already authoritative in `annotations/*.json`. The manifest therefore records the **2,968
526
+ scenes** ... and the question-level records are derived from the annotations on demand."*
527
+
528
+ Resume is spec-guarded. `ChangeFeatureCache.read(path, *, expect_spec_hash=...)` **refuses** a cache built
529
+ under a different spec:
530
+
531
+ ```python
532
+ if expect_spec_hash is not None and cache.spec_hash != expect_spec_hash:
533
+ raise FeatureExtractionError(
534
+ f"change feature cache at {p} was built under spec "
535
+ f"{cache.spec_hash!r} but the caller expects {expect_spec_hash!r}; "
536
+ f"re-extract rather than training on a representation that does not match serving"
537
+ )
538
+ ```
539
+
540
+ `TextFeatureCache.read` applies the same guard. The producer calls
541
+ `ChangeFeatureCache.read(cache_path, expect_spec_hash=extractor.spec_hash)` when resuming, so a
542
+ resolution change or a detector-state change invalidates the cache instead of silently mixing two
543
+ representations.
544
+
545
+ A scene whose imagery cannot be read is **skipped and reported** through `on_progress`, never silently
546
+ dropped: *"a scene missing from the cache is a training sample the model never sees, and the caller must be
547
+ able to count how many that was."*
548
+
549
+ ### 6.6 The serving path, and the spec-mismatch refusal
550
+
551
+ `ChangeVQASpecialist` (`specialists/change/vqa_specialist.py`) assembles three pieces: a trained head, a
552
+ change feature extractor, and a question text encoder. Its `execute`:
553
+
554
+ 1. Validates exactly two assets, both present.
555
+ 2. Calls `unavailable_reason()`; if anything is missing or mismatched, returns `_unavailable_result` with
556
+ **`answer=""`** and `degraded=True`.
557
+ 3. Otherwise loads the two images, extracts `(vector, facts)` and the text vector, resolves the question
558
+ type and temporal reference, and calls `predict_answers`.
559
+
560
+ **Why degraded produces no answer.** The module docstring is explicit: *"a random change map is visibly
561
+ noise, whereas an untrained 19-way classifier still emits a fluent, confident-looking `yes`. A user cannot
562
+ tell the second from a real answer, so the untrained case returns **no answer at all**."* The
563
+ `SpecialistResult` validator independently marks an empty change-VQA answer as degraded.
564
+
565
+ **The spec-mismatch refusal.** `feature_spec_mismatch()` compares the head's trained spec
566
+ (`head_metadata["change_cache_spec"]`) with the serving extractor's `spec_hash`. A mismatch is a
567
+ **refusal, not a warning**: *"A head is only meaningful for the feature distribution it was fitted on."*
568
+ The measured reason this matters: `configs/base.yaml`'s `change:` section carries **no `checkpoint_path`
569
+ key**, so the registry builds the change feature extractor with `checkpoint_path=None` and gets an
570
+ **untrained** STANet β€” while `scripts/prepare_change_vqa.py` defaults to the trained LEVIR checkpoint.
571
+ Training and serving would therefore consume different representations, and *"the head would still emit a
572
+ fluent `yes`."* The spec hash covers this because it encodes trained-vs-untrained, so the refusal catches
573
+ the detector-state mismatch as well as a resolution mismatch.
574
+
575
+ Other serving constants: `LOW_CONFIDENCE_THRESHOLD = 0.40` (a reporting threshold, not a calibration) and
576
+ `TEMPORAL_ORDER_NOTE` (T1=pre, T2=post; `label1=pre/label2=post` **proven** at agreement 1.0000 over 2,968
577
+ scenes; `im1=pre/im2=post` **supported statistically, not proven**).
578
+
579
+ **A config fact worth recording.** The specialist reads `change_vqa.image_size` (default 256),
580
+ `change_vqa.apply_type_mask` (default `True`) and `change_vqa.head_path` via `config.get(...)`, but
581
+ **`configs/base.yaml` has no `change_vqa` block at all**. Those keys therefore always take their code
582
+ defaults, and the frozen config does not name them. This is the same pattern as the grounding head's
583
+ `DEFAULT_HEAD_PATH` (a code constant rather than a config key, to avoid moving `Config.hash`).
584
+
585
+ ---
586
+
587
+ ## 7. The optical-SAR data path
588
+
589
+ ### 7.1 The serving path, step by step
590
+
591
+ `OpticalSarSpecialist.execute` (`specialists/optical_sar/specialist.py:320`):
592
+
593
+ 1. **Validate.** Exactly two assets; `_assess_pair` establishes one optical and one SAR (Β§9.2 of
594
+ [`GEOSPATIAL.md`](GEOSPATIAL.md)).
595
+ 2. **Descriptors.** `_descriptor_for` returns the asset's `SensorDescriptor` if present, else builds one
596
+ from the band count β€” `build_optical_adapter` for optical, `positional_fallback_descriptor` for SAR β€”
597
+ and appends a warning naming the assumption.
598
+ 3. **Read pixels.** `read_bands(optical_asset.path)` and `read_bands(sar_asset.path)`.
599
+ 4. **Pipeline.** `run_pipeline(...)` (below).
600
+ 5. **Confidence.** `_confidence_components` computes the four plan-section-26 components.
601
+ 6. **Answer + evidence.** `_compose_answer` builds prose only from computed facts; `_build_evidence`
602
+ emits the availability masks, the view items, the fused-representation item and the decision.
603
+
604
+ ### 7.2 `run_pipeline`
605
+
606
+ `specialists/optical_sar/inference.py::run_pipeline` is the seam between "an image on disk" and "three GAP
607
+ vectors plus two masks":
608
+
609
+ ```
610
+ optical raster -> sensor adapter -> (12, H, W) + mask[12]
611
+ SAR raster -> sensor adapter -> ( 2, H, W) + mask[2]
612
+ resize to CROMA's 120x120
613
+ CROMA(SAR_images=..., optical_images=...) <- no mask, ever
614
+ assemble_fusion_input <- mask enters HERE
615
+ fusion head -> logits -> probabilities
616
+ ```
617
+
618
+ Two properties are load-bearing:
619
+
620
+ - **Degradation is a first-class outcome.** `run_pipeline` *"NEVER raises for a missing model: the
621
+ degradation is reported, because 'we have no trained head' is a fact about the deployment, not a failure
622
+ of the request."* It raises only for a caller error β€” arrays that cannot be canonicalised at all.
623
+ - **The sensor side always runs.** *"The masks, the canonical channel placement, the availability
624
+ statistics and the modality validation are all real work that requires no weights, and they are exactly
625
+ the facts that tell an operator why a result is or is not trustworthy on Cartosat-2S + RISAT."*
626
+
627
+ `resize_to_canonical(array, resolution)` bilinearly resizes a `(C, H, W)` array to
628
+ `(C, resolution, resolution)`. Its docstring notes that nearest-neighbour would be wrong for continuous
629
+ optical reflectance channels, while for zero-filled channels *"interpolating zeros gives zeros."* It uses
630
+ `cv2.INTER_LINEAR` when OpenCV is available, falling back to `torch.nn.functional.interpolate`.
631
+
632
+ ### 7.3 The mask path
633
+
634
+ `build_masks` converts the two `SensorAdapterOutput` masks to `(1, C)` float32, ready for concatenation.
635
+ `mask_availability_stats` computes the fractions and counts that feed both the confidence components and
636
+ the evidence. The mask **never reaches CROMA** (finding C-1) β€” the specialist's evidence item records
637
+ `"croma_received_mask": false` and `"mask_consumed_by": "fusion_head"`.
638
+
639
+ ### 7.4 The fusion cache β€” `fusion_features`
640
+
641
+ `training/fusion/train.py` and `training/fusion/extract.py` implement the Phase 12 frozen-feature trainer
642
+ and its missing producer. The pipeline:
643
+
644
+ ```
645
+ paired patches -> normalise -> resize -> CROMA (frozen) -> FusionFeature
646
+ -> write_feature_cache -> train_fusion_head
647
+ ```
648
+
649
+ The cache constants:
650
+
651
+ | Constant | Value | Meaning |
652
+ |---|---|---|
653
+ | `CACHE_VERSION` | `"v1"` | Bumped when the encoder, resolution or stored fields change |
654
+ | `DEFAULT_FEATURE_CACHE_DIR` | `artifacts/optical_sar/fusion_features` | Under `artifacts/`, scoped to this specialist, **not** under the frozen `artifacts/change/` tree |
655
+
656
+ `feature_cache_path(cache_dir=None, tag="")` returns
657
+ `base / f"croma_features_{CACHE_VERSION}{suffix}.npz"`, where the suffix is `_<tag>` when a tag is given
658
+ (e.g. a split name). The Phase-12 CLI (`scripts/extract_fusion_features.py`) instead takes an explicit
659
+ `--out-cache`, documented in its own usage as `artifacts/optical_sar/fusion_features/train.npz`, and writes
660
+ a `.npz` plus a `.json` sidecar. So both naming conventions are in play: the training loop's default is
661
+ `croma_features_v1_<tag>.npz`, and the CLI's documented example is `<split>.npz`. Both are
662
+ `artifacts/optical_sar/fusion_features/` files, and `verify_feature_cache` checks `row_counts_agree` and
663
+ `sidecar_n_matches_rows`.
664
+
665
+ **The write is atomic-ish.** `write_feature_cache` writes both files to unique temporary names and renames
666
+ them into place with `os.replace`, so *"no reader ever sees a half-written `.npz`."* The two renames cannot
667
+ be one operation; a crash in the window leaves a complete-but-stale pair, which `verify_feature_cache`
668
+ **reports** rather than letting a truncated array load cleanly and be silently wrong. A crashed run may
669
+ leave a `.<name>.tmp-<pid>-<uuid>.npz` behind, and the function deliberately never deletes it β€” *"a delete
670
+ is not a safe operation to perform on someone else's filesystem."*
671
+
672
+ **The multi-label collapse is an explicit policy.** BigEarthNet v2.0 is multi-label; the head is a
673
+ single-label 19-class softmax. The collapse is a parameter, `label_policy: Sequence[str] -> int | None`,
674
+ and with no policy: exactly one label β†’ its index; more than one β†’ `AmbiguousLabelError`; zero β†’
675
+ `None` β†’ the sample is **skipped and counted**, never mapped to class 0 (*"'we do not know' and 'arable
676
+ land' are different facts and must not be merged"*).
677
+
678
+ **Resume is provenance-guarded.** Re-running with an existing cache loads it, skips the `sample_id`s
679
+ already present and appends the rest (`existing + new`, never a replacement). But *"appending rows from a
680
+ different arm, resolution or config would produce a cache whose metadata describes only half its rows"* β€”
681
+ so `check_resume_provenance` refuses that and names the field and both values; `resume=False` rebuilds.
682
+
683
+ **The dry run makes the same decision, not a similar one.** `plan_extraction` calls the same
684
+ `select_samples` the run calls, so `plan.n_selected` is `result.n_encoded` and `plan.n_would_write` is
685
+ `result.n_total`.
686
+
687
+ **Exit codes.** `scripts/extract_fusion_features.py` returns `0` on success (including a dry run), `2` for
688
+ a pre-flight or input error (nothing was encoded β€” a missing corpus, a leaked split, an unwritable
689
+ `--out-cache`, a missing CROMA checkpoint, an ambiguous label under `require_single_label`), and `3` for a
690
+ run that **started and then failed** with a typed error (e.g. a cached feature that is not 2318 wide). The
691
+ comment on the `FusionTrainingError` branch states the distinction: *"the run started, so this is 3, not a
692
+ pre-flight 2."*
693
+
694
+ ### 7.5 The normalisation arm β€” A / B
695
+
696
+ The Phase 12 trainer selects the normalisation arm through the **existing hash-exempt environment
697
+ channel** `SATQUERY_CROMA_USE_8_BIT`, **not** by editing `configs/base.yaml`:
698
+
699
+ ```python
700
+ ARM_CONTROL = "A"
701
+ ARM_VARIANT = "B"
702
+
703
+ ARMS: dict[str, Arm] = {
704
+ ARM_CONTROL: Arm(ARM_CONTROL, True, "... use_8_bit=true ..."),
705
+ ARM_VARIANT: Arm(ARM_VARIANT, False, "... use_8_bit=false ..."),
706
+ }
707
+ ```
708
+
709
+ The reason for the env channel is stated: *"Adding a `fusion_training:` block would move the frozen
710
+ `Config.hash` off `78f1e3700da15aa1` and detach the Phase 9 benchmark from its config."* `apply_arm` writes
711
+ the value and returns it so a caller can record it.
712
+
713
+ **A warning that must be repeated, not summarised.** The `B` in this dict is **NOT** the arm B registered
714
+ in `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§4. Both arms run the **same** per-channel `mean Β± 2Β·std`
715
+ encoder-input stretch; they differ **only** on whether the result then takes the uint8 round-trip. The
716
+ axis is the **uint8 axis**, not the registered "stretch vs no-stretch" axis. The originally preregistered
717
+ arm B (*"percentile/dB conditioning only, no encoder-input stretch"*) is **NOT CONSTRUCTIBLE** from this
718
+ repository: stage 1 is unimplemented, and `Arm` exposes no field that can express "skip the encoder-input
719
+ stretch" β€” the stretch in `training/fusion/extract.py` is applied unconditionally. There is **no arm C**.
720
+
721
+ **Blocker recorded, not crossed.** The arm is baked into the cached features, and
722
+ `artifacts/optical_sar/fusion_features/` holds **only the Arm-A cache** (`train.npz` and `train.json`;
723
+ `val.npz` and `test.npz` were pending as of the Phase-14 correction notes). Arm B therefore needs its own
724
+ feature cache before it can be trained.
725
+
726
+ ### 7.6 The extraction module's double-stretch guard
727
+
728
+ `training/fusion/extract.py` applies the DEV-2 stretch itself, and refuses an encoder that would apply it
729
+ again:
730
+
731
+ > `CROMAEncoder` applies the same stretch inside `encode` when `normalize_input=True` (its default).
732
+ > Passing such an encoder here would stretch twice β€” a silent distribution shift that produces a perfectly
733
+ > well-shaped cache full of wrong numbers. `_assert_raw_encoder` refuses it by name rather than trusting
734
+ > the caller; `build_extraction_encoder` constructs the raw encoder this module expects.
735
+
736
+ The CLI prints `normalize_input : <value> (must be False)`, so the guard is visible in the run output as
737
+ well as enforced in code.
738
+
739
+ ---
740
+
741
+ ## 8. The router data path
742
+
743
+ ### 8.1 The frozen encoder
744
+
745
+ `router/encoder.py::FrozenEncoder` wraps `SentenceTransformer` and holds it frozen. The verified contract
746
+ (Phase 4 probe, 2026-09-16, sentence-transformers 6.0.1) is recorded at the top of the module:
747
+
748
+ ```
749
+ SentenceTransformer(
750
+ model_name_or_path='sentence-transformers/all-MiniLM-L6-v2',
751
+ revision='1110a243fdf4', # pinned, verified reachable
752
+ device='cpu',
753
+ )
754
+ .get_sentence_embedding_dimension() -> 384
755
+ .tokenizer.model_max_length -> 256 (finding F4-1)
756
+ .max_seq_length -> 256
757
+ params -> 22,713,216
758
+ encode(64 queries, CPU) -> 0.118 s
759
+ ```
760
+
761
+ Two constants are declared: `VERIFIED_TOKENIZER_MAX_LENGTH = 256` and `VERIFIED_EMBEDDING_DIM = 384`. The
762
+ constructor **refuses** a `max_length` above 256 β€” *"truncation would be a silent no-op"* β€” and sets
763
+ `self._model.max_seq_length = max_length` so the configured truncation (128, from
764
+ `router.max_length`) is *"enforced rather than merely documented."* The encoder returns detached numpy
765
+ arrays with no autograd history, which is what makes cached-embedding training possible (finding F4-2).
766
+
767
+ ### 8.2 The corpus embedding cache
768
+
769
+ `router/train.py` is the consumer. `CORPUS_CACHE_VERSION = "v1"`, and `_corpus_fingerprint(corpus,
770
+ encoder_id, max_length)` hashes:
771
+
772
+ ```python
773
+ payload = {
774
+ "version": CORPUS_CACHE_VERSION,
775
+ "encoder": encoder_id, # f"{encoder.model_name}@{encoder.revision}"
776
+ "max_length": max_length,
777
+ "texts": sorted(e.text for e in corpus),
778
+ }
779
+ return hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
780
+ ```
781
+
782
+ `embed_corpus_cached` writes two files into the cache directory:
783
+
784
+ | File | Content |
785
+ |---|---|
786
+ | `embeddings_<fingerprint[:16]>.npy` | the `(n, 384)` float32 embedding matrix |
787
+ | `embeddings_<fingerprint[:16]>.json` | `{fingerprint, encoder, max_length, count, dim, seconds}` |
788
+
789
+ The cache key includes the encoder id **and** the corpus text set, so *"changing either invalidates it. A
790
+ stale cache silently training on the wrong vectors would be worse than no cache."* A cache hit requires
791
+ both `cached.shape[0] == len(corpus)` **and** `meta.get("fingerprint") == fingerprint`. A bad cache is a
792
+ **miss, not an error**: `except Exception: pass` falls through to re-encoding.
793
+
794
+ **This cache is a training-time artefact.** The serving path does not read it. `embed_corpus_cached` is
795
+ defined and used only in `router/train.py` (defined at `:232`, used at `:518`), while the serving path
796
+ calls `self.router.route(request.query)` (`core/controller.py:476`), which encodes the query live β€” a
797
+ single short string, measured at 0.118 s for 64 queries on CPU, so one query is sub-millisecond.
798
+
799
+ ### 8.3 The router's training discipline
800
+
801
+ The module records two findings that shape the data path:
802
+
803
+ - **F4-2 (frozen encoder β‡’ cached training).** *"The encoder is FROZEN. Embeddings are therefore a pure
804
+ function of the query text. So we embed the whole corpus ONCE, cache the vectors to disk, and train the
805
+ 50,822-parameter adapter on the cached matrix."* Router training needs **no GPU**.
806
+ - **F4-3 (splits by group, never by example).** *"Splitting by example would put 'show me the water body'
807
+ in train and 'show me the road' in val β€” same template, one token apart β€” and report a fake accuracy."*
808
+ The train function **refuses to proceed** if the split report is not clean.
809
+
810
+ The router's reported accuracy is **0.965116 on validation, ungated, n = 86**; the **test split was NOT
811
+ RUN**. This is a data-pipeline fact: the test split exists and is unused.
812
+
813
+ ---
814
+
815
+ ## 9. The grounding data path
816
+
817
+ ### 9.1 The path
818
+
819
+ `GroundingSpecialist.execute`:
820
+
821
+ 1. Validate exactly one image and a non-empty phrase; check the path exists.
822
+ 2. `load_pil_image(asset.path)` β€” the deterministic display decode (Β§2.3).
823
+ 3. `encoder.encode_image(image)` and `encoder.encode_text([query])` β€” RemoteCLIP ViT-B/32 at 224 px,
824
+ producing `(49, 512)` patch tokens and a 512-d text embedding.
825
+ 4. Decode: `_decode_with_head` when a trained head is loaded, else `_decode_zero_shot`.
826
+ 5. Confidence from measurable objectness statistics.
827
+ 6. Evidence: `BOUNDING_BOX` items in `NORMALIZED_0_1`, a `GEOLOCATION` item in `GEO` when the CRS allows,
828
+ a `STATISTIC` item, and a `STATISTIC` record of any dropped degenerate candidates.
829
+
830
+ ### 9.2 There is no feature cache on the grounding serving path
831
+
832
+ Unlike the change-VQA and optical-SAR paths, grounding has **no serving-time feature cache**. The encoder
833
+ runs live on each request. The Phase 7 resolution experiment and the Phase 8 evaluation used **cached**
834
+ patch features (the resolution experiment's artifacts are `per_sample_224_full.jsonl`,
835
+ `per_sample_448_full.jsonl` and `resolution_experiment_full.json` β€” 16,159 lines each), and the single
836
+ decode implementation exists precisely so a cached-feature path and a live path *"cannot disagree"*:
837
+
838
+ `decode_candidates_from_features` (`specialists/grounding/inference.py:151`) is *"THE single
839
+ implementation of the zero-shot decode. `ground_phrase` encodes an image then calls this; the evaluation
840
+ script reads a feature cache then calls this."* The module records the defect this closed: the Phase 8
841
+ eval originally built its own single-box baseline with `argmax_candidate` while Phase 7 measured through
842
+ `ground_phrase` β€” same 16,159 records, same metric, same cached features, but a different **decode**:
843
+
844
+ ```
845
+ Phase 7 via ground_phrase : mean best IoU 0.0972
846
+ eval via argmax_candidate : mean best IoU 0.0092
847
+ ```
848
+
849
+ A 10Γ— gap, from a duplicated decode. *"The duplication was the defect, so the duplication is what is
850
+ removed: there is now exactly one place that turns similarity into boxes."*
851
+
852
+ ### 9.3 The decode
853
+
854
+ `similarity_map` L2-normalises both sides before the dot product, because *"without that the dot product is
855
+ dominated by whichever vector happens to have the larger norm, which is a property of the features rather
856
+ than of the match."* `decode_candidates_from_features` then produces a threshold box (patches within
857
+ `DEFAULT_DELTA = 0.02` of the peak) plus up to `top_k` local maxima (a patch must beat its 4-neighbours,
858
+ and adjacent already-accepted peaks are skipped). If nothing qualifies, it falls back to `argmax_candidate`.
859
+
860
+ The head path uses `decode_cell_relative` + `sigmoid(objectness)` + `nms(iou_threshold=0.50)`, capped at
861
+ the specialist's `max_candidates` (6 by default, not the config's 20). The degenerate-box guard
862
+ (`DEGENERATE_AREA_FRACTION = 0.9`) applies to the zero-shot path only.
863
+
864
+ ---
865
+
866
+ ## 10. The VQA / caption data path
867
+
868
+ The VQA and caption paths share the display decode and the deterministic quality gate:
869
+
870
+ 1. `load_image_array(asset.path)` β†’ `(H, W, 3)` uint8.
871
+ 2. The quality gate (`assess_image_quality`) β€” a `NOISE` or `INVALID_VALUES` verdict **blocks the VLM
872
+ call**; a `FLAT` or `TOO_SMALL` verdict is usable but degraded. The gate is **wired into the serving
873
+ path**, not merely defined: `specialists/vqa/inference.py:190` computes it and returns
874
+ `self._refuse(...)` when `not quality.is_usable`, *before* the model is constructed or called. The
875
+ comment there states the reason β€” *"The VLM will produce a fluent, specific, entirely fabricated scene
876
+ description for ANY input it is handed, including uniform noise, even when the prompt explicitly
877
+ instructs it to decline. Measured, not theorised. The guard therefore has to sit here, before the
878
+ model."*
879
+ 3. The processor builds `pixel_values` with `processor_longest_edge: 512` pinned, so a 512-px tile is
880
+ **not** upscaled and split. The measured contrast is recorded in `configs/base.yaml`: default β†’
881
+ `pixel_values (1, 17, 3, 512, 512)`, 1142 prompt tokens; pinned β†’ `pixel_values (1, 1, 3, 512, 512)`.
882
+ 4. Prompts **must** go through `processor.apply_chat_template()` β€” hand-written prompt strings raise
883
+ `ValueError` because SmolVLM requires one `<image>` token per image (finding F5-3).
884
+ 5. `max_images_per_call: 1` β€” the VLM sees one image per call.
885
+
886
+ The VLM adapter is loaded from `SATQUERY_VLM_ADAPTER` when set (the resolution order is: explicit
887
+ argument, then that env var, then none). The adapter's metrics are **usable** (exact_match 0.963) but its
888
+ status is **ACCEPTANCE-REJECTED** β€” USABLE β‰  ACCEPTED.
889
+
890
+ ---
891
+
892
+ ## 11. Determinism and reproducibility
893
+
894
+ ### 11.1 What is deterministic
895
+
896
+ | Stage | Determinism | Mechanism |
897
+ |---|---|---|
898
+ | `load_image_array` | **Deterministic** | Pure percentile stretch; no sampling |
899
+ | `to_grayscale_float` | **Deterministic** | Pure percentile stretch |
900
+ | `assess_image_quality` | **Deterministic** | *"No model, no randomness, no thresholds learned from data"* |
901
+ | `normalise_for_croma` | **Deterministic** | *"No sampling, no learned statistics, no RNG, and no dependence on batch composition"* |
902
+ | `infer_modality` | **Deterministic** | Pure function of band count and explicit label |
903
+ | Sensor adapter | **Deterministic** | Pure band placement + zero-fill |
904
+ | `assemble_fusion_input` | **Deterministic** | Concatenation with asserted order |
905
+ | `resize_to_canonical` | **Deterministic** | Bilinear interpolation |
906
+ | Registration measurement | **Deterministic** | `cv2.phaseCorrelate`; no sampling |
907
+ | Post-processing (morphology, components) | **Deterministic** | `cv2` fixed operations |
908
+ | `channel_dropout` | **Seeded** | Takes an `rng`; *"Seeded deliberately: an unseeded training transform makes a run unreproducible"* |
909
+ | Router training | **Seeded** | `project.seed: 42` |
910
+ | Fusion training | **Seeded** | Explicit seeding, per `train_change_head`'s discipline |
911
+
912
+ ### 11.2 The global seed
913
+
914
+ `configs/base.yaml` declares `project.seed: 42`. `Config.seed` returns
915
+ `int(self.get("project.seed", 42))`. Training loops take it from the config.
916
+
917
+ ### 11.3 The config hash
918
+
919
+ `Config.hash` is `sha256(json.dumps(self._data, sort_keys=True, default=str).encode())[:16]`. The frozen
920
+ value is **`78f1e3700da15aa1`**. It is recorded in every evaluation run, and `scripts/eval_change.py`
921
+ **refuses to score on drift** (exit 3). This is why several pipeline decisions live in code constants or
922
+ environment channels rather than in `configs/base.yaml`: adding a key moves the hash and detaches the
923
+ published benchmark from its config.
924
+
925
+ ### 11.4 Environment channels β€” two distinct classes
926
+
927
+ There are two **different** kinds of environment override, and conflating them would be wrong:
928
+
929
+ **(a) Pre-hash overrides β€” these DO change `Config.hash`.** `load_config` applies these into the data
930
+ dictionary **before** `Config(data)` is constructed:
931
+
932
+ | Variable | Effect |
933
+ |---|---|
934
+ | `SATQUERY_PRECISION` | Sets `training.precision` |
935
+ | `SATQUERY_TORCH_COMPILE` | Sets `deployment.torch_compile` |
936
+
937
+ Because they are merged before hashing, a run using either has a **different** `Config.hash` from the
938
+ frozen one.
939
+
940
+ **(b) Hash-exempt overrides β€” these do NOT change `Config.hash`.** They are read at use time, outside the
941
+ config object:
942
+
943
+ | Variable | Read by | Purpose |
944
+ |---|---|---|
945
+ | `SATQUERY_DEVICE` | `core/config.py::Config.device_preference` | Device selection |
946
+ | `SATQUERY_CROMA_USE_8_BIT` | `specialists/optical_sar/radiometry.py::resolve_use_8_bit` | The uint8 arm (A/B) |
947
+ | `SATQUERY_CROMA_CHECKPOINT` | `specialists/optical_sar/croma.py::resolve_checkpoint_path` | CROMA checkpoint path (first in resolution order) |
948
+ | `SATQUERY_VLM_ADAPTER` | `specialists/vqa/model.py` | VLM adapter path |
949
+
950
+ The distinction matters for reproducibility: a run that sets `SATQUERY_DEVICE` is comparable to a run that
951
+ does not (same config identity), while a run that sets `SATQUERY_PRECISION` is **not** (different hash).
952
+
953
+ ### 11.5 Checkpoint resolution order
954
+
955
+ `specialists/optical_sar/croma.py::resolve_checkpoint_path` resolves the CROMA checkpoint
956
+ **offline-first**: `SATQUERY_CROMA_CHECKPOINT` (when set and non-empty, and the path must exist β€”
957
+ otherwise it raises naming the variable), then the Hugging Face Hub cache. The module records that the
958
+ `use_8_bit` resolution order is *"`SATQUERY_CROMA_USE_8_BIT` and then `croma.use_8_bit`."*
959
+
960
+ ### 11.6 What "reproducible" means per artefact
961
+
962
+ | Artefact | Reproducible? | How |
963
+ |---|---|---|
964
+ | Display array from a raster | βœ… byte-identical | Deterministic stretch |
965
+ | Availability mask | βœ… byte-identical | Deterministic placement |
966
+ | `normalise_for_croma` output | βœ… byte-identical | Pure function |
967
+ | Change feature cache | βœ… given the same spec hash and detector state | Spec-guarded |
968
+ | Text feature cache | βœ… given the same encoder + revision | Spec-guarded |
969
+ | Fusion feature cache | βœ… given the same arm, resolution and config | Provenance-guarded |
970
+ | Router corpus cache | βœ… given the same encoder id + text set | Fingerprint-guarded |
971
+ | A trained head | ❌ not byte-reproducible across hardware | Seeded, but GPU nondeterminism applies |
972
+
973
+ `UNKNOWN β€” not established from the available evidence` whether bit-exact reproducibility of a trained
974
+ checkpoint is achieved across GPU runs; the code seeds explicitly but does not claim cross-device
975
+ determinism.
976
+
977
+ ---
978
+
979
+ ## 12. Where intermediate artefacts live
980
+
981
+ All generated artefacts live under `artifacts/`, which is *"the repository's generated-artefact root (not
982
+ source, not config, not `data/` inputs)."*
983
+
984
+ ### 12.1 Caches (reproducible from a corpus)
985
+
986
+ | Artefact | Path | Producer | Key | Reproducible? |
987
+ |---|---|---|---|---|
988
+ | Router corpus embeddings | `<cache_dir>/embeddings_<fingerprint[:16]>.npy` + `.json` | `router/train.py::embed_corpus_cached` | corpus fingerprint (`v1` + encoder id + max_length + sorted texts) | βœ… |
989
+ | Change features | `artifacts/change_vqa/change_features.npz` | `scripts/prepare_change_vqa.py` | `feature_spec_hash(256, "stanet:trained")` = `c801326f85a185f8` | βœ… |
990
+ | Question features | `artifacts/change_vqa/text_features_<split>.npz` | `scripts/prepare_change_vqa.py` | `text_spec_hash(...)` = `d2801ea1a314354a` (pinned revision) | βœ… |
991
+ | Scene targets | `artifacts/change_vqa/scene_targets.jsonl` | `scripts/prepare_change_vqa.py` | `change_vqa_preproc_v1` | βœ… |
992
+ | Scene manifest | `artifacts/change_vqa/scene_manifest.jsonl` | `scripts/prepare_change_vqa.py` | `change_vqa_preproc_v1` | βœ… |
993
+ | Fusion features (Arm A) | `artifacts/optical_sar/fusion_features/train.npz` + `.json` | `scripts/extract_fusion_features.py` | `CACHE_VERSION="v1"` + arm + config + checkpoint digest | βœ… |
994
+ | Fusion features (Arm B) | β€” | β€” | β€” | **NOT PRODUCED** |
995
+
996
+ **A naming caution.** There is **no** `fusion_features_armB` directory (or any `armB`/`arm_b` literal) in
997
+ the repository β€” a grep for those spellings returns nothing. The arm is a field *inside* the cache's JSON
998
+ sidecar and the run record, not a directory name; the Arm-B cache would be a **separate cache built under
999
+ the same `artifacts/optical_sar/fusion_features/` root**, distinguished by its provenance record. Since
1000
+ that cache has not been built, the root holds only the Arm-A `train.npz` + `train.json`.
1001
+
1002
+ ### 12.2 Trained artefacts (not reproducible byte-for-byte)
1003
+
1004
+ | Artefact | Path | Notes |
1005
+ |---|---|---|
1006
+ | Grounding head | `artifacts/grounding/remoteclip_grounding_v001/head.pt` | The shipped default; `DEFAULT_HEAD_PATH` in code |
1007
+ | Change detector | `artifacts/change/levir_change_v001/head.pt` | Test pooled IoU 0.8122; **not wired into serving by default** |
1008
+ | Change-VQA head | `artifacts/change_vqa/run/head.pt` | `change_vqa_head_v1`, 1,453,912 params |
1009
+ | Fusion head (production) | `artifacts/optical_sar/fusion_head_production_v001/` | `pre_registered_115_metric.json` |
1010
+
1011
+ ### 12.3 Evidence artefacts
1012
+
1013
+ | Artefact | Path |
1014
+ |---|---|
1015
+ | CROMA forward pass | `artifacts/optical_sar/croma_forward.json` |
1016
+ | CDVQA imagery verification | `artifacts/cdvqa/imagery_verification.json` |
1017
+ | CDVQA/SECOND overlap | `artifacts/cdvqa/second_overlap.json` |
1018
+ | CDVQA temporal order | `artifacts/cdvqa/temporal_order_evidence_v2.json` |
1019
+ | Resolution experiment | `per_sample_224_full.jsonl`, `per_sample_448_full.jsonl`, `resolution_experiment_full.json` |
1020
+ | Phase 12 selection manifest | `artifacts/phase12_selection/selection_manifest_seed10.jsonl` |
1021
+
1022
+ ### 12.4 What the serving path writes
1023
+
1024
+ The serving path writes only **server-side diagnostic artefacts**, and their refs are **always `null`** to
1025
+ the client (F-16): the change map (`change_map_<stem>.tif`) and the optical/SAR views
1026
+ (`optical_view_<stem>.png`, `sar_view_<stem>.png`). *"The map is still WRITTEN. It is the operator's
1027
+ diagnostic ... What changes is only that its location is a server-side fact, not a client-facing one."*
1028
+ No `artifact://` URI is fabricated in its place.
1029
+
1030
+ ---
1031
+
1032
+ ## 13. What is NOT part of the pipeline
1033
+
1034
+ Each item below was checked against the code. Where a capability is absent, it is named rather than
1035
+ implied.
1036
+
1037
+ | Capability | Status | Evidence |
1038
+ |---|---|---|
1039
+ | **Reprojection** | **NOT part of the pipeline** | `compare_crs` reports `requires_reprojection`; no code reprojects. |
1040
+ | **Cloud masking** | **NOT part of the pipeline** | No cloud/QA/SCL handling anywhere. |
1041
+ | **Atmospheric correction** | **NOT part of the pipeline** | Rasters are consumed as delivered. |
1042
+ | **Mosaicking** | **NOT part of the pipeline** | No composite or seamline code. |
1043
+ | **Tiling / top-K tile selection** | **DECLARED, not executed** | Read only by `core/config.py`'s `top_k_tiles <= max_tiles` check. |
1044
+ | **Percentile / dB conditioning** | **DECLARED, not executed** | Read by no code; `radiometry.py` docstring; PHASE14 Β§1. |
1045
+ | **Nodata filling** | **NOT part of the pipeline** | `quality.py` refuses non-finite input; the fix is *"deferred to the tiling work."* |
1046
+ | **Pair-dimension reconciliation** | **NOT part of the pipeline** | The change path requires equal shapes; no reconciling resize. |
1047
+ | **A serving-time feature cache** | **NOT part of the pipeline** | Only the router's training corpus cache and the change-VQA/fusion preparation caches exist; serving encodes live. |
1048
+ | **Cross-step artefact hand-off** | **NOT part of the pipeline** | The change-VQA specialist *"accepts an optional `change_map` path in `request.params` for a future planner to populate. It is unused today"* β€” *"the controller does not hand artifacts between steps."* |
1049
+ | **A single-pass change+language request** | **NOT part of the pipeline** | Routing a change+language request plans **both** a `change` and a `change_vqa` step, so *"the detector runs twice for one request"* β€” the weights load once (the registry caches the instance), but the forward pass runs twice. |
1050
+ | **Image resolution decoding** | **NOT part of the pipeline** | CDVQA's `res_x`/`res_y` are the opaque strings `".1524m"` on every row; *"The field is of unknown semantics and is stored as an opaque string. It is not used anywhere."* |
1051
+
1052
+ ---
1053
+
1054
+ ## 14. What is NOT RUN / OPEN / BLOCKED for this topic
1055
+
1056
+ ### NOT RUN
1057
+
1058
+ - **The router test split.** Reported accuracy is **0.965116 on validation, ungated, n = 86**; the test
1059
+ split was **NOT RUN**.
1060
+ - **The optical-SAR forward pass with a trained head at serving time.** CROMA is not loaded locally and no
1061
+ trained head is wired; the specialist runs sensor-only.
1062
+ - **The CROMA normalisation experiment (option c).** Not run; the arm set is frozen at A/B and arm B (the
1063
+ registered one) is non-constructible.
1064
+ - **An end-to-end system benchmark.** No system-level accuracy is claimed.
1065
+ - **A change+language request end-to-end.** The planner would run the detector twice; the artefact
1066
+ hand-off that would make it once is not implemented.
1067
+
1068
+ ### OPEN
1069
+
1070
+ - **The percentile/dB conditioning stage.** Either implement it or retire the keys.
1071
+ - **The arm-constructibility question.** A human decision, deliberately unresolved.
1072
+ - **The optical-SAR ruling.** Accuracy 0.931 with macro-F1 0.434161; **OPEN**. Never quote the accuracy
1073
+ without the macro-F1.
1074
+ - **The change-VQA metric ruling.** `OPEN` and owner-gated (two test sets: test 0.697626/0.378373 and
1075
+ test2 0.651469/0.372309).
1076
+ - **Calibration.** ECE went **0.013755 β†’ 0.014929 β€” worse**. Retained only because it is in the frozen
1077
+ config.
1078
+ - **The VLM adapter.** Metrics usable (exact_match 0.963) but status **ACCEPTANCE-REJECTED**.
1079
+
1080
+ ### BLOCKED
1081
+
1082
+ - **The Arm-B fusion feature cache.** `artifacts/optical_sar/fusion_features/` holds only the Arm-A cache;
1083
+ `val.npz` and `test.npz` were pending. Arm B cannot be trained without its own cache.
1084
+ - **A clean change demo.** Blocked on a same-shape pair or a resize step.
1085
+
1086
+ ### DEFERRED
1087
+
1088
+ - **`codespace_name` trailing `\n`** in the `/api/health` payload (B-02). Cosmetic; the wake path is safe.
1089
+ - **Nodata filling from the raster profile.** Deferred to the tiling work, which is not implemented.
1090
+
1091
+ ---
1092
+
1093
+ ## 15. Where the evidence lives
1094
+
1095
+ **Source modules (the pipeline this chapter describes):**
1096
+
1097
+ | Path | What it defines |
1098
+ |---|---|
1099
+ | `preprocessing/raster.py` | `inspect_raster`, `read_bands`, `write_raster`, `infer_modality`, `file_sha256`, `pixel_area_m2` |
1100
+ | `preprocessing/imagery.py` | `load_image_array`, `load_pil_image` β€” the display decode |
1101
+ | `preprocessing/quality.py` | `assess_image_quality`, the deterministic gate |
1102
+ | `specialists/optical_sar/sensor_adapter.py` | Band mapping, zero-fill, availability mask |
1103
+ | `specialists/optical_sar/radiometry.py` | `normalise_for_croma`, `resolve_use_8_bit`, the zero-channel rule |
1104
+ | `specialists/optical_sar/inference.py` | `run_pipeline`, `resize_to_canonical`, `build_masks` |
1105
+ | `specialists/optical_sar/fusion_head.py` | `assemble_fusion_input`, `channel_dropout`, `expected_fusion_dim` |
1106
+ | `specialists/optical_sar/specialist.py` | `execute`, `_descriptor_for`, `_confidence_components` |
1107
+ | `specialists/change/specialist.py` | `_assess_pair`, `_predict`, `_write_artifacts` |
1108
+ | `specialists/change/postprocess.py` | `measure_registration`, `to_grayscale_float`, `postprocess_change_map` |
1109
+ | `specialists/change/vqa_specialist.py` | `feature_spec_mismatch`, `unavailable_reason`, `execute` |
1110
+ | `specialists/grounding/specialist.py` | `execute`, `_decode_with_head`, `_to_geo_boxes` |
1111
+ | `specialists/grounding/inference.py` | `decode_candidates_from_features`, `ground_phrase` |
1112
+ | `training/change_vqa/features.py` | `feature_spec_hash`, `text_spec_hash`, `ChangeFeatureExtractor`, the caches |
1113
+ | `training/change_vqa/dataset.py` | `DATASET_ID`, `PREPROCESSING_VERSION`, `SPLIT_MAP`, `LABEL_PALETTE` |
1114
+ | `training/fusion/train.py` | `CACHE_VERSION`, `DEFAULT_FEATURE_CACHE_DIR`, `feature_cache_path`, `write_feature_cache`, `ARMS` |
1115
+ | `training/fusion/extract.py` | The Phase 12 producer, the DEV-2 stretch, `_assert_raw_encoder` |
1116
+ | `router/encoder.py` | `FrozenEncoder`, the verified MiniLM contract |
1117
+ | `router/train.py` | `CORPUS_CACHE_VERSION`, `_corpus_fingerprint`, `embed_corpus_cached` |
1118
+ | `core/config.py` | `load_config`, `Config.hash`, `device_preference`, the env overrides |
1119
+ | `core/controller.py` | `_inspect_asset`, the nine-state machine |
1120
+
1121
+ **Scripts (the producers):**
1122
+
1123
+ - `scripts/prepare_change_vqa.py` β€” scene targets, manifest, change features, text features.
1124
+ - `scripts/extract_fusion_features.py` β€” the Phase 12 fusion feature cache (exit codes 0/2/3).
1125
+ - `scripts/train_fusion.py`, `scripts/train_change_vqa.py`, `scripts/train_grounding.py`,
1126
+ `scripts/train_router.py` β€” the training loops.
1127
+ - `scripts/diagnose_feature_cache.py` β€” cache diagnosis.
1128
+
1129
+ **Configuration:**
1130
+
1131
+ - `configs/base.yaml` β€” `project.seed`, `image.*`, `optical.*`, `sar.*`, `croma.*`, `fusion.*`,
1132
+ `grounding.*`, `change.*`, `router.*`, `vlm.*`, `training.*`.
1133
+
1134
+ **Documents:**
1135
+
1136
+ - `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` β€” the two ordered stages, the Gate F arm amendment, the
1137
+ Arm-B blocker.
1138
+ - `docs/PHASE12_ENTRY_GATE.md`, `docs/PHASE12_R08_TRUTHFUL_BACKFILL.md` β€” the Phase 12 run records and the
1139
+ `result_status` plumbing.
1140
+ - `docs/PHASE10_CDVQA_DATA_STATUS.md` β€” the CDVQA/SECOND layout, the temporal semantics, the four adapter
1141
+ defects found by execution.
1142
+ - `docs/PHASE7_RESOLUTION_DECISION.md` β€” the 224 decision (the cached-feature experiment).
1143
+ - `docs/API_CONTRACT.md` Β§2.5 β€” the asset allowlist and the upload contract.
1144
+ - `docs/FINAL_DELIVERY_REPORT.md` Β§4 (run ids), Β§5 (metrics), Β§6 (the 726Β²/736Β² degraded change demo).
1145
+
1146
+ **Sibling release chapters:**
1147
+
1148
+ - [`GEOSPATIAL.md`](GEOSPATIAL.md) β€” the raster contract, CRS, coordinate systems, sensor adapter, tiling
1149
+ policy, the 224 decision.
1150
+ - [`architecture/03-request-lifecycle.md`](architecture/03-request-lifecycle.md) β€” the nine controller
1151
+ states.
1152
+ - [`architecture/05-specialists.md`](architecture/05-specialists.md) β€” the six specialists.
1153
+ - [`architecture/07-configuration-freeze.md`](architecture/07-configuration-freeze.md) β€” the frozen hash
1154
+ and every config key.
1155
+ - [`DATASETS.md`](DATASETS.md), [`TRAINING.md`](TRAINING.md), [`REPRODUCIBILITY.md`](REPRODUCIBILITY.md).