Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
release: add docs/DATA_PIPELINE.md
Browse files- docs/DATA_PIPELINE.md +1155 -0
docs/DATA_PIPELINE.md
ADDED
|
@@ -0,0 +1,1155 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Data pipeline β deep reference
|
| 2 |
+
|
| 3 |
+
**Status tags used on every substantive claim:** `IMPLEMENTED` Β· `VERIFIED` Β· `MEASURED` Β· `ATTEMPTED`
|
| 4 |
+
Β· `NOT RUN` Β· `BLOCKED` Β· `DEFERRED` Β· `REJECTED` Β· `OPEN` Β· `RESOLVED` Β· `CLOSED`.
|
| 5 |
+
|
| 6 |
+
Every statement in this document was produced by reading a file in `C:/Users/anish/satquery-ai/`. Where a
|
| 7 |
+
fact is not established from the available evidence, this document writes
|
| 8 |
+
`UNKNOWN β not established from the available evidence`. Where a stage is *declared* in configuration but
|
| 9 |
+
read by no code, that is stated explicitly rather than presented as a working step. Where a cache spec
|
| 10 |
+
hash is cited, it was **recomputed** from the code that produces it, not copied from a report.
|
| 11 |
+
|
| 12 |
+
This chapter is the reference for the **end-to-end data path**: from an uploaded asset to the tensor a
|
| 13 |
+
specialist consumes, including every decode, every normalisation, every cache, and every intermediate
|
| 14 |
+
artefact. It is written to be sufficient to reconstruct the pipeline from the document alone.
|
| 15 |
+
|
| 16 |
+
**Sibling documents:** [`GEOSPATIAL.md`](GEOSPATIAL.md) (the raster contract, CRS, coordinate systems, the
|
| 17 |
+
sensor adapter, the tiling policy and the 224 decision), [`architecture/03-request-lifecycle.md`](architecture/03-request-lifecycle.md)
|
| 18 |
+
(the nine controller states), [`architecture/05-specialists.md`](architecture/05-specialists.md) (the six
|
| 19 |
+
specialists), [`architecture/07-configuration-freeze.md`](architecture/07-configuration-freeze.md) (the
|
| 20 |
+
frozen config identity), [`DATASETS.md`](DATASETS.md) (the corpora), [`REPRODUCIBILITY.md`](REPRODUCIBILITY.md)
|
| 21 |
+
(reproduction recipes), [`TRAINING.md`](TRAINING.md) (how each corpus is consumed).
|
| 22 |
+
|
| 23 |
+
---
|
| 24 |
+
|
| 25 |
+
## Table of contents
|
| 26 |
+
|
| 27 |
+
0. [How to read this document](#0-how-to-read-this-document)
|
| 28 |
+
1. [The end-to-end path](#1-the-end-to-end-path)
|
| 29 |
+
2. [Validation and decoding](#2-validation-and-decoding)
|
| 30 |
+
3. [Modality inference, and where it is consumed](#3-modality-inference-and-where-it-is-consumed)
|
| 31 |
+
4. [The normalisation stages](#4-the-normalisation-stages)
|
| 32 |
+
5. [Tiling and tile selection](#5-tiling-and-tile-selection)
|
| 33 |
+
6. [The change-VQA data path](#6-the-change-vqa-data-path)
|
| 34 |
+
7. [The optical-SAR data path](#7-the-optical-sar-data-path)
|
| 35 |
+
8. [The router data path](#8-the-router-data-path)
|
| 36 |
+
9. [The grounding data path](#9-the-grounding-data-path)
|
| 37 |
+
10. [The VQA / caption data path](#10-the-vqa--caption-data-path)
|
| 38 |
+
11. [Determinism and reproducibility](#11-determinism-and-reproducibility)
|
| 39 |
+
12. [Where intermediate artefacts live](#12-where-intermediate-artefacts-live)
|
| 40 |
+
13. [What is NOT part of the pipeline](#13-what-is-not-part-of-the-pipeline)
|
| 41 |
+
14. [What is NOT RUN / OPEN / BLOCKED for this topic](#14-what-is-not-run--open--blocked-for-this-topic)
|
| 42 |
+
15. [Where the evidence lives](#15-where-the-evidence-lives)
|
| 43 |
+
|
| 44 |
+
---
|
| 45 |
+
|
| 46 |
+
## 0. How to read this document
|
| 47 |
+
|
| 48 |
+
The pipeline has **two distinct halves**, and almost every confusion about it comes from conflating them:
|
| 49 |
+
|
| 50 |
+
1. **The serving path** β an uploaded asset becomes a specialist result, request by request. This is what
|
| 51 |
+
runs in production. It is *stateless* with respect to data: it reads a raster, prepares it, and calls
|
| 52 |
+
a model. No cache is consulted except the router's, and no intermediate artefact is required for a
|
| 53 |
+
result to be produced.
|
| 54 |
+
2. **The preparation / training path** β a corpus on disk becomes a *cache* of frozen features, and a head
|
| 55 |
+
is trained over that cache. This runs offline, once per corpus revision. It is where the
|
| 56 |
+
`mean Β± 2Β·std` radiometric stretch actually executes, where the `change_feat_v1` feature vector is
|
| 57 |
+
computed, and where the arm (A/B) is baked in.
|
| 58 |
+
|
| 59 |
+
The caches are the bridge: the training path produces them, and the **serving path does not read them**
|
| 60 |
+
(except the router's corpus cache, which is a training-time artefact, and the change-VQA text cache, which
|
| 61 |
+
is likewise training-time). A reader who assumes the serving path is cache-driven will misread the system;
|
| 62 |
+
the caches exist so that a *frozen* encoder is not re-run across an entire corpus on every training
|
| 63 |
+
experiment.
|
| 64 |
+
|
| 65 |
+
Three further distinctions are load-bearing and are stated at each point below:
|
| 66 |
+
|
| 67 |
+
- **Declared versus executed.** `configs/base.yaml` declares the tiling policy, the optical percentile
|
| 68 |
+
stretch and the SAR dB clip. All three are read by no code on the serving path.
|
| 69 |
+
- **Deterministic versus seeded.** Some stages are pure functions of their input (the display stretch, the
|
| 70 |
+
quality gate, the radiometric stretch). Others are stochastic and take a seeded RNG (channel dropout).
|
| 71 |
+
The distinction determines what "reproducible" means for each.
|
| 72 |
+
- **Cache-hit versus cache-miss.** Every cache in the project is keyed by a **spec hash** or a
|
| 73 |
+
**fingerprint**, and a mismatch is a **miss** or a **refusal**, never a silent stale read.
|
| 74 |
+
|
| 75 |
+
---
|
| 76 |
+
|
| 77 |
+
## 1. The end-to-end path
|
| 78 |
+
|
| 79 |
+
### 1.1 The nine controller states
|
| 80 |
+
|
| 81 |
+
The controller runs a fixed nine-state machine, declared in `configs/base.yaml` under `agent.states` and
|
| 82 |
+
mirrored by the `ControllerState` enum in `core/schemas.py:79`:
|
| 83 |
+
|
| 84 |
+
```
|
| 85 |
+
RECEIVE β PARSE β VALIDATE β PLAN β PREPROCESS β EXECUTE β AGGREGATE β VERIFY β RESPOND
|
| 86 |
+
```
|
| 87 |
+
|
| 88 |
+
The division of labour is stated once and holds throughout: *the router understands, the policy decides,
|
| 89 |
+
the specialists compute, the VLM explains, and the evidence proves.* The three states that matter for the
|
| 90 |
+
data pipeline are VALIDATE, PREPROCESS and EXECUTE.
|
| 91 |
+
|
| 92 |
+
### 1.2 The path, drawn
|
| 93 |
+
|
| 94 |
+
```
|
| 95 |
+
POST /v1/assets (gateway; five-type allowlist, two-layer size cap)
|
| 96 |
+
β
|
| 97 |
+
ββ asset_id (opaque, 128-bit; never a path)
|
| 98 |
+
β
|
| 99 |
+
βΌ
|
| 100 |
+
POST /v1/analyze { assets: [...], query }
|
| 101 |
+
β
|
| 102 |
+
ββ RECEIVE ββ resolve handles to server-side paths
|
| 103 |
+
β
|
| 104 |
+
ββ PARSE ββ router: query text β Intent (task, modality, temporal, spatial, language)
|
| 105 |
+
β
|
| 106 |
+
ββ VALIDATE ββ core/controller.py::_inspect_asset
|
| 107 |
+
β ββ preprocessing.raster.inspect_raster(path, max_pixels=..., ...)
|
| 108 |
+
β header-only: dims β bands β dtype β CRS β transform β bounds
|
| 109 |
+
β β nodata β modality β AssetMetadata
|
| 110 |
+
β
|
| 111 |
+
ββ PLAN ββ policy selects a workflow (which specialists, in what order)
|
| 112 |
+
β
|
| 113 |
+
ββ PREPROCESS ββ per-specialist preparation (this chapter's core)
|
| 114 |
+
β ββ decode: read_bands β (C, H, W); load_image_array β (H, W, 3) uint8
|
| 115 |
+
β ββ sensor adapter: band map + zero-fill + availability mask
|
| 116 |
+
β ββ normalisation: display stretch / encoder-input stretch
|
| 117 |
+
β ββ resize: to the model's canonical resolution
|
| 118 |
+
β
|
| 119 |
+
ββ EXECUTE ββ specialist forward pass β SpecialistResult
|
| 120 |
+
β
|
| 121 |
+
ββ AGGREGATE ββ combine specialist results
|
| 122 |
+
ββ VERIFY ββ consistency checks
|
| 123 |
+
ββ RESPOND ββ ResultEnvelope { run_id, result, trace }
|
| 124 |
+
```
|
| 125 |
+
|
| 126 |
+
`architecture/03-request-lifecycle.md` records a fact worth repeating here: the **PREPROCESS state is
|
| 127 |
+
never recorded** as a distinct `TraceStep` in the execution trace; the work happens inside each
|
| 128 |
+
specialist's `execute` rather than as a controller-level step. So the trace shows VALIDATE and EXECUTE but
|
| 129 |
+
not the preparation between them, and the preparation's parameters reach the trace through the
|
| 130 |
+
specialist's **evidence payloads** (the availability masks, the radiometry report, the decode metadata)
|
| 131 |
+
rather than through a PREPROCESS step.
|
| 132 |
+
|
| 133 |
+
### 1.3 What each state reads and writes
|
| 134 |
+
|
| 135 |
+
| State | Reads | Writes | Cache consulted |
|
| 136 |
+
|---|---|---|---|
|
| 137 |
+
| RECEIVE | the uploaded asset handle | a server-side path | none |
|
| 138 |
+
| PARSE | the query text | an `Intent` | router corpus cache (training only) |
|
| 139 |
+
| VALIDATE | raster **header** only | `AssetMetadata` | none |
|
| 140 |
+
| PLAN | the `Intent`, the `AssetMetadata` list | a workflow plan | none |
|
| 141 |
+
| PREPROCESS | raster **pixels** | model-ready arrays | none |
|
| 142 |
+
| EXECUTE | model-ready arrays | a `SpecialistResult` | none (serving) |
|
| 143 |
+
| AGGREGATE / VERIFY / RESPOND | specialist results | the envelope + trace | none |
|
| 144 |
+
|
| 145 |
+
The single most important cell in that table is VALIDATE's "raster header only": the inspection that
|
| 146 |
+
decides whether an asset is usable **does not decode it**. `inspect_raster`'s docstring states this β
|
| 147 |
+
*"Validate and describe a raster without loading its pixel data"* β and it is what keeps a 25-megapixel
|
| 148 |
+
rejection cheap.
|
| 149 |
+
|
| 150 |
+
---
|
| 151 |
+
|
| 152 |
+
## 2. Validation and decoding
|
| 153 |
+
|
| 154 |
+
### 2.1 Header-only validation β `inspect_raster`
|
| 155 |
+
|
| 156 |
+
Covered in full in [`GEOSPATIAL.md`](GEOSPATIAL.md) Β§1.3. The pipeline-relevant points:
|
| 157 |
+
|
| 158 |
+
- It reads dimensions, band count, dtypes, CRS, transform, bounds, nodata, resolution, driver and a tiled
|
| 159 |
+
flag, and constructs a `GeoMetadata` plus an `AssetMetadata`.
|
| 160 |
+
- The **pixel budget** is enforced here: `width * height > max_pixels` raises `OversizedImageError`, which
|
| 161 |
+
is recoverable β the caller may downscale. `max_pixels` comes from `image.max_pixels` (25,000,000).
|
| 162 |
+
- The **modality** is assigned here by `infer_modality` (Β§3).
|
| 163 |
+
- The **SHA-256** is computed here only when `compute_hash=True` (streaming, 1 MiB chunks) β it is needed
|
| 164 |
+
for manifests and dedup detection, and it is the field the change specialist uses to detect a repeated
|
| 165 |
+
acquisition.
|
| 166 |
+
|
| 167 |
+
### 2.2 Band decoding β `read_bands`
|
| 168 |
+
|
| 169 |
+
`preprocessing/raster.py::read_bands(path, *, indexes=None, out_dtype=None)` returns
|
| 170 |
+
`(bands, H, W)` plus a profile that retains `crs`, `transform` and `bounds`. The profile is the source of
|
| 171 |
+
georeferencing for every downstream write (change maps, masks) so *"downstream code never loses
|
| 172 |
+
georeferencing."* It accepts an optional `indexes` list and an optional `out_dtype` cast, and wraps every
|
| 173 |
+
failure as `RasterReadError`.
|
| 174 |
+
|
| 175 |
+
Consumers:
|
| 176 |
+
|
| 177 |
+
- `preprocessing/imagery.py::load_image_array` β for the VQA, caption and grounding paths.
|
| 178 |
+
- `specialists/optical_sar/specialist.py::execute` β reads the optical and SAR arrays raw, then hands them
|
| 179 |
+
to the sensor adapter.
|
| 180 |
+
- `specialists/change/specialist.py::_write_artifacts` β reads T1's profile so the change map is written
|
| 181 |
+
with T1's georeferencing.
|
| 182 |
+
|
| 183 |
+
### 2.3 Display decoding β `load_image_array`
|
| 184 |
+
|
| 185 |
+
`preprocessing/imagery.py::load_image_array` turns a raster into a displayable `(H, W, 3)` uint8 array:
|
| 186 |
+
|
| 187 |
+
1. `read_bands` β `(bands, H, W)`.
|
| 188 |
+
2. Transpose to `(H, W, bands)`; take the first three bands as RGB; repeat a single band three times; or
|
| 189 |
+
stack a 2-D array three times.
|
| 190 |
+
3. Per-**scene** 2/98 percentile stretch over finite values, clip to `[0, 1]`, scale to uint8.
|
| 191 |
+
|
| 192 |
+
Its docstring is explicit about what it does **not** do: *"It does not resample, crop, or reproject. Those
|
| 193 |
+
change the pixel grid, and the grounding specialist converts normalized boxes to pixel coordinates using
|
| 194 |
+
the ORIGINAL raster's dimensions β a silent resize here would put every box in the wrong place."* And
|
| 195 |
+
about determinism: *"The percentile stretch is deterministic: the same file always yields the same array,
|
| 196 |
+
so a grounding box and a VQA answer describe identical pixels."*
|
| 197 |
+
|
| 198 |
+
`load_pil_image` wraps the same array in a `PIL.Image`, and notes that georeferencing is not lost because
|
| 199 |
+
the caller's `AssetMetadata` already carries it.
|
| 200 |
+
|
| 201 |
+
### 2.4 The deterministic quality gate
|
| 202 |
+
|
| 203 |
+
`preprocessing/quality.py` is *"a deterministic input-quality gate"* that sits **upstream of the model**.
|
| 204 |
+
It exists because of a measured failure: a loaded SmolVLM-500M-Instruct, given uniform random noise and a
|
| 205 |
+
prompt that explicitly said *"answer only from what is visible"*, produced *"a fluent, specific, entirely
|
| 206 |
+
fabricated scene description."* The gate's rationale is stated directly: *"A 500M-parameter VLM will
|
| 207 |
+
describe *something* for any input it is given, and asking it to self-assess reliably is asking it to do
|
| 208 |
+
the thing it just failed at."*
|
| 209 |
+
|
| 210 |
+
The discriminating signal is **lag-1 spatial autocorrelation**, not variance. The measured separation:
|
| 211 |
+
|
| 212 |
+
| Input | Autocorrelation |
|
| 213 |
+
|---|---|
|
| 214 |
+
| Real remote-sensing imagery | 0.6 β 0.99 |
|
| 215 |
+
| Uniform random noise | ~0.00 |
|
| 216 |
+
| A constant (blank) image | undefined; variance ~0 |
|
| 217 |
+
|
| 218 |
+
Variance alone cannot separate a flat desert scene from an all-zero tile (both have near-zero variance);
|
| 219 |
+
autocorrelation is high for both uniform-but-real imagery and textured imagery, and collapses only for
|
| 220 |
+
noise.
|
| 221 |
+
|
| 222 |
+
The constants:
|
| 223 |
+
|
| 224 |
+
| Constant | Value | Meaning |
|
| 225 |
+
|---|---|---|
|
| 226 |
+
| `MIN_AUTOCORRELATION` | `0.10` | Below this, the array is not spatially coherent |
|
| 227 |
+
| `FLAT_STD_EPSILON` | `1e-6` | Below this normalised std, the array is constant |
|
| 228 |
+
| `NOISE_ENTROPY_BITS` | `7.8` | Shannon entropy above which the histogram is noise-like |
|
| 229 |
+
| `BLOCKING_VERDICTS` | `{NOISE, INVALID_VALUES}` | Verdicts that must block a VLM call |
|
| 230 |
+
|
| 231 |
+
`assess_image_quality` returns an `ImageQuality` with one of five verdicts:
|
| 232 |
+
|
| 233 |
+
| Verdict | Meaning | Effect |
|
| 234 |
+
|---|---|---|
|
| 235 |
+
| `STRUCTURED` | Spatial structure consistent with real imagery | Safe to analyse |
|
| 236 |
+
| `FLAT` | Effectively constant | Usable, but `is_degraded` β lower confidence |
|
| 237 |
+
| `NOISE` | Uncorrelated, near-uniform | **Blocks the VLM call** |
|
| 238 |
+
| `TOO_SMALL` | Fewer than 16 px on a side | Usable, flagged |
|
| 239 |
+
| `INVALID_VALUES` | NaN or infinite values present | **Blocks the VLM call** |
|
| 240 |
+
|
| 241 |
+
Two behaviours are worth quoting. First, the non-finite rule is deliberately **tolerance-free**: *"a fixed
|
| 242 |
+
fraction is size-dependent, so one NaN in a 64x64 array (0.99976) would pass while one NaN in a 10x10
|
| 243 |
+
array (0.99) would fail. The same defect must not be tolerated or rejected depending on image
|
| 244 |
+
dimensions."* Any non-finite value is disqualifying. Second, the entropy check is described as a
|
| 245 |
+
**redundant second signal**: *"MEASURED, and the margin here is thin: structured ramp+texture gives 7.581,
|
| 246 |
+
uniform noise gives 7.988. That is a 0.22-bit gap against a 7.8 threshold. The autocorrelation check is
|
| 247 |
+
what actually carries this gate (0.952 vs -0.008, a 0.96 margin against a 0.10 threshold)."*
|
| 248 |
+
|
| 249 |
+
The gate is a pure function β *"Nothing here is learned, sampled, or probabilistic. Same input, same
|
| 250 |
+
verdict."* A known limitation is recorded: a float GeoTIFF whose nodata sentinel is NaN lands in
|
| 251 |
+
`INVALID_VALUES` and is refused, and the comment records that *"Filling nodata from the raster profile
|
| 252 |
+
belongs in the tiling work"* β i.e. the fix is deferred to a stage that does not exist (Β§5).
|
| 253 |
+
|
| 254 |
+
### 2.5 Error types
|
| 255 |
+
|
| 256 |
+
Every failure in the validation and decoding path is a typed `SatQueryError` subclass. The relevant ones
|
| 257 |
+
for this pipeline:
|
| 258 |
+
|
| 259 |
+
| Error | Raised by | Meaning |
|
| 260 |
+
|---|---|---|
|
| 261 |
+
| `RasterReadError` | `inspect_raster`, `read_bands`, `write_raster` | The file is not a readable raster |
|
| 262 |
+
| `OversizedImageError` | `inspect_raster` | Exceeds the pixel budget; **recoverable** |
|
| 263 |
+
| `UnsupportedBandsError` | `inspect_raster`, the sensor adapter | Band count is zero, or bands cannot be placed |
|
| 264 |
+
| `MissingCRSError` | `parse_crs` | A CRS string is malformed |
|
| 265 |
+
| `CoordinateError` | `geospatial/transform.py` | A conversion lacks the context it needs |
|
| 266 |
+
| `SpecialistError` | specialists | A forward pass or preparation failed |
|
| 267 |
+
|
| 268 |
+
The design rule from `preprocessing/raster.py` applies throughout: *"Never raise a bare exception. Every
|
| 269 |
+
failure is a typed SatQueryError."*
|
| 270 |
+
|
| 271 |
+
---
|
| 272 |
+
|
| 273 |
+
## 3. Modality inference, and where it is consumed
|
| 274 |
+
|
| 275 |
+
`infer_modality(band_count, explicit=None)` (`preprocessing/raster.py:58`) is covered in full in
|
| 276 |
+
[`GEOSPATIAL.md`](GEOSPATIAL.md) Β§6. Its pipeline role is as the **first consumer** of the band count that
|
| 277 |
+
VALIDATE read:
|
| 278 |
+
|
| 279 |
+
- `inspect_raster` calls it and stores the result on `AssetMetadata.modality`.
|
| 280 |
+
- The optical-SAR specialist's `_assess_pair` reads the **declared** modality first and falls back to
|
| 281 |
+
`infer_modality(asset.geo.band_count)` only when it is `UNKNOWN`, appending a warning that names the
|
| 282 |
+
heuristic.
|
| 283 |
+
- The change specialist does **not** infer modality: it requires exactly two assets and compares their
|
| 284 |
+
CRS and registration, without asserting a modality for either.
|
| 285 |
+
|
| 286 |
+
The heuristic's limits are stated in its own comment: *"These are heuristics, not ground truth β the
|
| 287 |
+
sensor adapter is authoritative when a sensor descriptor is supplied."* A band count is not a sensor
|
| 288 |
+
declaration, and the adapter is the only component that can map a band to a canonical channel.
|
| 289 |
+
|
| 290 |
+
---
|
| 291 |
+
|
| 292 |
+
## 4. The normalisation stages
|
| 293 |
+
|
| 294 |
+
There are **three** distinct normalisation transforms in this codebase, plus one **declared but
|
| 295 |
+
unimplemented** stage. Conflating them is the most common misreading, so they are laid out side by side.
|
| 296 |
+
|
| 297 |
+
### 4.1 The display stretch β per-scene 2/98
|
| 298 |
+
|
| 299 |
+
**Where:** `preprocessing/imagery.py::load_image_array` (and its grayscale cousin in
|
| 300 |
+
`specialists/change/postprocess.py::to_grayscale_float`).
|
| 301 |
+
|
| 302 |
+
**What:** one `(lo, hi)` computed over **all finite values across all bands** of one scene, then
|
| 303 |
+
`(arr - lo) / (hi - lo)`, clipped to `[0, 1]`.
|
| 304 |
+
|
| 305 |
+
**Output:** uint8 `(H, W, 3)` for display (or float32 grayscale for registration).
|
| 306 |
+
|
| 307 |
+
**Purpose:** presentation β so a VQA answer and a grounding box describe identical pixels.
|
| 308 |
+
|
| 309 |
+
**Determinism:** pure; same file β same array.
|
| 310 |
+
|
| 311 |
+
**Not** the CROMA transform. `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§1.1 names three disqualifying
|
| 312 |
+
differences: per-scene rather than per-channel; uint8 3-band rather than float 12-channel; presentation
|
| 313 |
+
rather than radiometry. Reusing it for CROMA is explicitly **forbidden** by that ruling (Β§3.3).
|
| 314 |
+
|
| 315 |
+
### 4.2 The encoder-input stretch β per-channel `mean Β± 2Β·std`
|
| 316 |
+
|
| 317 |
+
**Where:** `specialists/optical_sar/radiometry.py::normalise_for_croma`.
|
| 318 |
+
|
| 319 |
+
**What:** for each channel `c` of each sample, `lower = mean_c - 2Β·std_c`, `upper = mean_c + 2Β·std_c`,
|
| 320 |
+
then `(x - lower) / (upper - lower)`, optionally through a uint8 round-trip (`Γ255`, clip, quantise,
|
| 321 |
+
`/255`). Computed **per sample** (not per batch) with `ddof=1`.
|
| 322 |
+
|
| 323 |
+
**Output:** float32 of the same shape, in `[0, 1]`.
|
| 324 |
+
|
| 325 |
+
**Purpose:** the transform the CROMA authors' own README instructs users to apply before the frozen
|
| 326 |
+
encoder.
|
| 327 |
+
|
| 328 |
+
**Determinism:** *"pure and deterministic: no sampling, no learned statistics, no RNG, and no dependence
|
| 329 |
+
on batch composition. Same input -> same output."*
|
| 330 |
+
|
| 331 |
+
**Where it runs:** the **extraction** path (`training/fusion/extract.py`), not the serving path (Β§7.5).
|
| 332 |
+
The extraction module applies it *"because the arm's `use_8_bit` has to be applied at exactly one
|
| 333 |
+
place"*, and it refuses an encoder that would stretch a second time.
|
| 334 |
+
|
| 335 |
+
**The zero-channel rule:** an all-zero channel (the sensor adapter's unavailability convention) has
|
| 336 |
+
`std = 0` and would divide by zero. Such channels are **skipped and left at exactly zero**, with status
|
| 337 |
+
`unavailable`; a present-but-constant channel is skipped with status `degenerate`. This is the C-1
|
| 338 |
+
discipline applied to radiometry β *"the same rule that forbids inventing a band forbids inventing a
|
| 339 |
+
dynamic range for a band that does not exist."*
|
| 340 |
+
|
| 341 |
+
### 4.3 The declared-but-unimplemented conditioning stage
|
| 342 |
+
|
| 343 |
+
**Where declared:** `configs/base.yaml`:
|
| 344 |
+
|
| 345 |
+
```yaml
|
| 346 |
+
optical:
|
| 347 |
+
normalization: percentile
|
| 348 |
+
lower_percentile: 2
|
| 349 |
+
upper_percentile: 98
|
| 350 |
+
sar:
|
| 351 |
+
representation: db
|
| 352 |
+
clip_min_db: -30
|
| 353 |
+
clip_max_db: 5
|
| 354 |
+
```
|
| 355 |
+
|
| 356 |
+
**Status:** `DECLARED (not read)`. No code applies a percentile/dB conditioning stage. The radiometry
|
| 357 |
+
module's docstring states the consequence: *"the two `optical.*` and three `sar.*` config keys are still
|
| 358 |
+
read by no code. That is recorded, not fixed."* `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§6 item 3
|
| 359 |
+
adds: *"A config key that no code reads is worse than a missing key: it reads as a satisfied
|
| 360 |
+
requirement."*
|
| 361 |
+
|
| 362 |
+
### 4.4 The side-by-side
|
| 363 |
+
|
| 364 |
+
| Transform | Granularity | Output | Purpose | Implemented | Runs on serving path |
|
| 365 |
+
|---|---|---|---|---|---|
|
| 366 |
+
| Display 2/98 | per-scene, all bands | uint8 (H,W,3) | Presentation | β
| β
|
|
| 367 |
+
| Encoder-input `mean Β± 2Β·std` | per-channel, per-sample | float32 `[0,1]` | Frozen-encoder interface | β
| β (extraction only) |
|
| 368 |
+
| Change grayscale 2/98 | per-scene, grayscale | float32 `[0,1]` | Registration correlation | β
| β
|
|
| 369 |
+
| Percentile / dB conditioning | per-scene, per config | unbounded rescale | `base.yaml`'s intent | β | β |
|
| 370 |
+
|
| 371 |
+
---
|
| 372 |
+
|
| 373 |
+
## 5. Tiling and tile selection
|
| 374 |
+
|
| 375 |
+
Covered in full in [`GEOSPATIAL.md`](GEOSPATIAL.md) Β§10. The pipeline-relevant summary:
|
| 376 |
+
|
| 377 |
+
- `image.max_pixels: 25000000` **is enforced**, at `inspect_raster` time, raising `OversizedImageError`.
|
| 378 |
+
- `image.tile_size: 512`, `image.tile_overlap: 128`, `image.max_tiles: 64`, `image.top_k_tiles: 4` are
|
| 379 |
+
**declared and read only by** `core/config.py`'s single relationship check
|
| 380 |
+
(`top_k_tiles <= max_tiles`). No serving-path code slices a raster into tiles or selects a top-K.
|
| 381 |
+
- `change.tile_size: 256`, `change.tile_overlap: 0` are likewise declared and not consumed as a tiling
|
| 382 |
+
step; 256 is the change model's training resolution.
|
| 383 |
+
|
| 384 |
+
**Consequence for the pipeline:** a raster within the pixel budget is handed to a specialist **whole**.
|
| 385 |
+
The only size-reduction steps that actually run are per-specialist resizes to a model's canonical
|
| 386 |
+
resolution (e.g. `resize_to_canonical` to 120Γ120 for CROMA, and the change feature extractor's bilinear
|
| 387 |
+
resize to 256Γ256). Neither is tiling: both resample the whole image rather than selecting a region.
|
| 388 |
+
|
| 389 |
+
The VLM path's `processor_longest_edge: 512` is a **control against** the processor's default 2048, not a
|
| 390 |
+
tiling step. `core/config.py` enforces `processor_longest_edge <= image.tile_size`, with the measured
|
| 391 |
+
rationale recorded: the default 2048 upscales a 512 tile 4Γ and splits it into ~17 sub-images, *"the real
|
| 392 |
+
figure is ~17x"* against the plan's estimated 4Γ.
|
| 393 |
+
|
| 394 |
+
---
|
| 395 |
+
|
| 396 |
+
## 6. The change-VQA data path
|
| 397 |
+
|
| 398 |
+
The change-VQA path (`Task.CHANGE_VQA`, R-02) is the most elaborate in the project: it has a frozen
|
| 399 |
+
detector, two frozen feature extractors, a trained head, and four cache artefacts. It is the clearest
|
| 400 |
+
illustration of the serving/preparation split.
|
| 401 |
+
|
| 402 |
+
### 6.1 The dataset and the preprocessing version tag
|
| 403 |
+
|
| 404 |
+
`training/change_vqa/dataset.py` defines the identity constants:
|
| 405 |
+
|
| 406 |
+
```python
|
| 407 |
+
DATASET_ID = "cdvqa"
|
| 408 |
+
PREPROCESSING_VERSION = "change_vqa_preproc_v1"
|
| 409 |
+
SPLIT_MAP: dict[str, str] = {
|
| 410 |
+
"Train": "train",
|
| 411 |
+
"Val": "val",
|
| 412 |
+
"Test": "test",
|
| 413 |
+
"Test2": "test",
|
| 414 |
+
}
|
| 415 |
+
```
|
| 416 |
+
|
| 417 |
+
`PREPROCESSING_VERSION` is *"bumped when the record shape or the target definition changes. Recorded in
|
| 418 |
+
every evaluation so two numbers computed under different definitions can never be silently compared."*
|
| 419 |
+
It is the **preprocessing version tag** for this path: it names the *record and target* definition, and is
|
| 420 |
+
distinct from the *feature spec* hashes below, which name the *tensor* definition.
|
| 421 |
+
|
| 422 |
+
`SPLIT_MAP` is a leakage control, not a convenience: `Test2` maps to `test` because Test2 is *"a SECOND
|
| 423 |
+
QUESTION SET over the SAME 968 scenes as Test, not an independent sample; it must never be pooled."*
|
| 424 |
+
`docs/PHASE10_CDVQA_DATA_STATUS.md` Β§4 measured this: Test and Test2 share **100% of their images** (968 /
|
| 425 |
+
968), while Train/Val/Test are cleanly disjoint.
|
| 426 |
+
|
| 427 |
+
`LABEL_PALETTE` decodes the six change classes (NVG surface, trees, low vegetation, water, buildings,
|
| 428 |
+
playgrounds) from a fixed colour palette; white `(255,255,255)` is a **shared background**, not a class.
|
| 429 |
+
|
| 430 |
+
`SceneTargets` derives supervision from `label1`/`label2` with four definitions, all fractions of total
|
| 431 |
+
pixels: `class_mag[c]` (magnitude), `class_delta[c]` (signed), `class_ratio[c]` (per-class ratio) and
|
| 432 |
+
`total_changed`. The `total_changed` definition was **validated against gold answers**: over 40 real
|
| 433 |
+
`change_ratio` rows, `count(label1 != label2) / total` reproduced the annotated bin **34/40 = 85%** of the
|
| 434 |
+
time, versus **1/40** for the competing definition. The residual 15% is recorded as *"annotator/map
|
| 435 |
+
disagreement, which is the noise floor of the target"* rather than smoothed away.
|
| 436 |
+
|
| 437 |
+
### 6.2 The change feature extractor
|
| 438 |
+
|
| 439 |
+
`training/change_vqa/features.py` defines the frozen change feature. The module explains why the features
|
| 440 |
+
are frozen: the plan budgets CDVQA at β€ 3 h on T4Γ2, and training end-to-end through STANet would
|
| 441 |
+
re-learn a representation the project already has. Freezing it drops the trainable parameter count from
|
| 442 |
+
~15.8 M to ~1.5 M.
|
| 443 |
+
|
| 444 |
+
`CHANGE_FEATURE_DIM = 1045`, and it is **three things**:
|
| 445 |
+
|
| 446 |
+
| Part | Dim | Composition | Answers |
|
| 447 |
+
|---|---|---|---|
|
| 448 |
+
| `level_pooled` | 1024 | 4 levels Γ 128 ch Γ {mean, max} | *"what does the change look like locally?"* |
|
| 449 |
+
| `change_stats` | 5 | mean, std, frac>0.3, frac>0.5, frac>0.7 of the change map | *"how much of the scene changed?"* |
|
| 450 |
+
| `change_grid` | 16 | adaptive 4Γ4 pool of the same map | *"WHERE did it change?"* |
|
| 451 |
+
|
| 452 |
+
The dimensions are **derived, never typed** β *"a literal here would drift silently from the pooling code
|
| 453 |
+
below and produce a shape error three files away"*:
|
| 454 |
+
|
| 455 |
+
```python
|
| 456 |
+
N_LEVELS = 4
|
| 457 |
+
LEVEL_CHANNELS = 128
|
| 458 |
+
LEVEL_POOLED_DIM = N_LEVELS * LEVEL_CHANNELS * 2 # 1024
|
| 459 |
+
CHANGE_STATS_DIM = 2 + len(CHANGE_STAT_THRESHOLDS) # 5
|
| 460 |
+
CHANGE_GRID_DIM = CHANGE_GRID * CHANGE_GRID # 16
|
| 461 |
+
CHANGE_FEATURE_DIM = LEVEL_POOLED_DIM + CHANGE_STATS_DIM + CHANGE_GRID_DIM # 1045
|
| 462 |
+
```
|
| 463 |
+
|
| 464 |
+
The spatial parts are load-bearing: *"Without `change_grid` the head would see a scene-global average and
|
| 465 |
+
could not distinguish 'buildings changed in one corner' from 'buildings changed everywhere', which is
|
| 466 |
+
precisely the difference between `largest_change` and `smallest_change`."*
|
| 467 |
+
|
| 468 |
+
The default working resolution is `DEFAULT_IMAGE_SIZE = 256` β *"the change model's OWN training
|
| 469 |
+
resolution (`configs/base.yaml: change.tile_size: 256`; LEVIR-CD patches are 256x256). Feeding the native
|
| 470 |
+
512 would push the frozen encoder outside the distribution it was trained on."* 512 is offered as an
|
| 471 |
+
explicit, recorded alternative.
|
| 472 |
+
|
| 473 |
+
`ChangeFeatureExtractor.extract(t1, t2)` re-runs the detector's own submodules to capture the intermediate
|
| 474 |
+
fused levels that `STANetStyleChangeDetector.forward` does not return. That is *"a second implementation
|
| 475 |
+
of one forward pass, so it is a drift hazard β and it is closed by a test that asserts the probability map
|
| 476 |
+
produced here is **bit-identical** to `STANetStyleChangeDetector.forward(...).probabilities`."*
|
| 477 |
+
|
| 478 |
+
### 6.3 The text feature extractor
|
| 479 |
+
|
| 480 |
+
`TextFeatureExtractor` wraps `sentence-transformers/all-MiniLM-L6-v2` β the same model the router uses, so
|
| 481 |
+
*"the text side of the reasoning head adds no new model to the deployment."* It produces `TEXT_FEATURE_DIM
|
| 482 |
+
= 384` normalised vectors and raises `FeatureExtractionError` if the encoder's output width disagrees with
|
| 483 |
+
the declared constant.
|
| 484 |
+
|
| 485 |
+
### 6.4 The cache spec hashes
|
| 486 |
+
|
| 487 |
+
Two independent hash functions key the two caches. Both are `sha256(...)[:16]`.
|
| 488 |
+
|
| 489 |
+
`feature_spec_hash(image_size, *, extractor)` hashes the feature definition:
|
| 490 |
+
`{spec, image_size, extractor, levels, level_channels, thresholds, grid, text_encoder}`, sorted keys.
|
| 491 |
+
|
| 492 |
+
`text_spec_hash(*, encoder, revision=None)` hashes the question-feature definition:
|
| 493 |
+
`{spec: "change_vqa_text_v1", encoder, revision, dim, normalised}`, sorted keys. The separate hash is
|
| 494 |
+
deliberate: *"One combined hash would force a full re-extraction of both whenever either changed, and β
|
| 495 |
+
worse β would let a stale text cache pass a check it should fail."*
|
| 496 |
+
|
| 497 |
+
**Recomputed values** (this document computed them from the functions above, with the constants as read
|
| 498 |
+
from the source):
|
| 499 |
+
|
| 500 |
+
| Input | Spec hash |
|
| 501 |
+
|---|---|
|
| 502 |
+
| `feature_spec_hash(256, extractor="stanet:trained")` | `c801326f85a185f8` |
|
| 503 |
+
| `feature_spec_hash(256, extractor="stanet:untrained")` | `714efac5b0e6a6d1` |
|
| 504 |
+
| `text_spec_hash(encoder="sentence-transformers/all-MiniLM-L6-v2", revision="1110a243fdf4")` | `d2801ea1a314354a` |
|
| 505 |
+
| `text_spec_hash(encoder="sentence-transformers/all-MiniLM-L6-v2")` (unpinned) | `f87f5226c4c4a69b` |
|
| 506 |
+
|
| 507 |
+
So the change-feature cache spec for the trained detector at 256 px is **`c801326f85a185f8`**, and the
|
| 508 |
+
question-feature cache spec with the pinned MiniLM revision (`configs/base.yaml: router.revision:
|
| 509 |
+
1110a243fdf4`) is **`d2801ea1a314354a`**. The spec hash **encodes trained-vs-untrained**, which is how the
|
| 510 |
+
serving specialist detects a detector-state mismatch (Β§6.6).
|
| 511 |
+
|
| 512 |
+
### 6.5 Cache layout, resume, and the artefacts
|
| 513 |
+
|
| 514 |
+
`scripts/prepare_change_vqa.py` is the producer. Its default output directory is
|
| 515 |
+
`DEFAULT_OUT = REPO_ROOT / "artifacts" / "change_vqa"`, and it writes:
|
| 516 |
+
|
| 517 |
+
| Artefact | Shape | Content |
|
| 518 |
+
|---|---|---|
|
| 519 |
+
| `<out>/scene_manifest.jsonl` | one record per `(native split, scene)` | identity, integrity, split, question counts |
|
| 520 |
+
| `<out>/scene_targets.jsonl` | one record per scene | label-derived `SceneTargets` |
|
| 521 |
+
| `<out>/change_features.npz` | `(n_scenes, 1045)` + spec hash + extractor config | frozen STANet features, keyed by `scene_key` |
|
| 522 |
+
| `<out>/text_features_<split>.npz` | `(n_questions, 384)` + spec hash | question features, keyed by `question_id` |
|
| 523 |
+
|
| 524 |
+
The manifest is **scene-level, deliberately**: *"153,130 question rows would produce a ~50 MB JSONL that
|
| 525 |
+
duplicates text already authoritative in `annotations/*.json`. The manifest therefore records the **2,968
|
| 526 |
+
scenes** ... and the question-level records are derived from the annotations on demand."*
|
| 527 |
+
|
| 528 |
+
Resume is spec-guarded. `ChangeFeatureCache.read(path, *, expect_spec_hash=...)` **refuses** a cache built
|
| 529 |
+
under a different spec:
|
| 530 |
+
|
| 531 |
+
```python
|
| 532 |
+
if expect_spec_hash is not None and cache.spec_hash != expect_spec_hash:
|
| 533 |
+
raise FeatureExtractionError(
|
| 534 |
+
f"change feature cache at {p} was built under spec "
|
| 535 |
+
f"{cache.spec_hash!r} but the caller expects {expect_spec_hash!r}; "
|
| 536 |
+
f"re-extract rather than training on a representation that does not match serving"
|
| 537 |
+
)
|
| 538 |
+
```
|
| 539 |
+
|
| 540 |
+
`TextFeatureCache.read` applies the same guard. The producer calls
|
| 541 |
+
`ChangeFeatureCache.read(cache_path, expect_spec_hash=extractor.spec_hash)` when resuming, so a
|
| 542 |
+
resolution change or a detector-state change invalidates the cache instead of silently mixing two
|
| 543 |
+
representations.
|
| 544 |
+
|
| 545 |
+
A scene whose imagery cannot be read is **skipped and reported** through `on_progress`, never silently
|
| 546 |
+
dropped: *"a scene missing from the cache is a training sample the model never sees, and the caller must be
|
| 547 |
+
able to count how many that was."*
|
| 548 |
+
|
| 549 |
+
### 6.6 The serving path, and the spec-mismatch refusal
|
| 550 |
+
|
| 551 |
+
`ChangeVQASpecialist` (`specialists/change/vqa_specialist.py`) assembles three pieces: a trained head, a
|
| 552 |
+
change feature extractor, and a question text encoder. Its `execute`:
|
| 553 |
+
|
| 554 |
+
1. Validates exactly two assets, both present.
|
| 555 |
+
2. Calls `unavailable_reason()`; if anything is missing or mismatched, returns `_unavailable_result` with
|
| 556 |
+
**`answer=""`** and `degraded=True`.
|
| 557 |
+
3. Otherwise loads the two images, extracts `(vector, facts)` and the text vector, resolves the question
|
| 558 |
+
type and temporal reference, and calls `predict_answers`.
|
| 559 |
+
|
| 560 |
+
**Why degraded produces no answer.** The module docstring is explicit: *"a random change map is visibly
|
| 561 |
+
noise, whereas an untrained 19-way classifier still emits a fluent, confident-looking `yes`. A user cannot
|
| 562 |
+
tell the second from a real answer, so the untrained case returns **no answer at all**."* The
|
| 563 |
+
`SpecialistResult` validator independently marks an empty change-VQA answer as degraded.
|
| 564 |
+
|
| 565 |
+
**The spec-mismatch refusal.** `feature_spec_mismatch()` compares the head's trained spec
|
| 566 |
+
(`head_metadata["change_cache_spec"]`) with the serving extractor's `spec_hash`. A mismatch is a
|
| 567 |
+
**refusal, not a warning**: *"A head is only meaningful for the feature distribution it was fitted on."*
|
| 568 |
+
The measured reason this matters: `configs/base.yaml`'s `change:` section carries **no `checkpoint_path`
|
| 569 |
+
key**, so the registry builds the change feature extractor with `checkpoint_path=None` and gets an
|
| 570 |
+
**untrained** STANet β while `scripts/prepare_change_vqa.py` defaults to the trained LEVIR checkpoint.
|
| 571 |
+
Training and serving would therefore consume different representations, and *"the head would still emit a
|
| 572 |
+
fluent `yes`."* The spec hash covers this because it encodes trained-vs-untrained, so the refusal catches
|
| 573 |
+
the detector-state mismatch as well as a resolution mismatch.
|
| 574 |
+
|
| 575 |
+
Other serving constants: `LOW_CONFIDENCE_THRESHOLD = 0.40` (a reporting threshold, not a calibration) and
|
| 576 |
+
`TEMPORAL_ORDER_NOTE` (T1=pre, T2=post; `label1=pre/label2=post` **proven** at agreement 1.0000 over 2,968
|
| 577 |
+
scenes; `im1=pre/im2=post` **supported statistically, not proven**).
|
| 578 |
+
|
| 579 |
+
**A config fact worth recording.** The specialist reads `change_vqa.image_size` (default 256),
|
| 580 |
+
`change_vqa.apply_type_mask` (default `True`) and `change_vqa.head_path` via `config.get(...)`, but
|
| 581 |
+
**`configs/base.yaml` has no `change_vqa` block at all**. Those keys therefore always take their code
|
| 582 |
+
defaults, and the frozen config does not name them. This is the same pattern as the grounding head's
|
| 583 |
+
`DEFAULT_HEAD_PATH` (a code constant rather than a config key, to avoid moving `Config.hash`).
|
| 584 |
+
|
| 585 |
+
---
|
| 586 |
+
|
| 587 |
+
## 7. The optical-SAR data path
|
| 588 |
+
|
| 589 |
+
### 7.1 The serving path, step by step
|
| 590 |
+
|
| 591 |
+
`OpticalSarSpecialist.execute` (`specialists/optical_sar/specialist.py:320`):
|
| 592 |
+
|
| 593 |
+
1. **Validate.** Exactly two assets; `_assess_pair` establishes one optical and one SAR (Β§9.2 of
|
| 594 |
+
[`GEOSPATIAL.md`](GEOSPATIAL.md)).
|
| 595 |
+
2. **Descriptors.** `_descriptor_for` returns the asset's `SensorDescriptor` if present, else builds one
|
| 596 |
+
from the band count β `build_optical_adapter` for optical, `positional_fallback_descriptor` for SAR β
|
| 597 |
+
and appends a warning naming the assumption.
|
| 598 |
+
3. **Read pixels.** `read_bands(optical_asset.path)` and `read_bands(sar_asset.path)`.
|
| 599 |
+
4. **Pipeline.** `run_pipeline(...)` (below).
|
| 600 |
+
5. **Confidence.** `_confidence_components` computes the four plan-section-26 components.
|
| 601 |
+
6. **Answer + evidence.** `_compose_answer` builds prose only from computed facts; `_build_evidence`
|
| 602 |
+
emits the availability masks, the view items, the fused-representation item and the decision.
|
| 603 |
+
|
| 604 |
+
### 7.2 `run_pipeline`
|
| 605 |
+
|
| 606 |
+
`specialists/optical_sar/inference.py::run_pipeline` is the seam between "an image on disk" and "three GAP
|
| 607 |
+
vectors plus two masks":
|
| 608 |
+
|
| 609 |
+
```
|
| 610 |
+
optical raster -> sensor adapter -> (12, H, W) + mask[12]
|
| 611 |
+
SAR raster -> sensor adapter -> ( 2, H, W) + mask[2]
|
| 612 |
+
resize to CROMA's 120x120
|
| 613 |
+
CROMA(SAR_images=..., optical_images=...) <- no mask, ever
|
| 614 |
+
assemble_fusion_input <- mask enters HERE
|
| 615 |
+
fusion head -> logits -> probabilities
|
| 616 |
+
```
|
| 617 |
+
|
| 618 |
+
Two properties are load-bearing:
|
| 619 |
+
|
| 620 |
+
- **Degradation is a first-class outcome.** `run_pipeline` *"NEVER raises for a missing model: the
|
| 621 |
+
degradation is reported, because 'we have no trained head' is a fact about the deployment, not a failure
|
| 622 |
+
of the request."* It raises only for a caller error β arrays that cannot be canonicalised at all.
|
| 623 |
+
- **The sensor side always runs.** *"The masks, the canonical channel placement, the availability
|
| 624 |
+
statistics and the modality validation are all real work that requires no weights, and they are exactly
|
| 625 |
+
the facts that tell an operator why a result is or is not trustworthy on Cartosat-2S + RISAT."*
|
| 626 |
+
|
| 627 |
+
`resize_to_canonical(array, resolution)` bilinearly resizes a `(C, H, W)` array to
|
| 628 |
+
`(C, resolution, resolution)`. Its docstring notes that nearest-neighbour would be wrong for continuous
|
| 629 |
+
optical reflectance channels, while for zero-filled channels *"interpolating zeros gives zeros."* It uses
|
| 630 |
+
`cv2.INTER_LINEAR` when OpenCV is available, falling back to `torch.nn.functional.interpolate`.
|
| 631 |
+
|
| 632 |
+
### 7.3 The mask path
|
| 633 |
+
|
| 634 |
+
`build_masks` converts the two `SensorAdapterOutput` masks to `(1, C)` float32, ready for concatenation.
|
| 635 |
+
`mask_availability_stats` computes the fractions and counts that feed both the confidence components and
|
| 636 |
+
the evidence. The mask **never reaches CROMA** (finding C-1) β the specialist's evidence item records
|
| 637 |
+
`"croma_received_mask": false` and `"mask_consumed_by": "fusion_head"`.
|
| 638 |
+
|
| 639 |
+
### 7.4 The fusion cache β `fusion_features`
|
| 640 |
+
|
| 641 |
+
`training/fusion/train.py` and `training/fusion/extract.py` implement the Phase 12 frozen-feature trainer
|
| 642 |
+
and its missing producer. The pipeline:
|
| 643 |
+
|
| 644 |
+
```
|
| 645 |
+
paired patches -> normalise -> resize -> CROMA (frozen) -> FusionFeature
|
| 646 |
+
-> write_feature_cache -> train_fusion_head
|
| 647 |
+
```
|
| 648 |
+
|
| 649 |
+
The cache constants:
|
| 650 |
+
|
| 651 |
+
| Constant | Value | Meaning |
|
| 652 |
+
|---|---|---|
|
| 653 |
+
| `CACHE_VERSION` | `"v1"` | Bumped when the encoder, resolution or stored fields change |
|
| 654 |
+
| `DEFAULT_FEATURE_CACHE_DIR` | `artifacts/optical_sar/fusion_features` | Under `artifacts/`, scoped to this specialist, **not** under the frozen `artifacts/change/` tree |
|
| 655 |
+
|
| 656 |
+
`feature_cache_path(cache_dir=None, tag="")` returns
|
| 657 |
+
`base / f"croma_features_{CACHE_VERSION}{suffix}.npz"`, where the suffix is `_<tag>` when a tag is given
|
| 658 |
+
(e.g. a split name). The Phase-12 CLI (`scripts/extract_fusion_features.py`) instead takes an explicit
|
| 659 |
+
`--out-cache`, documented in its own usage as `artifacts/optical_sar/fusion_features/train.npz`, and writes
|
| 660 |
+
a `.npz` plus a `.json` sidecar. So both naming conventions are in play: the training loop's default is
|
| 661 |
+
`croma_features_v1_<tag>.npz`, and the CLI's documented example is `<split>.npz`. Both are
|
| 662 |
+
`artifacts/optical_sar/fusion_features/` files, and `verify_feature_cache` checks `row_counts_agree` and
|
| 663 |
+
`sidecar_n_matches_rows`.
|
| 664 |
+
|
| 665 |
+
**The write is atomic-ish.** `write_feature_cache` writes both files to unique temporary names and renames
|
| 666 |
+
them into place with `os.replace`, so *"no reader ever sees a half-written `.npz`."* The two renames cannot
|
| 667 |
+
be one operation; a crash in the window leaves a complete-but-stale pair, which `verify_feature_cache`
|
| 668 |
+
**reports** rather than letting a truncated array load cleanly and be silently wrong. A crashed run may
|
| 669 |
+
leave a `.<name>.tmp-<pid>-<uuid>.npz` behind, and the function deliberately never deletes it β *"a delete
|
| 670 |
+
is not a safe operation to perform on someone else's filesystem."*
|
| 671 |
+
|
| 672 |
+
**The multi-label collapse is an explicit policy.** BigEarthNet v2.0 is multi-label; the head is a
|
| 673 |
+
single-label 19-class softmax. The collapse is a parameter, `label_policy: Sequence[str] -> int | None`,
|
| 674 |
+
and with no policy: exactly one label β its index; more than one β `AmbiguousLabelError`; zero β
|
| 675 |
+
`None` β the sample is **skipped and counted**, never mapped to class 0 (*"'we do not know' and 'arable
|
| 676 |
+
land' are different facts and must not be merged"*).
|
| 677 |
+
|
| 678 |
+
**Resume is provenance-guarded.** Re-running with an existing cache loads it, skips the `sample_id`s
|
| 679 |
+
already present and appends the rest (`existing + new`, never a replacement). But *"appending rows from a
|
| 680 |
+
different arm, resolution or config would produce a cache whose metadata describes only half its rows"* β
|
| 681 |
+
so `check_resume_provenance` refuses that and names the field and both values; `resume=False` rebuilds.
|
| 682 |
+
|
| 683 |
+
**The dry run makes the same decision, not a similar one.** `plan_extraction` calls the same
|
| 684 |
+
`select_samples` the run calls, so `plan.n_selected` is `result.n_encoded` and `plan.n_would_write` is
|
| 685 |
+
`result.n_total`.
|
| 686 |
+
|
| 687 |
+
**Exit codes.** `scripts/extract_fusion_features.py` returns `0` on success (including a dry run), `2` for
|
| 688 |
+
a pre-flight or input error (nothing was encoded β a missing corpus, a leaked split, an unwritable
|
| 689 |
+
`--out-cache`, a missing CROMA checkpoint, an ambiguous label under `require_single_label`), and `3` for a
|
| 690 |
+
run that **started and then failed** with a typed error (e.g. a cached feature that is not 2318 wide). The
|
| 691 |
+
comment on the `FusionTrainingError` branch states the distinction: *"the run started, so this is 3, not a
|
| 692 |
+
pre-flight 2."*
|
| 693 |
+
|
| 694 |
+
### 7.5 The normalisation arm β A / B
|
| 695 |
+
|
| 696 |
+
The Phase 12 trainer selects the normalisation arm through the **existing hash-exempt environment
|
| 697 |
+
channel** `SATQUERY_CROMA_USE_8_BIT`, **not** by editing `configs/base.yaml`:
|
| 698 |
+
|
| 699 |
+
```python
|
| 700 |
+
ARM_CONTROL = "A"
|
| 701 |
+
ARM_VARIANT = "B"
|
| 702 |
+
|
| 703 |
+
ARMS: dict[str, Arm] = {
|
| 704 |
+
ARM_CONTROL: Arm(ARM_CONTROL, True, "... use_8_bit=true ..."),
|
| 705 |
+
ARM_VARIANT: Arm(ARM_VARIANT, False, "... use_8_bit=false ..."),
|
| 706 |
+
}
|
| 707 |
+
```
|
| 708 |
+
|
| 709 |
+
The reason for the env channel is stated: *"Adding a `fusion_training:` block would move the frozen
|
| 710 |
+
`Config.hash` off `78f1e3700da15aa1` and detach the Phase 9 benchmark from its config."* `apply_arm` writes
|
| 711 |
+
the value and returns it so a caller can record it.
|
| 712 |
+
|
| 713 |
+
**A warning that must be repeated, not summarised.** The `B` in this dict is **NOT** the arm B registered
|
| 714 |
+
in `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§4. Both arms run the **same** per-channel `mean Β± 2Β·std`
|
| 715 |
+
encoder-input stretch; they differ **only** on whether the result then takes the uint8 round-trip. The
|
| 716 |
+
axis is the **uint8 axis**, not the registered "stretch vs no-stretch" axis. The originally preregistered
|
| 717 |
+
arm B (*"percentile/dB conditioning only, no encoder-input stretch"*) is **NOT CONSTRUCTIBLE** from this
|
| 718 |
+
repository: stage 1 is unimplemented, and `Arm` exposes no field that can express "skip the encoder-input
|
| 719 |
+
stretch" β the stretch in `training/fusion/extract.py` is applied unconditionally. There is **no arm C**.
|
| 720 |
+
|
| 721 |
+
**Blocker recorded, not crossed.** The arm is baked into the cached features, and
|
| 722 |
+
`artifacts/optical_sar/fusion_features/` holds **only the Arm-A cache** (`train.npz` and `train.json`;
|
| 723 |
+
`val.npz` and `test.npz` were pending as of the Phase-14 correction notes). Arm B therefore needs its own
|
| 724 |
+
feature cache before it can be trained.
|
| 725 |
+
|
| 726 |
+
### 7.6 The extraction module's double-stretch guard
|
| 727 |
+
|
| 728 |
+
`training/fusion/extract.py` applies the DEV-2 stretch itself, and refuses an encoder that would apply it
|
| 729 |
+
again:
|
| 730 |
+
|
| 731 |
+
> `CROMAEncoder` applies the same stretch inside `encode` when `normalize_input=True` (its default).
|
| 732 |
+
> Passing such an encoder here would stretch twice β a silent distribution shift that produces a perfectly
|
| 733 |
+
> well-shaped cache full of wrong numbers. `_assert_raw_encoder` refuses it by name rather than trusting
|
| 734 |
+
> the caller; `build_extraction_encoder` constructs the raw encoder this module expects.
|
| 735 |
+
|
| 736 |
+
The CLI prints `normalize_input : <value> (must be False)`, so the guard is visible in the run output as
|
| 737 |
+
well as enforced in code.
|
| 738 |
+
|
| 739 |
+
---
|
| 740 |
+
|
| 741 |
+
## 8. The router data path
|
| 742 |
+
|
| 743 |
+
### 8.1 The frozen encoder
|
| 744 |
+
|
| 745 |
+
`router/encoder.py::FrozenEncoder` wraps `SentenceTransformer` and holds it frozen. The verified contract
|
| 746 |
+
(Phase 4 probe, 2026-09-16, sentence-transformers 6.0.1) is recorded at the top of the module:
|
| 747 |
+
|
| 748 |
+
```
|
| 749 |
+
SentenceTransformer(
|
| 750 |
+
model_name_or_path='sentence-transformers/all-MiniLM-L6-v2',
|
| 751 |
+
revision='1110a243fdf4', # pinned, verified reachable
|
| 752 |
+
device='cpu',
|
| 753 |
+
)
|
| 754 |
+
.get_sentence_embedding_dimension() -> 384
|
| 755 |
+
.tokenizer.model_max_length -> 256 (finding F4-1)
|
| 756 |
+
.max_seq_length -> 256
|
| 757 |
+
params -> 22,713,216
|
| 758 |
+
encode(64 queries, CPU) -> 0.118 s
|
| 759 |
+
```
|
| 760 |
+
|
| 761 |
+
Two constants are declared: `VERIFIED_TOKENIZER_MAX_LENGTH = 256` and `VERIFIED_EMBEDDING_DIM = 384`. The
|
| 762 |
+
constructor **refuses** a `max_length` above 256 β *"truncation would be a silent no-op"* β and sets
|
| 763 |
+
`self._model.max_seq_length = max_length` so the configured truncation (128, from
|
| 764 |
+
`router.max_length`) is *"enforced rather than merely documented."* The encoder returns detached numpy
|
| 765 |
+
arrays with no autograd history, which is what makes cached-embedding training possible (finding F4-2).
|
| 766 |
+
|
| 767 |
+
### 8.2 The corpus embedding cache
|
| 768 |
+
|
| 769 |
+
`router/train.py` is the consumer. `CORPUS_CACHE_VERSION = "v1"`, and `_corpus_fingerprint(corpus,
|
| 770 |
+
encoder_id, max_length)` hashes:
|
| 771 |
+
|
| 772 |
+
```python
|
| 773 |
+
payload = {
|
| 774 |
+
"version": CORPUS_CACHE_VERSION,
|
| 775 |
+
"encoder": encoder_id, # f"{encoder.model_name}@{encoder.revision}"
|
| 776 |
+
"max_length": max_length,
|
| 777 |
+
"texts": sorted(e.text for e in corpus),
|
| 778 |
+
}
|
| 779 |
+
return hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
|
| 780 |
+
```
|
| 781 |
+
|
| 782 |
+
`embed_corpus_cached` writes two files into the cache directory:
|
| 783 |
+
|
| 784 |
+
| File | Content |
|
| 785 |
+
|---|---|
|
| 786 |
+
| `embeddings_<fingerprint[:16]>.npy` | the `(n, 384)` float32 embedding matrix |
|
| 787 |
+
| `embeddings_<fingerprint[:16]>.json` | `{fingerprint, encoder, max_length, count, dim, seconds}` |
|
| 788 |
+
|
| 789 |
+
The cache key includes the encoder id **and** the corpus text set, so *"changing either invalidates it. A
|
| 790 |
+
stale cache silently training on the wrong vectors would be worse than no cache."* A cache hit requires
|
| 791 |
+
both `cached.shape[0] == len(corpus)` **and** `meta.get("fingerprint") == fingerprint`. A bad cache is a
|
| 792 |
+
**miss, not an error**: `except Exception: pass` falls through to re-encoding.
|
| 793 |
+
|
| 794 |
+
**This cache is a training-time artefact.** The serving path does not read it. `embed_corpus_cached` is
|
| 795 |
+
defined and used only in `router/train.py` (defined at `:232`, used at `:518`), while the serving path
|
| 796 |
+
calls `self.router.route(request.query)` (`core/controller.py:476`), which encodes the query live β a
|
| 797 |
+
single short string, measured at 0.118 s for 64 queries on CPU, so one query is sub-millisecond.
|
| 798 |
+
|
| 799 |
+
### 8.3 The router's training discipline
|
| 800 |
+
|
| 801 |
+
The module records two findings that shape the data path:
|
| 802 |
+
|
| 803 |
+
- **F4-2 (frozen encoder β cached training).** *"The encoder is FROZEN. Embeddings are therefore a pure
|
| 804 |
+
function of the query text. So we embed the whole corpus ONCE, cache the vectors to disk, and train the
|
| 805 |
+
50,822-parameter adapter on the cached matrix."* Router training needs **no GPU**.
|
| 806 |
+
- **F4-3 (splits by group, never by example).** *"Splitting by example would put 'show me the water body'
|
| 807 |
+
in train and 'show me the road' in val β same template, one token apart β and report a fake accuracy."*
|
| 808 |
+
The train function **refuses to proceed** if the split report is not clean.
|
| 809 |
+
|
| 810 |
+
The router's reported accuracy is **0.965116 on validation, ungated, n = 86**; the **test split was NOT
|
| 811 |
+
RUN**. This is a data-pipeline fact: the test split exists and is unused.
|
| 812 |
+
|
| 813 |
+
---
|
| 814 |
+
|
| 815 |
+
## 9. The grounding data path
|
| 816 |
+
|
| 817 |
+
### 9.1 The path
|
| 818 |
+
|
| 819 |
+
`GroundingSpecialist.execute`:
|
| 820 |
+
|
| 821 |
+
1. Validate exactly one image and a non-empty phrase; check the path exists.
|
| 822 |
+
2. `load_pil_image(asset.path)` β the deterministic display decode (Β§2.3).
|
| 823 |
+
3. `encoder.encode_image(image)` and `encoder.encode_text([query])` β RemoteCLIP ViT-B/32 at 224 px,
|
| 824 |
+
producing `(49, 512)` patch tokens and a 512-d text embedding.
|
| 825 |
+
4. Decode: `_decode_with_head` when a trained head is loaded, else `_decode_zero_shot`.
|
| 826 |
+
5. Confidence from measurable objectness statistics.
|
| 827 |
+
6. Evidence: `BOUNDING_BOX` items in `NORMALIZED_0_1`, a `GEOLOCATION` item in `GEO` when the CRS allows,
|
| 828 |
+
a `STATISTIC` item, and a `STATISTIC` record of any dropped degenerate candidates.
|
| 829 |
+
|
| 830 |
+
### 9.2 There is no feature cache on the grounding serving path
|
| 831 |
+
|
| 832 |
+
Unlike the change-VQA and optical-SAR paths, grounding has **no serving-time feature cache**. The encoder
|
| 833 |
+
runs live on each request. The Phase 7 resolution experiment and the Phase 8 evaluation used **cached**
|
| 834 |
+
patch features (the resolution experiment's artifacts are `per_sample_224_full.jsonl`,
|
| 835 |
+
`per_sample_448_full.jsonl` and `resolution_experiment_full.json` β 16,159 lines each), and the single
|
| 836 |
+
decode implementation exists precisely so a cached-feature path and a live path *"cannot disagree"*:
|
| 837 |
+
|
| 838 |
+
`decode_candidates_from_features` (`specialists/grounding/inference.py:151`) is *"THE single
|
| 839 |
+
implementation of the zero-shot decode. `ground_phrase` encodes an image then calls this; the evaluation
|
| 840 |
+
script reads a feature cache then calls this."* The module records the defect this closed: the Phase 8
|
| 841 |
+
eval originally built its own single-box baseline with `argmax_candidate` while Phase 7 measured through
|
| 842 |
+
`ground_phrase` β same 16,159 records, same metric, same cached features, but a different **decode**:
|
| 843 |
+
|
| 844 |
+
```
|
| 845 |
+
Phase 7 via ground_phrase : mean best IoU 0.0972
|
| 846 |
+
eval via argmax_candidate : mean best IoU 0.0092
|
| 847 |
+
```
|
| 848 |
+
|
| 849 |
+
A 10Γ gap, from a duplicated decode. *"The duplication was the defect, so the duplication is what is
|
| 850 |
+
removed: there is now exactly one place that turns similarity into boxes."*
|
| 851 |
+
|
| 852 |
+
### 9.3 The decode
|
| 853 |
+
|
| 854 |
+
`similarity_map` L2-normalises both sides before the dot product, because *"without that the dot product is
|
| 855 |
+
dominated by whichever vector happens to have the larger norm, which is a property of the features rather
|
| 856 |
+
than of the match."* `decode_candidates_from_features` then produces a threshold box (patches within
|
| 857 |
+
`DEFAULT_DELTA = 0.02` of the peak) plus up to `top_k` local maxima (a patch must beat its 4-neighbours,
|
| 858 |
+
and adjacent already-accepted peaks are skipped). If nothing qualifies, it falls back to `argmax_candidate`.
|
| 859 |
+
|
| 860 |
+
The head path uses `decode_cell_relative` + `sigmoid(objectness)` + `nms(iou_threshold=0.50)`, capped at
|
| 861 |
+
the specialist's `max_candidates` (6 by default, not the config's 20). The degenerate-box guard
|
| 862 |
+
(`DEGENERATE_AREA_FRACTION = 0.9`) applies to the zero-shot path only.
|
| 863 |
+
|
| 864 |
+
---
|
| 865 |
+
|
| 866 |
+
## 10. The VQA / caption data path
|
| 867 |
+
|
| 868 |
+
The VQA and caption paths share the display decode and the deterministic quality gate:
|
| 869 |
+
|
| 870 |
+
1. `load_image_array(asset.path)` β `(H, W, 3)` uint8.
|
| 871 |
+
2. The quality gate (`assess_image_quality`) β a `NOISE` or `INVALID_VALUES` verdict **blocks the VLM
|
| 872 |
+
call**; a `FLAT` or `TOO_SMALL` verdict is usable but degraded. The gate is **wired into the serving
|
| 873 |
+
path**, not merely defined: `specialists/vqa/inference.py:190` computes it and returns
|
| 874 |
+
`self._refuse(...)` when `not quality.is_usable`, *before* the model is constructed or called. The
|
| 875 |
+
comment there states the reason β *"The VLM will produce a fluent, specific, entirely fabricated scene
|
| 876 |
+
description for ANY input it is handed, including uniform noise, even when the prompt explicitly
|
| 877 |
+
instructs it to decline. Measured, not theorised. The guard therefore has to sit here, before the
|
| 878 |
+
model."*
|
| 879 |
+
3. The processor builds `pixel_values` with `processor_longest_edge: 512` pinned, so a 512-px tile is
|
| 880 |
+
**not** upscaled and split. The measured contrast is recorded in `configs/base.yaml`: default β
|
| 881 |
+
`pixel_values (1, 17, 3, 512, 512)`, 1142 prompt tokens; pinned β `pixel_values (1, 1, 3, 512, 512)`.
|
| 882 |
+
4. Prompts **must** go through `processor.apply_chat_template()` β hand-written prompt strings raise
|
| 883 |
+
`ValueError` because SmolVLM requires one `<image>` token per image (finding F5-3).
|
| 884 |
+
5. `max_images_per_call: 1` β the VLM sees one image per call.
|
| 885 |
+
|
| 886 |
+
The VLM adapter is loaded from `SATQUERY_VLM_ADAPTER` when set (the resolution order is: explicit
|
| 887 |
+
argument, then that env var, then none). The adapter's metrics are **usable** (exact_match 0.963) but its
|
| 888 |
+
status is **ACCEPTANCE-REJECTED** β USABLE β ACCEPTED.
|
| 889 |
+
|
| 890 |
+
---
|
| 891 |
+
|
| 892 |
+
## 11. Determinism and reproducibility
|
| 893 |
+
|
| 894 |
+
### 11.1 What is deterministic
|
| 895 |
+
|
| 896 |
+
| Stage | Determinism | Mechanism |
|
| 897 |
+
|---|---|---|
|
| 898 |
+
| `load_image_array` | **Deterministic** | Pure percentile stretch; no sampling |
|
| 899 |
+
| `to_grayscale_float` | **Deterministic** | Pure percentile stretch |
|
| 900 |
+
| `assess_image_quality` | **Deterministic** | *"No model, no randomness, no thresholds learned from data"* |
|
| 901 |
+
| `normalise_for_croma` | **Deterministic** | *"No sampling, no learned statistics, no RNG, and no dependence on batch composition"* |
|
| 902 |
+
| `infer_modality` | **Deterministic** | Pure function of band count and explicit label |
|
| 903 |
+
| Sensor adapter | **Deterministic** | Pure band placement + zero-fill |
|
| 904 |
+
| `assemble_fusion_input` | **Deterministic** | Concatenation with asserted order |
|
| 905 |
+
| `resize_to_canonical` | **Deterministic** | Bilinear interpolation |
|
| 906 |
+
| Registration measurement | **Deterministic** | `cv2.phaseCorrelate`; no sampling |
|
| 907 |
+
| Post-processing (morphology, components) | **Deterministic** | `cv2` fixed operations |
|
| 908 |
+
| `channel_dropout` | **Seeded** | Takes an `rng`; *"Seeded deliberately: an unseeded training transform makes a run unreproducible"* |
|
| 909 |
+
| Router training | **Seeded** | `project.seed: 42` |
|
| 910 |
+
| Fusion training | **Seeded** | Explicit seeding, per `train_change_head`'s discipline |
|
| 911 |
+
|
| 912 |
+
### 11.2 The global seed
|
| 913 |
+
|
| 914 |
+
`configs/base.yaml` declares `project.seed: 42`. `Config.seed` returns
|
| 915 |
+
`int(self.get("project.seed", 42))`. Training loops take it from the config.
|
| 916 |
+
|
| 917 |
+
### 11.3 The config hash
|
| 918 |
+
|
| 919 |
+
`Config.hash` is `sha256(json.dumps(self._data, sort_keys=True, default=str).encode())[:16]`. The frozen
|
| 920 |
+
value is **`78f1e3700da15aa1`**. It is recorded in every evaluation run, and `scripts/eval_change.py`
|
| 921 |
+
**refuses to score on drift** (exit 3). This is why several pipeline decisions live in code constants or
|
| 922 |
+
environment channels rather than in `configs/base.yaml`: adding a key moves the hash and detaches the
|
| 923 |
+
published benchmark from its config.
|
| 924 |
+
|
| 925 |
+
### 11.4 Environment channels β two distinct classes
|
| 926 |
+
|
| 927 |
+
There are two **different** kinds of environment override, and conflating them would be wrong:
|
| 928 |
+
|
| 929 |
+
**(a) Pre-hash overrides β these DO change `Config.hash`.** `load_config` applies these into the data
|
| 930 |
+
dictionary **before** `Config(data)` is constructed:
|
| 931 |
+
|
| 932 |
+
| Variable | Effect |
|
| 933 |
+
|---|---|
|
| 934 |
+
| `SATQUERY_PRECISION` | Sets `training.precision` |
|
| 935 |
+
| `SATQUERY_TORCH_COMPILE` | Sets `deployment.torch_compile` |
|
| 936 |
+
|
| 937 |
+
Because they are merged before hashing, a run using either has a **different** `Config.hash` from the
|
| 938 |
+
frozen one.
|
| 939 |
+
|
| 940 |
+
**(b) Hash-exempt overrides β these do NOT change `Config.hash`.** They are read at use time, outside the
|
| 941 |
+
config object:
|
| 942 |
+
|
| 943 |
+
| Variable | Read by | Purpose |
|
| 944 |
+
|---|---|---|
|
| 945 |
+
| `SATQUERY_DEVICE` | `core/config.py::Config.device_preference` | Device selection |
|
| 946 |
+
| `SATQUERY_CROMA_USE_8_BIT` | `specialists/optical_sar/radiometry.py::resolve_use_8_bit` | The uint8 arm (A/B) |
|
| 947 |
+
| `SATQUERY_CROMA_CHECKPOINT` | `specialists/optical_sar/croma.py::resolve_checkpoint_path` | CROMA checkpoint path (first in resolution order) |
|
| 948 |
+
| `SATQUERY_VLM_ADAPTER` | `specialists/vqa/model.py` | VLM adapter path |
|
| 949 |
+
|
| 950 |
+
The distinction matters for reproducibility: a run that sets `SATQUERY_DEVICE` is comparable to a run that
|
| 951 |
+
does not (same config identity), while a run that sets `SATQUERY_PRECISION` is **not** (different hash).
|
| 952 |
+
|
| 953 |
+
### 11.5 Checkpoint resolution order
|
| 954 |
+
|
| 955 |
+
`specialists/optical_sar/croma.py::resolve_checkpoint_path` resolves the CROMA checkpoint
|
| 956 |
+
**offline-first**: `SATQUERY_CROMA_CHECKPOINT` (when set and non-empty, and the path must exist β
|
| 957 |
+
otherwise it raises naming the variable), then the Hugging Face Hub cache. The module records that the
|
| 958 |
+
`use_8_bit` resolution order is *"`SATQUERY_CROMA_USE_8_BIT` and then `croma.use_8_bit`."*
|
| 959 |
+
|
| 960 |
+
### 11.6 What "reproducible" means per artefact
|
| 961 |
+
|
| 962 |
+
| Artefact | Reproducible? | How |
|
| 963 |
+
|---|---|---|
|
| 964 |
+
| Display array from a raster | β
byte-identical | Deterministic stretch |
|
| 965 |
+
| Availability mask | β
byte-identical | Deterministic placement |
|
| 966 |
+
| `normalise_for_croma` output | β
byte-identical | Pure function |
|
| 967 |
+
| Change feature cache | β
given the same spec hash and detector state | Spec-guarded |
|
| 968 |
+
| Text feature cache | β
given the same encoder + revision | Spec-guarded |
|
| 969 |
+
| Fusion feature cache | β
given the same arm, resolution and config | Provenance-guarded |
|
| 970 |
+
| Router corpus cache | β
given the same encoder id + text set | Fingerprint-guarded |
|
| 971 |
+
| A trained head | β not byte-reproducible across hardware | Seeded, but GPU nondeterminism applies |
|
| 972 |
+
|
| 973 |
+
`UNKNOWN β not established from the available evidence` whether bit-exact reproducibility of a trained
|
| 974 |
+
checkpoint is achieved across GPU runs; the code seeds explicitly but does not claim cross-device
|
| 975 |
+
determinism.
|
| 976 |
+
|
| 977 |
+
---
|
| 978 |
+
|
| 979 |
+
## 12. Where intermediate artefacts live
|
| 980 |
+
|
| 981 |
+
All generated artefacts live under `artifacts/`, which is *"the repository's generated-artefact root (not
|
| 982 |
+
source, not config, not `data/` inputs)."*
|
| 983 |
+
|
| 984 |
+
### 12.1 Caches (reproducible from a corpus)
|
| 985 |
+
|
| 986 |
+
| Artefact | Path | Producer | Key | Reproducible? |
|
| 987 |
+
|---|---|---|---|---|
|
| 988 |
+
| Router corpus embeddings | `<cache_dir>/embeddings_<fingerprint[:16]>.npy` + `.json` | `router/train.py::embed_corpus_cached` | corpus fingerprint (`v1` + encoder id + max_length + sorted texts) | β
|
|
| 989 |
+
| Change features | `artifacts/change_vqa/change_features.npz` | `scripts/prepare_change_vqa.py` | `feature_spec_hash(256, "stanet:trained")` = `c801326f85a185f8` | β
|
|
| 990 |
+
| Question features | `artifacts/change_vqa/text_features_<split>.npz` | `scripts/prepare_change_vqa.py` | `text_spec_hash(...)` = `d2801ea1a314354a` (pinned revision) | β
|
|
| 991 |
+
| Scene targets | `artifacts/change_vqa/scene_targets.jsonl` | `scripts/prepare_change_vqa.py` | `change_vqa_preproc_v1` | β
|
|
| 992 |
+
| Scene manifest | `artifacts/change_vqa/scene_manifest.jsonl` | `scripts/prepare_change_vqa.py` | `change_vqa_preproc_v1` | β
|
|
| 993 |
+
| Fusion features (Arm A) | `artifacts/optical_sar/fusion_features/train.npz` + `.json` | `scripts/extract_fusion_features.py` | `CACHE_VERSION="v1"` + arm + config + checkpoint digest | β
|
|
| 994 |
+
| Fusion features (Arm B) | β | β | β | **NOT PRODUCED** |
|
| 995 |
+
|
| 996 |
+
**A naming caution.** There is **no** `fusion_features_armB` directory (or any `armB`/`arm_b` literal) in
|
| 997 |
+
the repository β a grep for those spellings returns nothing. The arm is a field *inside* the cache's JSON
|
| 998 |
+
sidecar and the run record, not a directory name; the Arm-B cache would be a **separate cache built under
|
| 999 |
+
the same `artifacts/optical_sar/fusion_features/` root**, distinguished by its provenance record. Since
|
| 1000 |
+
that cache has not been built, the root holds only the Arm-A `train.npz` + `train.json`.
|
| 1001 |
+
|
| 1002 |
+
### 12.2 Trained artefacts (not reproducible byte-for-byte)
|
| 1003 |
+
|
| 1004 |
+
| Artefact | Path | Notes |
|
| 1005 |
+
|---|---|---|
|
| 1006 |
+
| Grounding head | `artifacts/grounding/remoteclip_grounding_v001/head.pt` | The shipped default; `DEFAULT_HEAD_PATH` in code |
|
| 1007 |
+
| Change detector | `artifacts/change/levir_change_v001/head.pt` | Test pooled IoU 0.8122; **not wired into serving by default** |
|
| 1008 |
+
| Change-VQA head | `artifacts/change_vqa/run/head.pt` | `change_vqa_head_v1`, 1,453,912 params |
|
| 1009 |
+
| Fusion head (production) | `artifacts/optical_sar/fusion_head_production_v001/` | `pre_registered_115_metric.json` |
|
| 1010 |
+
|
| 1011 |
+
### 12.3 Evidence artefacts
|
| 1012 |
+
|
| 1013 |
+
| Artefact | Path |
|
| 1014 |
+
|---|---|
|
| 1015 |
+
| CROMA forward pass | `artifacts/optical_sar/croma_forward.json` |
|
| 1016 |
+
| CDVQA imagery verification | `artifacts/cdvqa/imagery_verification.json` |
|
| 1017 |
+
| CDVQA/SECOND overlap | `artifacts/cdvqa/second_overlap.json` |
|
| 1018 |
+
| CDVQA temporal order | `artifacts/cdvqa/temporal_order_evidence_v2.json` |
|
| 1019 |
+
| Resolution experiment | `per_sample_224_full.jsonl`, `per_sample_448_full.jsonl`, `resolution_experiment_full.json` |
|
| 1020 |
+
| Phase 12 selection manifest | `artifacts/phase12_selection/selection_manifest_seed10.jsonl` |
|
| 1021 |
+
|
| 1022 |
+
### 12.4 What the serving path writes
|
| 1023 |
+
|
| 1024 |
+
The serving path writes only **server-side diagnostic artefacts**, and their refs are **always `null`** to
|
| 1025 |
+
the client (F-16): the change map (`change_map_<stem>.tif`) and the optical/SAR views
|
| 1026 |
+
(`optical_view_<stem>.png`, `sar_view_<stem>.png`). *"The map is still WRITTEN. It is the operator's
|
| 1027 |
+
diagnostic ... What changes is only that its location is a server-side fact, not a client-facing one."*
|
| 1028 |
+
No `artifact://` URI is fabricated in its place.
|
| 1029 |
+
|
| 1030 |
+
---
|
| 1031 |
+
|
| 1032 |
+
## 13. What is NOT part of the pipeline
|
| 1033 |
+
|
| 1034 |
+
Each item below was checked against the code. Where a capability is absent, it is named rather than
|
| 1035 |
+
implied.
|
| 1036 |
+
|
| 1037 |
+
| Capability | Status | Evidence |
|
| 1038 |
+
|---|---|---|
|
| 1039 |
+
| **Reprojection** | **NOT part of the pipeline** | `compare_crs` reports `requires_reprojection`; no code reprojects. |
|
| 1040 |
+
| **Cloud masking** | **NOT part of the pipeline** | No cloud/QA/SCL handling anywhere. |
|
| 1041 |
+
| **Atmospheric correction** | **NOT part of the pipeline** | Rasters are consumed as delivered. |
|
| 1042 |
+
| **Mosaicking** | **NOT part of the pipeline** | No composite or seamline code. |
|
| 1043 |
+
| **Tiling / top-K tile selection** | **DECLARED, not executed** | Read only by `core/config.py`'s `top_k_tiles <= max_tiles` check. |
|
| 1044 |
+
| **Percentile / dB conditioning** | **DECLARED, not executed** | Read by no code; `radiometry.py` docstring; PHASE14 Β§1. |
|
| 1045 |
+
| **Nodata filling** | **NOT part of the pipeline** | `quality.py` refuses non-finite input; the fix is *"deferred to the tiling work."* |
|
| 1046 |
+
| **Pair-dimension reconciliation** | **NOT part of the pipeline** | The change path requires equal shapes; no reconciling resize. |
|
| 1047 |
+
| **A serving-time feature cache** | **NOT part of the pipeline** | Only the router's training corpus cache and the change-VQA/fusion preparation caches exist; serving encodes live. |
|
| 1048 |
+
| **Cross-step artefact hand-off** | **NOT part of the pipeline** | The change-VQA specialist *"accepts an optional `change_map` path in `request.params` for a future planner to populate. It is unused today"* β *"the controller does not hand artifacts between steps."* |
|
| 1049 |
+
| **A single-pass change+language request** | **NOT part of the pipeline** | Routing a change+language request plans **both** a `change` and a `change_vqa` step, so *"the detector runs twice for one request"* β the weights load once (the registry caches the instance), but the forward pass runs twice. |
|
| 1050 |
+
| **Image resolution decoding** | **NOT part of the pipeline** | CDVQA's `res_x`/`res_y` are the opaque strings `".1524m"` on every row; *"The field is of unknown semantics and is stored as an opaque string. It is not used anywhere."* |
|
| 1051 |
+
|
| 1052 |
+
---
|
| 1053 |
+
|
| 1054 |
+
## 14. What is NOT RUN / OPEN / BLOCKED for this topic
|
| 1055 |
+
|
| 1056 |
+
### NOT RUN
|
| 1057 |
+
|
| 1058 |
+
- **The router test split.** Reported accuracy is **0.965116 on validation, ungated, n = 86**; the test
|
| 1059 |
+
split was **NOT RUN**.
|
| 1060 |
+
- **The optical-SAR forward pass with a trained head at serving time.** CROMA is not loaded locally and no
|
| 1061 |
+
trained head is wired; the specialist runs sensor-only.
|
| 1062 |
+
- **The CROMA normalisation experiment (option c).** Not run; the arm set is frozen at A/B and arm B (the
|
| 1063 |
+
registered one) is non-constructible.
|
| 1064 |
+
- **An end-to-end system benchmark.** No system-level accuracy is claimed.
|
| 1065 |
+
- **A change+language request end-to-end.** The planner would run the detector twice; the artefact
|
| 1066 |
+
hand-off that would make it once is not implemented.
|
| 1067 |
+
|
| 1068 |
+
### OPEN
|
| 1069 |
+
|
| 1070 |
+
- **The percentile/dB conditioning stage.** Either implement it or retire the keys.
|
| 1071 |
+
- **The arm-constructibility question.** A human decision, deliberately unresolved.
|
| 1072 |
+
- **The optical-SAR ruling.** Accuracy 0.931 with macro-F1 0.434161; **OPEN**. Never quote the accuracy
|
| 1073 |
+
without the macro-F1.
|
| 1074 |
+
- **The change-VQA metric ruling.** `OPEN` and owner-gated (two test sets: test 0.697626/0.378373 and
|
| 1075 |
+
test2 0.651469/0.372309).
|
| 1076 |
+
- **Calibration.** ECE went **0.013755 β 0.014929 β worse**. Retained only because it is in the frozen
|
| 1077 |
+
config.
|
| 1078 |
+
- **The VLM adapter.** Metrics usable (exact_match 0.963) but status **ACCEPTANCE-REJECTED**.
|
| 1079 |
+
|
| 1080 |
+
### BLOCKED
|
| 1081 |
+
|
| 1082 |
+
- **The Arm-B fusion feature cache.** `artifacts/optical_sar/fusion_features/` holds only the Arm-A cache;
|
| 1083 |
+
`val.npz` and `test.npz` were pending. Arm B cannot be trained without its own cache.
|
| 1084 |
+
- **A clean change demo.** Blocked on a same-shape pair or a resize step.
|
| 1085 |
+
|
| 1086 |
+
### DEFERRED
|
| 1087 |
+
|
| 1088 |
+
- **`codespace_name` trailing `\n`** in the `/api/health` payload (B-02). Cosmetic; the wake path is safe.
|
| 1089 |
+
- **Nodata filling from the raster profile.** Deferred to the tiling work, which is not implemented.
|
| 1090 |
+
|
| 1091 |
+
---
|
| 1092 |
+
|
| 1093 |
+
## 15. Where the evidence lives
|
| 1094 |
+
|
| 1095 |
+
**Source modules (the pipeline this chapter describes):**
|
| 1096 |
+
|
| 1097 |
+
| Path | What it defines |
|
| 1098 |
+
|---|---|
|
| 1099 |
+
| `preprocessing/raster.py` | `inspect_raster`, `read_bands`, `write_raster`, `infer_modality`, `file_sha256`, `pixel_area_m2` |
|
| 1100 |
+
| `preprocessing/imagery.py` | `load_image_array`, `load_pil_image` β the display decode |
|
| 1101 |
+
| `preprocessing/quality.py` | `assess_image_quality`, the deterministic gate |
|
| 1102 |
+
| `specialists/optical_sar/sensor_adapter.py` | Band mapping, zero-fill, availability mask |
|
| 1103 |
+
| `specialists/optical_sar/radiometry.py` | `normalise_for_croma`, `resolve_use_8_bit`, the zero-channel rule |
|
| 1104 |
+
| `specialists/optical_sar/inference.py` | `run_pipeline`, `resize_to_canonical`, `build_masks` |
|
| 1105 |
+
| `specialists/optical_sar/fusion_head.py` | `assemble_fusion_input`, `channel_dropout`, `expected_fusion_dim` |
|
| 1106 |
+
| `specialists/optical_sar/specialist.py` | `execute`, `_descriptor_for`, `_confidence_components` |
|
| 1107 |
+
| `specialists/change/specialist.py` | `_assess_pair`, `_predict`, `_write_artifacts` |
|
| 1108 |
+
| `specialists/change/postprocess.py` | `measure_registration`, `to_grayscale_float`, `postprocess_change_map` |
|
| 1109 |
+
| `specialists/change/vqa_specialist.py` | `feature_spec_mismatch`, `unavailable_reason`, `execute` |
|
| 1110 |
+
| `specialists/grounding/specialist.py` | `execute`, `_decode_with_head`, `_to_geo_boxes` |
|
| 1111 |
+
| `specialists/grounding/inference.py` | `decode_candidates_from_features`, `ground_phrase` |
|
| 1112 |
+
| `training/change_vqa/features.py` | `feature_spec_hash`, `text_spec_hash`, `ChangeFeatureExtractor`, the caches |
|
| 1113 |
+
| `training/change_vqa/dataset.py` | `DATASET_ID`, `PREPROCESSING_VERSION`, `SPLIT_MAP`, `LABEL_PALETTE` |
|
| 1114 |
+
| `training/fusion/train.py` | `CACHE_VERSION`, `DEFAULT_FEATURE_CACHE_DIR`, `feature_cache_path`, `write_feature_cache`, `ARMS` |
|
| 1115 |
+
| `training/fusion/extract.py` | The Phase 12 producer, the DEV-2 stretch, `_assert_raw_encoder` |
|
| 1116 |
+
| `router/encoder.py` | `FrozenEncoder`, the verified MiniLM contract |
|
| 1117 |
+
| `router/train.py` | `CORPUS_CACHE_VERSION`, `_corpus_fingerprint`, `embed_corpus_cached` |
|
| 1118 |
+
| `core/config.py` | `load_config`, `Config.hash`, `device_preference`, the env overrides |
|
| 1119 |
+
| `core/controller.py` | `_inspect_asset`, the nine-state machine |
|
| 1120 |
+
|
| 1121 |
+
**Scripts (the producers):**
|
| 1122 |
+
|
| 1123 |
+
- `scripts/prepare_change_vqa.py` β scene targets, manifest, change features, text features.
|
| 1124 |
+
- `scripts/extract_fusion_features.py` β the Phase 12 fusion feature cache (exit codes 0/2/3).
|
| 1125 |
+
- `scripts/train_fusion.py`, `scripts/train_change_vqa.py`, `scripts/train_grounding.py`,
|
| 1126 |
+
`scripts/train_router.py` β the training loops.
|
| 1127 |
+
- `scripts/diagnose_feature_cache.py` β cache diagnosis.
|
| 1128 |
+
|
| 1129 |
+
**Configuration:**
|
| 1130 |
+
|
| 1131 |
+
- `configs/base.yaml` β `project.seed`, `image.*`, `optical.*`, `sar.*`, `croma.*`, `fusion.*`,
|
| 1132 |
+
`grounding.*`, `change.*`, `router.*`, `vlm.*`, `training.*`.
|
| 1133 |
+
|
| 1134 |
+
**Documents:**
|
| 1135 |
+
|
| 1136 |
+
- `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` β the two ordered stages, the Gate F arm amendment, the
|
| 1137 |
+
Arm-B blocker.
|
| 1138 |
+
- `docs/PHASE12_ENTRY_GATE.md`, `docs/PHASE12_R08_TRUTHFUL_BACKFILL.md` β the Phase 12 run records and the
|
| 1139 |
+
`result_status` plumbing.
|
| 1140 |
+
- `docs/PHASE10_CDVQA_DATA_STATUS.md` β the CDVQA/SECOND layout, the temporal semantics, the four adapter
|
| 1141 |
+
defects found by execution.
|
| 1142 |
+
- `docs/PHASE7_RESOLUTION_DECISION.md` β the 224 decision (the cached-feature experiment).
|
| 1143 |
+
- `docs/API_CONTRACT.md` Β§2.5 β the asset allowlist and the upload contract.
|
| 1144 |
+
- `docs/FINAL_DELIVERY_REPORT.md` Β§4 (run ids), Β§5 (metrics), Β§6 (the 726Β²/736Β² degraded change demo).
|
| 1145 |
+
|
| 1146 |
+
**Sibling release chapters:**
|
| 1147 |
+
|
| 1148 |
+
- [`GEOSPATIAL.md`](GEOSPATIAL.md) β the raster contract, CRS, coordinate systems, sensor adapter, tiling
|
| 1149 |
+
policy, the 224 decision.
|
| 1150 |
+
- [`architecture/03-request-lifecycle.md`](architecture/03-request-lifecycle.md) β the nine controller
|
| 1151 |
+
states.
|
| 1152 |
+
- [`architecture/05-specialists.md`](architecture/05-specialists.md) β the six specialists.
|
| 1153 |
+
- [`architecture/07-configuration-freeze.md`](architecture/07-configuration-freeze.md) β the frozen hash
|
| 1154 |
+
and every config key.
|
| 1155 |
+
- [`DATASETS.md`](DATASETS.md), [`TRAINING.md`](TRAINING.md), [`REPRODUCIBILITY.md`](REPRODUCIBILITY.md).
|