thundercode commited on
Commit
cf51713
Β·
verified Β·
1 Parent(s): 64bac50

release: add docs/GEOSPATIAL.md

Browse files
Files changed (1) hide show
  1. docs/GEOSPATIAL.md +1747 -0
docs/GEOSPATIAL.md ADDED
@@ -0,0 +1,1747 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Geospatial β€” deep reference
2
+
3
+ **Status tags used on every substantive claim:** `IMPLEMENTED` Β· `VERIFIED` Β· `MEASURED` Β· `ATTEMPTED`
4
+ Β· `NOT RUN` Β· `BLOCKED` Β· `DEFERRED` Β· `REJECTED` Β· `OPEN` Β· `RESOLVED` Β· `CLOSED`.
5
+
6
+ Every statement in this document was produced by reading a file in `C:/Users/anish/satquery-ai/`. Where
7
+ a fact is not established from the available evidence, this document writes
8
+ `UNKNOWN β€” not established from the available evidence`. Where a transform is *declared* in configuration
9
+ but read by no code, that is stated explicitly rather than presented as a working feature. Where a
10
+ capability is absent, it is listed in [Β§13](#13-what-is-not-implemented-exhaustive) rather than implied.
11
+
12
+ This chapter is the reference for the **spatial contract**: how a raster becomes an analysed image, what
13
+ georeferencing survives, what coordinate system a box is in, how a sensor's bands become CROMA's canonical
14
+ channels, and which of those steps are actually implemented on the serving path. It is written to be
15
+ sufficient to reconstruct the subsystem from the document alone.
16
+
17
+ **Sibling documents:** [`DATA_PIPELINE.md`](DATA_PIPELINE.md) (the end-to-end data path, caches and
18
+ artefacts), [`architecture/03-request-lifecycle.md`](architecture/03-request-lifecycle.md) (where the
19
+ PREPROCESS and VALIDATE states sit), [`architecture/05-specialists.md`](architecture/05-specialists.md)
20
+ (the six specialists in full), [`architecture/07-configuration-freeze.md`](architecture/07-configuration-freeze.md)
21
+ (the frozen config identity and every key), [`DATASETS.md`](DATASETS.md) (the corpora and their measured
22
+ shapes).
23
+
24
+ ---
25
+
26
+ ## Table of contents
27
+
28
+ 0. [How to read this document](#0-how-to-read-this-document)
29
+ 1. [The raster input contract](#1-the-raster-input-contract)
30
+ 2. [CRS handling](#2-crs-handling-geospatialcrs-py)
31
+ 3. [Affine transforms and coordinate frames](#3-affine-transforms-and-coordinate-frames)
32
+ 4. [The coordinate-system contract](#4-the-coordinate-system-contract)
33
+ 5. [Benchmark box scale β€” 0–100 versus 0–1](#5-benchmark-box-scale--0100-versus-01)
34
+ 6. [Band-count modality inference](#6-band-count-modality-inference)
35
+ 7. [Radiometric normalisation and representation](#7-radiometric-normalisation-and-representation)
36
+ 8. [The CROMA sensor adapter contract](#8-the-croma-sensor-adapter-contract)
37
+ 9. [Pair-compatibility rules](#9-pair-compatibility-rules)
38
+ 10. [Large-image tiling policy](#10-large-image-tiling-policy)
39
+ 11. [Grounding resolution β€” the frozen 224 decision](#11-grounding-resolution--the-frozen-224-decision)
40
+ 12. [Worked examples](#12-worked-examples)
41
+ 13. [What is NOT implemented β€” exhaustive](#13-what-is-not-implemented-exhaustive)
42
+ 14. [What is NOT RUN / OPEN / BLOCKED for this topic](#14-what-is-not-run--open--blocked-for-this-topic)
43
+ 15. [Where the evidence lives](#15-where-the-evidence-lives)
44
+
45
+ ---
46
+
47
+ ## 0. How to read this document
48
+
49
+ Three distinctions are load-bearing throughout, and conflating them is the single most common way to
50
+ misread this system:
51
+
52
+ 1. **Declared versus executed.** `configs/base.yaml` is the authoritative registry, and `core/config.py`
53
+ validates and hashes it at load time. But *declaring* a policy and *executing* it are different
54
+ claims. The tiling policy, the optical percentile stretch and the SAR dB clip are all declared in
55
+ `configs/base.yaml` and are all read by no code on the serving path. This document says so at each
56
+ point rather than presenting the key as a feature.
57
+ 2. **Adapter versus encoder.** The CROMA sensor adapter (`specialists/optical_sar/sensor_adapter.py`)
58
+ performs band mapping, zero-fill and availability masking. CROMA the encoder does **not** receive the
59
+ availability mask (finding C-1). The mask is consumed by the fusion head. These are separate contracts
60
+ and Β§8.9 treats the distinction in full.
61
+ 3. **Presentation versus radiometry.** There are two different percentile stretches in this codebase:
62
+ a per-*scene* 2/98 display stretch in `preprocessing/imagery.py` (and a grayscale cousin in
63
+ `specialists/change/postprocess.py`), and a per-*channel* `mean Β± 2Β·std` encoder-input stretch in
64
+ `specialists/optical_sar/radiometry.py`. They serve different purposes and are **not**
65
+ interchangeable. Β§7.5 is explicit about why.
66
+
67
+ The status vocabulary is used on every substantive claim. `MEASURED` means a number was produced by
68
+ executing code; `VERIFIED` means the claim was checked against a second independent source; `IMPLEMENTED`
69
+ means the code exists on the serving path; `DECLARED (not read)` means the value lives in
70
+ `configs/base.yaml` and no code reads it.
71
+
72
+ ---
73
+
74
+ ## 1. The raster input contract
75
+
76
+ ### 1.1 What arrives, and what is accepted
77
+
78
+ An asset reaches the analysis path through `POST /v1/assets`, which mints an opaque `asset_id`. The
79
+ gateway enforces a **closed content-type allowlist of exactly five types**
80
+ (`docs/API_CONTRACT.md` Β§2.5):
81
+
82
+ | Content type | Notes |
83
+ |---|---|
84
+ | `image/tiff` | **The type the geospatial specialists need.** A client that uploads only PNG/JPEG can serve `vqa`, `caption` and `grounding`, but not `change` or `optical_sar` |
85
+ | `image/geotiff` | Accepted; media-type parameters are ignored, so `image/tiff; charset=binary` is accepted |
86
+ | `image/png` | Accepted |
87
+ | `image/jpeg` | Accepted |
88
+ | `application/octet-stream` | Accepted, but the raster reader still has to be able to open the bytes |
89
+
90
+ Two behaviours are specified rather than defaulted: a request with **no** declared content type is
91
+ **refused** rather than defaulted (defaulting is how a PDF reaches a raster reader), and a disallowed
92
+ type gets `415`. The stored filename is derived from the **content type**, never from the client's
93
+ filename, so a client-supplied `../../` cannot influence where bytes land (`docs/API_CONTRACT.md` Β§2.5).
94
+ The size cap is enforced at two layers β€” the gateway and the Space β€” and both now refuse an unparsable
95
+ or non-positive value rather than silently falling back to a default
96
+ (`docs/API_CONTRACT.md` Β§2.5, "Size cap").
97
+
98
+ `artifact_ref` on every `Evidence` item is **always `null` in v1** (`core/schemas.py:225`): v1 exposes no
99
+ artifact-serving endpoint, so a reference a client cannot retrieve is not fabricated. The schema field's
100
+ own description states this β€” *"Reference to an externally retrievable artifact. NEVER a filesystem path
101
+ (F-16, owner ruling 2026-09-23)"* β€” and the change and optical-SAR specialists both emit `null` on every
102
+ path (Β§4.3).
103
+
104
+ ### 1.2 The validation chain
105
+
106
+ `preprocessing/raster.py` is the first thing every uploaded asset meets. Its module docstring names the
107
+ chain it implements, verbatim:
108
+
109
+ ```
110
+ file -> dimensions -> bands -> dtype -> CRS -> transform -> bounds -> nodata
111
+ -> modality -> temporal metadata
112
+ ```
113
+
114
+ The module's three design rules are stated at the top of the file and are worth quoting because they
115
+ govern every failure mode below:
116
+
117
+ > * Never raise a bare exception. Every failure is a typed SatQueryError.
118
+ > * Never silently drop geospatial metadata. If the source had a CRS, the returned AssetMetadata says so.
119
+ > * A missing CRS degrades to non-geospatial mode; it does not abort.
120
+
121
+ ### 1.3 `inspect_raster` β€” header-only validation
122
+
123
+ `inspect_raster` (`preprocessing/raster.py:96`) validates and describes a raster **without loading its
124
+ pixel data**. Its exact signature:
125
+
126
+ ```python
127
+ def inspect_raster(
128
+ path: str | Path,
129
+ *,
130
+ max_pixels: int | None = None,
131
+ explicit_modality: str | None = None,
132
+ compute_hash: bool = False,
133
+ sensor: SensorDescriptor | None = None,
134
+ ) -> AssetMetadata:
135
+ ```
136
+
137
+ The behaviour, field by field:
138
+
139
+ - The path is checked with `Path(path).exists()` and `is_file()`; a missing path raises `RasterReadError`.
140
+ - `rasterio` is imported lazily; an `ImportError` raises `RasterReadError` rather than a bare error.
141
+ - Inside `rasterio.open(p)`, the reader extracts `width`, `height`, `count`, the sorted set of `dtypes`,
142
+ the CRS as a string (`ds.crs.to_string() if ds.crs else None`), the affine transform truncated to its
143
+ first six coefficients (`list(ds.transform)[:6] if ds.transform else None`), the bounds
144
+ (`list(ds.bounds)`), `nodata`, `resolution` (`list(ds.res)`), the `driver` string, and a tiled flag
145
+ read through `_safe_is_tiled`.
146
+ - `_safe_is_tiled` (`preprocessing/raster.py:42`) exists because `DatasetReader.is_tiled` is scheduled for
147
+ removal; it suppresses the pending deprecation and **degrades to `False`** rather than aborting, since
148
+ the flag is informational metadata only.
149
+ - Degenerate dimensions (`width <= 0 or height <= 0`) raise `RasterReadError`.
150
+ - If `max_pixels` is given and `width * height > max_pixels`, it raises `OversizedImageError` β€” and the
151
+ docstring records that this is **recoverable**: *"the caller may downscale"*.
152
+ - A band count `<= 0` raises `UnsupportedBandsError`.
153
+ - Anything else rasterio raises is wrapped: `RasterReadError(f"could not open raster '{p}': {exc}")`.
154
+
155
+ The return is an `AssetMetadata` (see Β§1.4) with `modality` from `infer_modality` (Β§6) and, when
156
+ `compute_hash=True`, the streaming SHA-256 of the file (`file_sha256`, `preprocessing/raster.py:82`,
157
+ chunked at 1 MiB).
158
+
159
+ `inspect_raster` is called from `core/controller.py::_inspect_asset` during the VALIDATE state, with the
160
+ configured pixel budget passed through. That means the **header-only** nature of the inspection is a real
161
+ property of the serving path: a request never decodes a raster just to learn its shape.
162
+
163
+ ### 1.4 `AssetMetadata` and `GeoMetadata`
164
+
165
+ `GeoMetadata` (`core/schemas.py:120`) is the geospatial detail carrier. It is `extra="allow"`, which is
166
+ deliberate β€” the optical-SAR specialist adds cross-sensor keys to it (Β§4.3):
167
+
168
+ | Field | Type | Meaning |
169
+ |---|---|---|
170
+ | `crs` | `str \| None` | CRS as a string (e.g. `EPSG:4326`), or `None` |
171
+ | `transform` | `list[float] \| None` | Affine coefficients, row-major, six values |
172
+ | `bounds` | `list[float] \| None` | `[minx, miny, maxx, maxy]` |
173
+ | `width`, `height` | `int \| None` | Raster dimensions in pixels |
174
+ | `band_count` | `int \| None` | Number of bands |
175
+ | `dtype` | `str \| None` | Comma-joined distinct dtypes when the bands disagree |
176
+ | `nodata` | `float \| None` | The nodata sentinel, or `None` |
177
+ | `resolution` | `list[float] \| None` | `(xres, yres)` |
178
+ | `has_crs` | `bool` | `crs_value is not None` |
179
+ | `is_georeferenced` | `bool` | `bool(crs_value and transform)` |
180
+
181
+ `AssetMetadata` (`core/schemas.py:153`) is `extra="forbid"` and carries:
182
+
183
+ | Field | Type | Meaning |
184
+ |---|---|---|
185
+ | `asset_id` | `str` | Minted `asset_<12 hex>`, opaque |
186
+ | `path` | `str` | Resolved local path |
187
+ | `modality` | `Modality` | From `infer_modality`, or declared |
188
+ | `sha256` | `str \| None` | Present only when `compute_hash=True` |
189
+ | `geo` | `GeoMetadata` | The block above |
190
+ | `sensor` | `SensorDescriptor \| None` | Present only when a sensor descriptor was supplied |
191
+ | `acquisition_date` | `str \| None` | Temporal metadata |
192
+ | `scene_id`, `dataset_id`, `split` | `str \| None` | Provenance for leakage control |
193
+
194
+ Note the distinction between `has_crs` and `is_georeferenced`: a raster can carry a CRS but no affine
195
+ transform (or vice versa), and both are recorded. Only when **both** are present is
196
+ `is_georeferenced=True`. `geospatial/transform.py::to_geo` requires the transform, so a raster with
197
+ `has_crs=True` but `transform=None` cannot be converted to geographic coordinates β€” it will raise
198
+ `CoordinateError("cannot convert to geo: raster has no affine transform")`.
199
+
200
+ ### 1.5 Band reading, writing, and pixel area
201
+
202
+ `read_bands` (`preprocessing/raster.py:188`) returns `(bands, H, W)` plus a profile that **retains**
203
+ `crs`, `transform` and `bounds` as plain values, so downstream code never loses georeferencing. The
204
+ profile is a copy of `ds.profile` with those three keys overwritten from the live dataset.
205
+
206
+ `write_raster` (`preprocessing/raster.py:221`) writes an array using a profile captured by `read_bands`,
207
+ and re-applies `crs` and `transform` explicitly in the `rasterio.open(..., "w", ...)` call. It is used by
208
+ evidence generation so change maps stay georeferenced β€” **except** when the pair is suppressed for poor
209
+ registration, in which case the transform is deliberately dropped (Β§9.1).
210
+
211
+ `pixel_area_m2` (`preprocessing/raster.py:259`) returns the area of one pixel in square metres **when the
212
+ CRS is metric**, and `None` otherwise:
213
+
214
+ ```python
215
+ def pixel_area_m2(meta: AssetMetadata) -> float | None:
216
+ from geospatial.crs import is_metric
217
+ if not is_metric(meta.geo.crs):
218
+ return None
219
+ if not meta.geo.resolution or len(meta.geo.resolution) < 2:
220
+ return None
221
+ xres, yres = meta.geo.resolution[0], meta.geo.resolution[1]
222
+ return abs(xres * yres)
223
+ ```
224
+
225
+ The docstring is explicit: *"Returns None for geographic CRS or unknown resolution β€” callers must not
226
+ guess."* A degree-based CRS has no meaningful pixel area in metres, so the function refuses rather than
227
+ returning a plausible-looking number.
228
+
229
+ ---
230
+
231
+ ## 2. CRS handling (`geospatial/crs.py`)
232
+
233
+ CRS comparison is kept **separate** from pixel arithmetic so that *"can these two rasters be compared at
234
+ all?"* is answerable without touching pixels (`geospatial/crs.py` docstring). The module depends on
235
+ `pyproj`.
236
+
237
+ ### 2.1 `parse_crs`
238
+
239
+ ```python
240
+ def parse_crs(value: str | None) -> CRS | None:
241
+ if value is None or value == "":
242
+ return None
243
+ try:
244
+ return CRS.from_user_input(value)
245
+ except PyProjCRSError as exc:
246
+ raise MissingCRSError(f"unparseable CRS '{value}': {exc}") from exc
247
+ ```
248
+
249
+ Absent is `None`; malformed raises. That distinction matters: a missing CRS is a legitimate state
250
+ (degrades to non-geospatial), while a *malformed* one is a defect that must surface.
251
+
252
+ ### 2.2 `describe_crs`
253
+
254
+ `describe_crs` returns a dict for the execution trace: `present`, `srs`, `name`, `is_geographic`,
255
+ `is_projected`, and `axis_units` (the first axis's `unit_name`). For an absent CRS it returns
256
+ `{"present": False}`.
257
+
258
+ ### 2.3 `compare_crs` and `CRSCompatibility`
259
+
260
+ `CRSCompatibility` (`geospatial/crs.py:19`) is a frozen dataclass with `compatible`, `identical`,
261
+ `same_units`, `left`, `right`, `reason`, and a derived property:
262
+
263
+ ```python
264
+ @property
265
+ def requires_reprojection(self) -> bool:
266
+ return self.compatible and not self.identical
267
+ ```
268
+
269
+ `compare_crs` (`geospatial/crs.py:59`) decides:
270
+
271
+ - **A missing CRS on either side is NOT compatible.** The reason string is
272
+ `"one or both rasters lack a CRS; spatial comparison is unsafe"`. The docstring records that this is
273
+ *"not fatal for non-spatial tasks"*.
274
+ - **Identical CRS** (`crs_l.equals(crs_r)`) returns `identical=True`, reason `"identical CRS"`.
275
+ - **Same kind** (both projected or both geographic) returns `compatible=True, identical=False`, reason
276
+ `"reprojection required (<left name> -> <right name>)"`.
277
+ - **Mixed projected/geographic** returns `compatible=True, identical=False`, with a reason that adds
278
+ *"reprojection required and should be verified"*. The comment records the rationale: *"Mixing a
279
+ projected CRS with a geographic one is legal to reproject but signals a pipeline mistake."*
280
+
281
+ `same_units` compares the first axis's `unit_name` on each side (defaulting to `"unknown"` when there is
282
+ no axis info).
283
+
284
+ **Crucially: `compare_crs` never reprojects.** It reports whether reprojection is *required*; the actual
285
+ transform is not implemented anywhere (Β§13).
286
+
287
+ ### 2.4 `is_metric`
288
+
289
+ ```python
290
+ def is_metric(crs_value: str | None) -> bool:
291
+ crs = parse_crs(crs_value)
292
+ if crs is None:
293
+ return False
294
+ if not crs.is_projected:
295
+ return False
296
+ if not crs.axis_info:
297
+ return False
298
+ return crs.axis_info[0].unit_name in {"metre", "meter", "m"}
299
+ ```
300
+
301
+ A geographic CRS is never metric here, and the unit check is a closed set of three spellings.
302
+
303
+ ### 2.5 Where CRS comparison is used
304
+
305
+ The only consumer on the serving path is the change specialist's `_assess_pair`
306
+ (`specialists/change/specialist.py:251`):
307
+
308
+ ```python
309
+ compat = compare_crs(t1.geo.crs, t2.geo.crs)
310
+ if not compat.compatible:
311
+ warnings.append(
312
+ f"CRS comparison unavailable ({compat.reason}); change is "
313
+ f"reported in normalized pixel coordinates only"
314
+ )
315
+ elif compat.requires_reprojection:
316
+ warnings.append(
317
+ f"T1 and T2 use different CRS ({compat.reason}); change is "
318
+ f"valid only if the rasters already share a pixel grid"
319
+ )
320
+ ```
321
+
322
+ The comment above it states the policy precisely: *"No CRS on either side is NOT fatal for pixel-domain
323
+ change detection β€” the two rasters still tile to the same grid. It IS fatal for any claim about ground
324
+ coordinates, so the warning is recorded and the geospatial block is withheld downstream."* The
325
+ withholding happens in `ChangeSpecialist._geospatial`, which drops the transform when registration is
326
+ unusable (Β§9.1).
327
+
328
+ ---
329
+
330
+ ## 3. Affine transforms and coordinate frames
331
+
332
+ `geospatial/transform.py` is where a normalised 0–1 box, a pixel box and a geographic box are converted
333
+ between. Its module docstring states the non-negotiable rules (citing `docs/ARCHITECTURE_FREEZE.md`
334
+ Β§2.6):
335
+
336
+ > * CRS and affine transform are preserved, never silently stripped.
337
+ > * Conversions are explicit about their source and target frames.
338
+ > * A conversion that cannot be performed raises, rather than returning a plausible lie.
339
+
340
+ That last line is the module's governing principle, and it is why `to_normalized`, `to_pixel` and
341
+ `to_geo` all raise `CoordinateError` when they lack the context a conversion needs.
342
+
343
+ ### 3.1 `PixelWindow` and `GeoBounds`
344
+
345
+ `PixelWindow` (`geospatial/transform.py:29`) is a frozen dataclass of `[col_min, row_min, col_max,
346
+ row_max]`. Its `__post_init__` **raises** on an inverted window:
347
+
348
+ ```python
349
+ if self.col_max < self.col_min or self.row_max < self.row_min:
350
+ raise CoordinateError(f"pixel window is inverted: {self.as_list()}")
351
+ ```
352
+
353
+ It exposes `width`, `height` and `area` properties. `GeoBounds` (`geospatial/transform.py:60`) is the
354
+ geographic cousin: `(minx, miny, maxx, maxy)`.
355
+
356
+ ### 3.2 Normalised ↔ pixel
357
+
358
+ `normalized_to_pixel(box, width, height)` multiplies each coordinate by the corresponding dimension. Its
359
+ docstring records a deliberate choice:
360
+
361
+ > Values are NOT clipped: a box slightly outside the frame is preserved so the caller can decide whether
362
+ > that is a bug or a legitimate edge case.
363
+
364
+ It raises `CoordinateError` if `len(box) != 4` or if either dimension is `<= 0`.
365
+ `pixel_to_normalized` is the inverse, dividing by the dimensions.
366
+
367
+ ### 3.3 Pixel ↔ geographic
368
+
369
+ `pixel_to_geo(window, transform)` applies the affine transform to two corners. The docstring explains the
370
+ axis convention:
371
+
372
+ > rasterio/affine transforms map (col, row) -> (x, y). The y axis usually points down in pixel space, so
373
+ > the geographic miny comes from the BOTTOM row.
374
+
375
+ It uses the non-deprecated matmul form `transform @ (col, row)` and then takes `min`/`max` across the two
376
+ corners so the returned `GeoBounds` is ordered. It raises `CoordinateError` if the transform is `None`.
377
+
378
+ `geo_to_pixel(bounds, transform)` inverts the transform with `~transform`, raising `CoordinateError` if
379
+ the inverse cannot be computed (`"affine transform is not invertible: ..."`).
380
+
381
+ ### 3.4 Schema-aware conversions
382
+
383
+ `to_normalized`, `to_pixel` and `to_geo` operate on `Box` objects and are schema-aware β€” they read
384
+ `box.coordinate_system` and return a `Box` with the target system set. The context each requires is
385
+ documented in the docstring:
386
+
387
+ - `pixel -> normalized` requires `width` and `height`.
388
+ - `geo -> normalized` requires `width`, `height` **and** the affine `transform`.
389
+
390
+ Each returns early when the box is already in the target frame, and otherwise uses
391
+ `box.model_copy(update={...})` to produce a new box with the new coordinates and the new
392
+ `coordinate_system`.
393
+
394
+ ### 3.5 Intersection and IoU
395
+
396
+ `intersect(a, b)` returns the axis-aligned intersection of two `PixelWindow`s, or `None` when they are
397
+ disjoint. `iou(a, b)` computes intersection-over-union in `[0, 1]`, returning `0.0` for a disjoint pair
398
+ or a zero-area union. These are used by the grounding NMS path and by change-region overlap.
399
+
400
+ ### 3.6 `affine_from_metadata`
401
+
402
+ ```python
403
+ def affine_from_metadata(geo: GeoMetadata) -> Affine | None:
404
+ if not geo.transform or len(geo.transform) != 6:
405
+ return None
406
+ a, b, c, d, e, f = geo.transform
407
+ return Affine(a, b, c, d, e, f)
408
+ ```
409
+
410
+ A `GeoMetadata` whose transform is missing or not exactly six values yields `None`; callers treat that as
411
+ "cannot georeference" rather than as an error. This is the bridge the grounding specialist uses in
412
+ `_to_geo_boxes` (Β§4.3).
413
+
414
+ ---
415
+
416
+ ## 4. The coordinate-system contract
417
+
418
+ ### 4.1 The enum
419
+
420
+ `CoordinateSystem` (`core/schemas.py:57`) has exactly three members, and the class docstring is a warning
421
+ in itself β€” *"Never omit this. A bare box is meaningless without it. (C-5)"*:
422
+
423
+ ```python
424
+ class CoordinateSystem(str, Enum):
425
+ NORMALIZED_0_1 = "normalized_0_1"
426
+ PIXEL = "pixel"
427
+ GEO = "geo"
428
+ ```
429
+
430
+ `docs/API_CONTRACT.md` Β§3.2 fixes the meaning of each value:
431
+
432
+ | Value | Meaning | Rendering |
433
+ |---|---|---|
434
+ | `normalized_0_1` | `[0, 1]`, origin **top-left** | Multiply by image width/height |
435
+ | `pixel` | Absolute pixel coordinates | Use directly |
436
+ | `geo` | CRS coordinates (usually EPSG:4326) | Requires a map, not a 2-D canvas |
437
+
438
+ The contract adds a hard warning: *"The frontend must read the `coordinate_system` field on each
439
+ `Box`/`Region` and must not assume one convention. A box drawn with the wrong assumption lands in
440
+ plausible-looking wrong places."* It also records that earlier drafts used shorthand (`normalized`,
441
+ `geographic`) and that those strings are **wrong** β€” they would fail schema validation, because the
442
+ models are `extra="forbid"` and the enum is closed.
443
+
444
+ ### 4.2 Why it is mandatory β€” the schema validator
445
+
446
+ The `Box` model (`core/schemas.py:171`) defaults `coordinate_system` to
447
+ `CoordinateSystem.NORMALIZED_0_1`, and its `_ordered` validator rejects a box whose coordinates do not
448
+ satisfy `x2 >= x1` and `y2 >= y1`. `Region` (`core/schemas.py:189`) and `ChangeRegion`
449
+ (`core/schemas.py:202`) both carry the same field with the same default.
450
+
451
+ The enforcement that matters is on `Evidence`. `Evidence._spatial_needs_crs`
452
+ (`core/schemas.py:238`) **rejects a spatial evidence item that carries coordinates but no
453
+ coordinate system**:
454
+
455
+ ```python
456
+ @model_validator(mode="after")
457
+ def _spatial_needs_crs(self) -> "Evidence":
458
+ spatial = {
459
+ EvidenceType.BOUNDING_BOX,
460
+ EvidenceType.MASK,
461
+ EvidenceType.CHANGE_MAP,
462
+ EvidenceType.TILE,
463
+ EvidenceType.IMAGE_CROP,
464
+ EvidenceType.JOINT_FEATURE_REGION,
465
+ }
466
+ if self.type in spatial and self.coordinates and self.coordinate_system is None:
467
+ raise ValueError(
468
+ f"evidence type '{self.type.value}' carries coordinates "
469
+ "but no coordinate_system"
470
+ )
471
+ return self
472
+ ```
473
+
474
+ So a bare box cannot be published as evidence. The validator fires only when the item both is of a
475
+ spatial type **and** carries coordinates β€” a `STATISTIC` or `GEOLOCATION` item with no coordinates is
476
+ unaffected, and a spatial item with no coordinates (e.g. an availability-mask record) is also unaffected.
477
+ `GEOLOCATION` is deliberately **not** in the spatial set: a geolocation is a point reference, not a
478
+ rectangle, and the grounding specialist sets its `coordinate_system` to `GEO` explicitly anyway.
479
+
480
+ Note the interaction with `extra="forbid"`: `Evidence` forbids extra keys, so an author cannot smuggle a
481
+ convention through an undeclared field. `GeoMetadata`, by contrast, is `extra="allow"`, which is how the
482
+ optical-SAR specialist attaches cross-sensor keys (Β§4.3).
483
+
484
+ ### 4.3 What each specialist emits
485
+
486
+ | Specialist | Boxes / regions | Evidence coordinate systems |
487
+ |---|---|---|
488
+ | `grounding` | `Box` and `Region` in `NORMALIZED_0_1` | `BOUNDING_BOX` in `NORMALIZED_0_1`; a `GEOLOCATION` item in `GEO` when the CRS allows; `STATISTIC` items with no coordinates |
489
+ | `change` | `ChangeRegion` and `Region` in `NORMALIZED_0_1` | `CHANGE_MAP` and `BOUNDING_BOX` in `NORMALIZED_0_1`; `STATISTIC` items with no coordinates |
490
+ | `optical_sar` | No `regions` (empty list) | `AVAILABILITY_MASK` items (no coordinates); `OPTICAL_VIEW` / `SAR_VIEW` with `coordinates=[0,0,1,1]` in `NORMALIZED_0_1`; a `JOINT_FEATURE_REGION` in `NORMALIZED_0_1`; `STATISTIC` items |
491
+
492
+ **Grounding.** `GroundingSpecialist.execute` (`specialists/grounding/specialist.py:253`) builds every box
493
+ with an explicit system:
494
+
495
+ ```python
496
+ boxes = [
497
+ Box(
498
+ x1=d.box[0], y1=d.box[1], x2=d.box[2], y2=d.box[3],
499
+ score=d.score, label=request.query,
500
+ coordinate_system=CoordinateSystem.NORMALIZED_0_1,
501
+ )
502
+ for d in decoded
503
+ ]
504
+ ```
505
+
506
+ Geographic boxes are **derived, never replacements**: `_to_geo_boxes` (`specialists/grounding/specialist.py:431`)
507
+ converts the normalised boxes to geographic bounds *only* when the raster has a CRS, a transform and
508
+ dimensions, and emits a `GEOLOCATION` evidence item with `coordinate_system=CoordinateSystem.GEO`. When
509
+ the raster has no CRS/transform, the method appends a warning β€”
510
+ *"raster has no CRS/transform; boxes are reported in normalized coordinates only and cannot be placed on a
511
+ map"* β€” and returns an empty list. The result's boxes stay normalised; the geo boxes are additional
512
+ evidence for a consumer with a CRS. Its own comment states the rule: *"Pixel and geo coordinates are
513
+ DERIVED, never replacements."*
514
+
515
+ **Change.** `regions_to_schema` (`specialists/change/postprocess.py:443`) builds `ChangeRegion` objects in
516
+ the requested system. It supports `NORMALIZED_0_1` and `PIXEL` and **raises** for anything else:
517
+
518
+ ```python
519
+ else:
520
+ raise ValueError(
521
+ f"regions_to_schema cannot produce {coordinate_system.value} "
522
+ f"coordinates; convert after this call"
523
+ )
524
+ ```
525
+
526
+ The change specialist calls it with the default (`NORMALIZED_0_1`). Geographic conversion is a deliberate
527
+ post-step β€” the docstring records: *"Callers that want geographic coordinates convert afterwards via
528
+ `geospatial.transform`, which keeps this function free of rasterio and CRS concerns."*
529
+
530
+ **Optical-SAR.** The specialist emits no `regions`. Its view evidence items carry
531
+ `coordinates=[0.0, 0.0, 1.0, 1.0]` in `NORMALIZED_0_1` β€” the whole frame β€” because the view is a
532
+ presentation of the canonical tensor, not a localisation. Its `geospatial` block is the **optical**
533
+ asset's georeferencing (chosen because it is the asset a reader can navigate by), with cross-sensor keys
534
+ added via `GeoMetadata`'s `extra="allow"`: `optical_crs`, `sar_crs`, `sar_bounds`, `sar_resolution`,
535
+ `optical_availability`, `sar_availability`, `optical_sensor`, `sar_sensor`, and
536
+ `resolution_ratio_sar_to_optical` when both resolutions are known
537
+ (`specialists/optical_sar/specialist.py:597`). The comment records the reason the two footprints are
538
+ carried separately rather than fused: *"RISAT and Cartosat-2S are different missions with different
539
+ orbits."*
540
+
541
+ ---
542
+
543
+ ## 5. Benchmark box scale β€” 0–100 versus 0–1
544
+
545
+ This is a small contract with an outsized failure mode, and the codebase treats it as such.
546
+
547
+ ### 5.1 The two conventions
548
+
549
+ - **VRSBench** (the grounding benchmark) states verbatim that *"all box coordinates are normalized to
550
+ 0-100"*.
551
+ - **This project** stores every box normalised to **0–1**.
552
+
553
+ Treating a 0–100 box as a pixel box, or as a 0–1 box, is a silent, catastrophic bug: the numbers are all
554
+ in range, the schema accepts them, and the box lands in the wrong place.
555
+
556
+ ### 5.2 The conversion functions
557
+
558
+ `geospatial/transform.py:286` provides the conversion, and its docstring names the failure mode it exists
559
+ to prevent:
560
+
561
+ ```python
562
+ def benchmark_boxes_to_normalized(
563
+ boxes: Iterable[Sequence[float]], scale: float = 100.0
564
+ ) -> list[list[float]]:
565
+ """Convert benchmark boxes (e.g. VRSBench's 0-100) to our internal 0-1.
566
+
567
+ Finding from Phase 0: VRSBench states verbatim that "all box coordinates are
568
+ normalized to 0-100". Treating those numbers as pixels is a silent, catastrophic
569
+ bug β€” this function exists so that mistake can be made once, explicitly.
570
+ """
571
+ if scale <= 0:
572
+ raise CoordinateError(f"box scale must be positive, got {scale}")
573
+ ...
574
+ out.append([float(v) / scale for v in b])
575
+ ```
576
+
577
+ `normalized_boxes_to_benchmark` is the exact inverse (multiply by `scale`). Both raise on a
578
+ non-positive scale and on a box that is not four values.
579
+
580
+ A second, independent conversion lives in the evaluation metrics:
581
+ `evaluation/metrics/grounding.py::benchmark_to_normalized(box, scale=100.0)` divides by the scale and
582
+ raises `ValueError(f"scale must be positive, got {scale}")` on a non-positive value. `docs/API_CONTRACT.md`
583
+ Β§3.2 names it explicitly as the site of the VRSBench conversion.
584
+
585
+ ### 5.3 The declared scale
586
+
587
+ `configs/base.yaml` declares the conversion factor in the grounding block, with a comment that restates
588
+ the convention:
589
+
590
+ ```yaml
591
+ grounding:
592
+ ...
593
+ # VRSBench stores boxes normalised to 0-100; we store 0-1.
594
+ benchmark_box_scale: 100.0
595
+ coordinate_system: normalized_0_1
596
+ ```
597
+
598
+ `benchmark_box_scale: 100.0` is therefore the **declared** value, and `100.0` is also the **default
599
+ argument** in both conversion functions. The two agree.
600
+
601
+ ### 5.4 The double-application risk
602
+
603
+ Because the scale appears in three places β€” the config key, the `geospatial.transform` default and the
604
+ `evaluation.metrics.grounding` default β€” there is a real hazard that a caller converts 0–100 β†’ 0–1 and
605
+ then a downstream stage divides by 100 again, producing a box 100Γ— too small in each coordinate. The
606
+ codebase's mitigations are structural rather than documentary:
607
+
608
+ 1. **The conversion is named for its direction.** `benchmark_boxes_to_normalized` and
609
+ `normalized_boxes_to_benchmark` are inverses with self-describing names, so a reader cannot easily
610
+ apply the wrong one without noticing.
611
+ 2. **The default is the correct scale, not `1.0`.** A caller that forgets to pass `scale` gets `100.0`,
612
+ which is right for VRSBench and would produce an obviously wrong box (values > 1) for already-0–1
613
+ input β€” a loud failure, not a silent one.
614
+ 3. **The schema rejects out-of-range values indirectly.** A `Box` built from un-divided VRSBench
615
+ coordinates would have `score` in range but coordinates in `[0, 100]`; the `_ordered` validator would
616
+ pass it (the ordering is still valid), so the schema alone does **not** catch this. The catch is that
617
+ the metrics module's own `benchmark_to_normalized` is the documented entry point, and the resolution
618
+ experiment and the grounding evaluation both go through it.
619
+
620
+ **Residual risk β€” stated honestly.** There is no runtime assertion that a `Box` in
621
+ `NORMALIZED_0_1` actually has coordinates within `[0, 1]`. `Box._ordered` checks only ordering. A
622
+ mis-scaled box would therefore pass schema validation. `UNKNOWN β€” not established from the available
623
+ evidence` whether any production call site currently constructs a `Box` without going through one of the
624
+ two conversion functions; the serving path (`GroundingSpecialist.execute`) builds boxes from the decoder's
625
+ output, which is already in `[0, 1]`, so the serving path is not exposed. The risk is confined to
626
+ evaluation and experiment code that reads VRSBench directly.
627
+
628
+ ---
629
+
630
+ ## 6. Band-count modality inference
631
+
632
+ ### 6.1 The heuristics
633
+
634
+ `preprocessing/raster.py` defines two closed sets:
635
+
636
+ ```python
637
+ _OPTICAL_BAND_COUNTS = {3, 4, 8, 11, 12, 13}
638
+ _SAR_BAND_COUNTS = {1, 2}
639
+ ```
640
+
641
+ with the comment: *"Band-count heuristics for modality detection when explicit metadata is absent. These
642
+ are heuristics, not ground truth β€” the sensor adapter is authoritative when a sensor descriptor is
643
+ supplied."*
644
+
645
+ ### 6.2 `infer_modality`
646
+
647
+ ```python
648
+ def infer_modality(band_count: int, explicit: str | None = None) -> Modality:
649
+ if explicit:
650
+ try:
651
+ return Modality(explicit.lower())
652
+ except ValueError:
653
+ pass
654
+
655
+ if band_count in _SAR_BAND_COUNTS:
656
+ return Modality.SAR
657
+ if band_count in _OPTICAL_BAND_COUNTS:
658
+ return Modality.OPTICAL
659
+ # 12-band optical and 2-band SAR are both plausible; default to optical only
660
+ # when the count clearly favours it.
661
+ if band_count >= 4:
662
+ return Modality.OPTICAL
663
+ return Modality.UNKNOWN
664
+ ```
665
+
666
+ The decision order is worth reading carefully:
667
+
668
+ 1. An **explicit** label wins if it parses to a `Modality` member. An explicit label that does *not* parse
669
+ (e.g. `"radar"`) is silently ignored and inference proceeds β€” the `except ValueError: pass` is
670
+ deliberate.
671
+ 2. `{1, 2}` β‡’ `SAR`.
672
+ 3. `{3, 4, 8, 11, 12, 13}` β‡’ `OPTICAL`.
673
+ 4. `>= 4` (and not already matched) β‡’ `OPTICAL`.
674
+ 5. Otherwise β‡’ `UNKNOWN`.
675
+
676
+ So the effective mapping is: 1–2 β‡’ SAR; 3–13 β‡’ optical (by set or by the `>= 4` fallback); 0 or a
677
+ non-positive count β‡’ `UNKNOWN` (though `inspect_raster` rejects `band_count <= 0` before this point).
678
+
679
+ ### 6.3 The declared-wins rule
680
+
681
+ In the optical-SAR specialist, `_assess_pair` (`specialists/optical_sar/specialist.py:223`) uses the
682
+ asset's **declared** modality first, falling back to band-count inference only when it is `UNKNOWN`:
683
+
684
+ ```python
685
+ for asset in assets:
686
+ modality = asset.modality
687
+ if modality is Modality.UNKNOWN:
688
+ inferred = self._infer_modality(asset, warnings)
689
+ resolved.append(inferred)
690
+ else:
691
+ resolved.append(modality)
692
+ ```
693
+
694
+ `_infer_modality` appends a warning that names the heuristic for what it is: *"A band count is a
695
+ heuristic, not a sensor declaration β€” label the asset explicitly if this is wrong."* The docstring above
696
+ `_assess_pair` states the rationale: *"The declared value wins when present, because the caller may know
697
+ something the band count does not β€” a 2-band Cartosat stack, or a 12-band decomposition product."*
698
+
699
+ ### 6.4 Why the heuristic is not authoritative
700
+
701
+ The comment on the sets and the warning text both point at the same thing: a band count does not identify
702
+ a *sensor*. A 2-band file could be a dual-pol SAR pair or a two-band optical stack; a 4-band file could be
703
+ RGB+NIR or a four-polarisation SAR product. The adapter (Β§8) is the authoritative path, because it maps
704
+ named bands onto CROMA's canonical channels and records what it could not place.
705
+
706
+ ---
707
+
708
+ ## 7. Radiometric normalisation and representation
709
+
710
+ This section covers a genuine specification gap and its partial resolution. It is presented in full
711
+ because the gap is load-bearing and the honest state is more useful than a tidy summary.
712
+
713
+ ### 7.1 What configuration declares
714
+
715
+ `configs/base.yaml` declares two radiometric transforms:
716
+
717
+ ```yaml
718
+ optical:
719
+ normalization: percentile
720
+ lower_percentile: 2
721
+ upper_percentile: 98
722
+ # CROMA expects exactly 12 optical channels.
723
+ canonical_channels: 12
724
+
725
+ sar:
726
+ representation: db
727
+ clip_min_db: -30
728
+ clip_max_db: 5
729
+ # CROMA expects exactly 2 SAR channels (VV, VH).
730
+ canonical_channels: 2
731
+ ```
732
+
733
+ So the declared optical normalisation is a **percentile stretch at the 2nd and 98th percentiles**, and the
734
+ declared SAR representation is **decibels clipped to [βˆ’30, +5]**.
735
+
736
+ ### 7.2 The DEV-2 ruling β€” two ordered stages
737
+
738
+ `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` records the ruling that governs this. The gap it confirmed:
739
+ CROMA's own example preprocessing is a **per-channel** dynamic-range stretch
740
+ (`mean Β± 2Β·std β†’ [0, 1]`, optionally through a uint8 round-trip), while `configs/base.yaml` specifies
741
+ something structurally different. The ruling's verification table records that a grep of
742
+ `specialists/optical_sar/` for `NORM_*`, `clip_min_db`, `lower_percentile`, `use_8_bit` returned **six
743
+ hits, all constant definitions/labels** at `sensor_adapter.py:125,126,210,323,547,548`, and **zero
744
+ application sites**.
745
+
746
+ The ruling resolves this by declaring the pipeline to have **two ordered stages**, and it states
747
+ explicitly that they are *not alternatives*:
748
+
749
+ | Stage | Transform | Owner | Status |
750
+ |---|---|---|---|
751
+ | **Radiometric conditioning** | per-scene percentile / dB-clip, per `configs/base.yaml` | preprocessing / sensor adapter | **specified in the ruling, must be implemented** |
752
+ | **Encoder-input normalisation** | per-channel `mean Β± 2Β·std` β†’ `[0,1]`, then optional `Γ—255` uint8 β†’ `/255` | a new function consumed by `croma.py` | **mandatory, immediate** |
753
+
754
+ The rationale for the CROMA stage being mandatory: *"it is the only stage that bounds the input range.
755
+ Without it the encoder receives unbounded input."* The `use_8_bit` branch is the safer default because
756
+ the uint8 round-trip **guarantees** `[0, 1]`, whereas the float path's clip applies to a stretch that is
757
+ unbounded by construction (`mean Β± 2Β·std` is not a min/max clamp). The decision on `use_8_bit` is
758
+ **ENABLED**.
759
+
760
+ ### 7.3 The implemented stage β€” `specialists/optical_sar/radiometry.py`
761
+
762
+ The **encoder-input** stage is implemented, in full, at `specialists/optical_sar/radiometry.py`. Its
763
+ module docstring states the boundary between what is and is not implemented:
764
+
765
+ > Implemented: the **encoder-input** stage β€” per-channel `mean +/- 2*std` -> `[0, 1]`, the transform the
766
+ > CROMA authors' own README instructs users to apply.
767
+ >
768
+ > NOT implemented: the percentile / dB **conditioning** stage that `base.yaml` names.
769
+
770
+ The transform is quoted verbatim from the CROMA README in the docstring:
771
+
772
+ ```
773
+ min_value = x[:, c].mean() - 2 * x[:, c].std()
774
+ max_value = x[:, c].mean() + 2 * x[:, c].std()
775
+ img = (x[:, c] - min_value) / (max_value - min_value) * 255.0
776
+ img = clip(img, 0, 255).to(uint8)
777
+ # before the forward pass:
778
+ x = x.float() / 255
779
+ ```
780
+
781
+ The module makes two **deliberate deviations** from the README, both documented:
782
+
783
+ 1. **Per sample, not per batch.** The README computes `x[:, c].mean()` over the whole batch, which makes
784
+ one image's encoding a function of its neighbours. This module computes the window **per sample**,
785
+ because *"a serving system whose whole point is reproducibility"* must not produce a different answer
786
+ because another request was batched alongside it. At `N = 1` the two are identical, and the README's
787
+ own example runs at `N = 1`.
788
+ 2. **Unbiased standard deviation (`ddof=1`).** Upstream is torch, whose `.std()` defaults to the unbiased
789
+ estimator; NumPy's defaults to `ddof=0`. The module uses `ddof=1` to match the reference
790
+ implementation, and states the numerical immateriality (~1.00003 factor at 120Γ—120) so nobody has to
791
+ guess which convention a number came from.
792
+
793
+ The key constants:
794
+
795
+ | Constant | Value | Meaning |
796
+ |---|---|---|
797
+ | `TRANSFORM_NAME` | `"per_channel_mean_pm_2std"` | Recorded on every trace |
798
+ | `WINDOW_SIGMAS` | `2.0` | The window half-width |
799
+ | `DEFAULT_USE_8_BIT` | `True` | The ruling's decision |
800
+ | `ENV_USE_8_BIT` | `"SATQUERY_CROMA_USE_8_BIT"` | Hash-exempt override channel |
801
+ | `MIN_FINITE_PIXELS` | `2` | Below this, a std is not meaningful |
802
+ | `STATUS_NORMALISED` / `STATUS_UNAVAILABLE` / `STATUS_DEGENERATE` | strings | Per-channel status |
803
+
804
+ `normalise_for_croma(array, *, use_8_bit, modality)` is **pure and deterministic** β€” the docstring
805
+ states: *"no sampling, no learned statistics, no RNG, and no dependence on batch composition. Same input
806
+ -> same output."* It accepts `(B, C, H, W)` or `(C, H, W)`, returns a `float32` array of the same shape
807
+ plus a `RadiometryReport`, and leaves skipped channels **exactly zero**.
808
+
809
+ ### 7.4 The zero-channel rule
810
+
811
+ This is the load-bearing part, and it is the C-1 discipline applied to radiometry. An unavailable channel
812
+ is all zeros (the sensor adapter's convention). Such a channel has `std = 0`, so `max_value - min_value
813
+ == 0` and the stretch divides by zero. The module therefore **skips and leaves at exactly zero** any
814
+ channel with no dynamic range, and records a **different status** for each reason:
815
+
816
+ - `unavailable` β€” the channel is all zeros (`if not np.any(channel)`), i.e. the sensor did not measure it.
817
+ - `degenerate` β€” the channel is present but constant (`_channel_window` returns `None` because
818
+ `upper - lower <= 0`), i.e. the sensor measured a flat band.
819
+ - `normalised` β€” the real case.
820
+
821
+ The docstring explains why the distinction matters: *"'the sensor did not measure this' is not the same
822
+ fact as 'the sensor measured a flat band', and the trace must not merge them."* The mechanism is the
823
+ degenerate-window test, which needs **no mask** β€” deliberately, because `CROMAEncoder.encode` has a
824
+ pinned signature of exactly `(self, optical, sar)` and finding C-1 says CROMA never receives a mask.
825
+ Inferring unavailability from the data keeps that guard intact.
826
+
827
+ A non-finite pixel (NaN/Inf) is *"a defective measurement, not a measurement of zero"*; it is pinned to
828
+ `0.0` rather than allowed to propagate a NaN into the frozen encoder, where *"one NaN would poison the
829
+ whole forward pass."* The `finite_pixels` count in the report keeps the gap detectable.
830
+
831
+ `resolve_use_8_bit(config)` resolves the flag in a documented order β€” environment variable first, then
832
+ `croma.use_8_bit` when the key exists and is a bool, then `DEFAULT_USE_8_BIT` β€” and returns
833
+ `(value, source)` where `source` is `"env"`, `"config"` or `"default"`. The reason it is not simply a
834
+ config key is recorded: adding `croma.use_8_bit` to `configs/base.yaml` would move `Config.hash` away
835
+ from the frozen `78f1e3700da15aa1` and detach the Phase 9 benchmark. The hash-exempt channel is the same
836
+ pattern already used for `SATQUERY_DEVICE`.
837
+
838
+ `summarise(optical_report, sar_report)` bundles both modalities under one `use_8_bit` and **asserts** they
839
+ agree, raising `ValueError` if they do not: *"upstream applies one transform to both."*
840
+
841
+ ### 7.5 The display stretches β€” a different transform
842
+
843
+ Two other percentile stretches exist in the codebase. Neither is the CROMA transform, and the DEV-2
844
+ ruling names reusing one of them as **forbidden**.
845
+
846
+ **`preprocessing/imagery.py::load_image_array`** produces a displayable `(H, W, 3)` uint8 array:
847
+
848
+ ```python
849
+ arr = array.astype(np.float32)
850
+ finite = arr[np.isfinite(arr)]
851
+ if finite.size:
852
+ lo, hi = np.percentile(finite, (2, 98))
853
+ if hi > lo:
854
+ arr = (arr - lo) / (hi - lo)
855
+ arr = np.clip(arr, 0.0, 1.0)
856
+ return (arr * 255).astype(np.uint8)
857
+ ```
858
+
859
+ Its docstring is explicit about its scope: it *"does not resample, crop, or reproject. Those change the
860
+ pixel grid, and the grounding specialist converts normalized boxes to pixel coordinates using the
861
+ ORIGINAL raster's dimensions β€” a silent resize here would put every box in the wrong place."* The
862
+ determinism claim is stated: *"the same file always yields the same array, so a grounding box and a VQA
863
+ answer describe identical pixels."*
864
+
865
+ `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§1.1 records the three disqualifying differences between
866
+ this stretch and CROMA's: it is **per-scene** (one `(lo, hi)` over all bands) rather than **per-channel**;
867
+ it returns **uint8 3-band** rather than float 12-channel; and its **purpose** is presentation, not
868
+ radiometry.
869
+
870
+ **`specialists/change/postprocess.py::to_grayscale_float`** applies a per-scene 2/98 percentile stretch
871
+ to a grayscale collapse, so that *"a uint16 raster and a uint8 raster of the same scene correlate
872
+ identically."* Its comment records a measured defect fixed in place: `np.percentile` returns float64
873
+ scalars even for float32 input, so the array is promoted; without the explicit
874
+ `.astype(np.float32, copy=False)` the function returns float64 and `cv2.phaseCorrelate` fails its own
875
+ type assertion against a `CV_32F` Hanning window β€” *"Measured, not theorised."*
876
+
877
+ ### 7.6 The unimplemented stage, and its consequence
878
+
879
+ The percentile/dB conditioning stage is **not implemented**. The radiometry module's docstring states the
880
+ consequence plainly:
881
+
882
+ > **Consequence: the two `optical.*` and three `sar.*` config keys are still read by no code.** That is
883
+ > recorded, not fixed.
884
+
885
+ `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§6 item 3 adds: *"A config key that no code reads is worse
886
+ than a missing key: it reads as a satisfied requirement."* The ruling's follow-on table assigns the
887
+ implementation to the engineer, with the alternative of explicitly retiring the keys.
888
+
889
+ The status of the whole chain is therefore:
890
+
891
+ | Stage | Declared | Implemented | Read by code |
892
+ |---|---|---|---|
893
+ | Optical percentile 2/98 conditioning | βœ… `optical.*` | ❌ | ❌ |
894
+ | SAR dB clip βˆ’30/+5 conditioning | βœ… `sar.*` | ❌ | ❌ |
895
+ | Per-channel `mean Β± 2Β·std` β†’ [0,1] | βœ… (ruling Β§3.1) | βœ… `radiometry.py` | βœ… (via `training/fusion/extract.py`; see below) |
896
+
897
+ **Where the implemented stretch actually runs.** The serving-path pipeline
898
+ (`specialists/optical_sar/inference.py::run_pipeline`) does **not** call `normalise_for_croma`. It
899
+ canonicalises, resizes to 120Γ—120 and calls `encoder.encode(...)`. The stretch is applied by the
900
+ **extraction** path (`training/fusion/extract.py`), which states: *"This pipeline applies the DEV-2
901
+ encoder-input stretch itself (`radiometry.normalise_for_croma`), because the arm's `use_8_bit` has to be
902
+ applied at exactly one place."* `extract.py` also refuses an encoder that would stretch a second time:
903
+ `_assert_raw_encoder` *"refuses it by name rather than trusting the caller."* So the implemented
904
+ radiometric transform is a **training-time** transform that feeds the cached features, and the serving
905
+ path's input scaling is whatever `CROMAEncoder.encode` does internally (see Β§8.9 and the CROMA module).
906
+
907
+ `UNKNOWN β€” not established from the available evidence` whether the serving path's encoder applies the
908
+ same `mean Β± 2Β·std` stretch internally; `specialists/optical_sar/croma.py::CROMAEncoder.encode` was
909
+ verified to have the pinned signature `(self, optical, sar)` with no mask parameter, and the extraction
910
+ module's guard implies the encoder *can* apply a stretch when `normalize_input=True` (its default), but
911
+ the exact serving configuration of that flag is not established from the files read for this chapter.
912
+
913
+ ---
914
+
915
+ ## 8. The CROMA sensor adapter contract
916
+
917
+ `specialists/optical_sar/sensor_adapter.py` is *"the one place where 'some sensor's bands' becomes
918
+ 'CROMA's canonical channels', and it exists so that translation is explicit, inspectable and testable."*
919
+ Its module docstring states **the two hard rules** from the architecture plan:
920
+
921
+ > "No invented missing bands."
922
+ > "Do not fabricate spectral bands."
923
+
924
+ The failure mode the module prevents is described with unusual clarity:
925
+
926
+ > A 4-band Cartosat-2S scene maps cleanly onto canonical channels 1-4. Filling channels 5-12 with
927
+ > *something* β€” a copy of B3 as a fake "red-edge", a zero-order hold, an interpolation β€” produces a tensor
928
+ > whose shape is correct and whose statistics look plausible. CROMA would run. Nothing would raise. And
929
+ > every downstream number would be computed from channels that no sensor ever measured, presented as if
930
+ > they were measurements.
931
+
932
+ So: **missing channels are zero-filled, and the availability mask says which ones were real.**
933
+
934
+ ### 8.1 `SensorDescriptor`
935
+
936
+ The schema object (`core/schemas.py:138`) is `extra="forbid"`, described as the *"Frozen sensor-adapter
937
+ contract (plan section 18)"*:
938
+
939
+ | Field | Type | Default | Meaning |
940
+ |---|---|---|---|
941
+ | `sensor` | `str` | required | Sensor identifier, e.g. `"cartosat_2s"` |
942
+ | `available_bands` | `list[str]` | `[]` | The bands the sensor actually delivers |
943
+ | `band_map` | `dict[str, str]` | `{}` | Source band name β†’ canonical band name |
944
+ | `normalization` | `str` | `"percentile"` | Identifier recorded for the trace |
945
+ | `availability_mask` | `list[bool]` | `[]` | Which canonical channels are real |
946
+ | `resolution` | `float \| None` | `None` | Ground sample distance in metres |
947
+ | `canonical_channels` | `int` | `0` | Target channel count (12 optical, 2 SAR) |
948
+ | `missing_channels_zero_filled` | `bool` | `True` | The zero-fill convention |
949
+
950
+ **Why `band_map` is canonical-channel-keyed.** The module docstring addresses a direction question: the
951
+ plan's example writes `band_map` sensor-keyed (`{"B1": "c1", ...}`), and the schema types it as
952
+ `dict[str, str]`. This module uses the schema's direction because the question asked at load time is
953
+ forward-looking: *"canonical channel 3 β€” did I get a real band for it, and if so which one?"*
954
+ `canonical_to_sensor` is exposed as the inverse for reporting.
955
+
956
+ ### 8.2 The canonical channel orders
957
+
958
+ ```python
959
+ OPTICAL_CANONICAL: tuple[str, ...] = (
960
+ "B01", # coastal aerosol
961
+ "B02", # blue
962
+ "B03", # green
963
+ "B04", # red
964
+ "B05", # red edge 1
965
+ "B06", # red edge 2
966
+ "B07", # red edge 3
967
+ "B08", # NIR
968
+ "B8A", # narrow NIR
969
+ "B09", # water vapour
970
+ "B11", # SWIR 1
971
+ "B12", # SWIR 2
972
+ )
973
+
974
+ SAR_CANONICAL: tuple[str, ...] = ("VV", "VH")
975
+ ```
976
+
977
+ The comment above `OPTICAL_CANONICAL` states the contract: *"The ORDER is part of the contract: channel
978
+ index i must always mean the same physical band, or a pretrained encoder's weights would be applied to
979
+ the wrong input."* This is Sentinel-2 L2A with the cirrus band removed, per CROMA's README: *"Sentinel-2
980
+ must be 12 channels (remove the cirrus band if necessary)."* Note that `B10` (cirrus) is **not** in the
981
+ canonical order β€” that is the removal.
982
+
983
+ `SAR_CANONICAL` is described as *"Sentinel-1 dual-pol order CROMA was pretrained with. Positional fallback
984
+ ONLY."*
985
+
986
+ `_BAND_ALIASES` maps common spellings onto the canonical form: `B1`β†’`B01`, `BLUE`β†’`B02`, `NIR`β†’`B08`,
987
+ `SWIR1`β†’`B11`, `VVPOL`β†’`VV`, and so on. The comment is a caution: *"Kept small and explicit: a guess here
988
+ becomes a silently mislabelled band."* `_canonicalise` strips whitespace and underscores and uppercases
989
+ before lookup, and **never invents** a band β€” an unmapped name passes through unchanged and is later
990
+ rejected.
991
+
992
+ ### 8.3 Building an optical adapter
993
+
994
+ `build_optical_adapter(sensor, *, available_bands, canonical_channels=12, normalization="percentile",
995
+ resolution=None)` (`specialists/optical_sar/sensor_adapter.py:205`):
996
+
997
+ - When `available_bands` is `None`, the sensor is looked up in `_KNOWN_OPTICAL_SENSORS`; an unknown
998
+ sensor with no explicit band list raises `UnsupportedBandsError` β€” *"refusing to guess a band
999
+ arrangement."*
1000
+ - Each band is canonicalised. Any canonical name **not** in `OPTICAL_CANONICAL` raises
1001
+ `UnsupportedBandsError`: *"do not map onto the canonical Sentinel-2 order ... supply an explicit
1002
+ band_map rather than guessing a position."*
1003
+ - Duplicate bands raise.
1004
+ - More bands than `canonical_channels` raises.
1005
+ - `band_map` is built as `{source: canonical for source, canonical in zip(bands, canonical)}` β€” keyed by
1006
+ the **source** name, because *"that is what an array row can actually be looked up by."*
1007
+ - `availability` is `[canonical_name in band_map.values() for canonical_name in OPTICAL_CANONICAL[:canonical_channels]]`.
1008
+
1009
+ The returned descriptor has `missing_channels_zero_filled=True`.
1010
+
1011
+ `_KNOWN_OPTICAL_SENSORS` is *"deliberately tiny and explicit. Adding one is a code change, not a guess at
1012
+ runtime"*:
1013
+
1014
+ | Sensor | Bands |
1015
+ |---|---|
1016
+ | `sentinel_2` | All 12 canonical |
1017
+ | `cartosat_2s` | `["B1", "B2", "B3", "B4"]` |
1018
+ | `cartosat_2` | `["B1", "B2", "B3", "B4"]` |
1019
+ | `landsat_8` | `["B1", "B2", "B3", "B4"]` |
1020
+
1021
+ The comment on `cartosat_2s` is a standing instruction: *"Cartosat-2S is a 4-band VNIR imager. It is NOT a
1022
+ Sentinel-2 clone and must not be treated as one: the 8 channels it lacks stay masked-off."*
1023
+
1024
+ ### 8.4 Describing a SAR sensor β€” inspect, do not assume
1025
+
1026
+ `describe_sar(sensor, *, available_bands, canonical_channels=2, normalization="db", resolution=None)`
1027
+ (`specialists/optical_sar/sensor_adapter.py:318`) implements the plan's instruction:
1028
+
1029
+ > "RISAT imagery may vary in acquisition/polarization characteristics... the SAR adapter must inspect
1030
+ > actual available channels rather than assuming one fixed polarization pair."
1031
+
1032
+ The module docstring explains why a hardcoded `["VV", "VH"]` would be wrong: *"ISRO describes RISAT-1 as
1033
+ a C-band SAR mission with multiple polarization configurations (single HH, single VV, dual HH+HV, dual
1034
+ VV+VH, ...). A hardcoded `["VV", "VH"]` would silently mislabel an HH/HV scene, and the resulting tensor
1035
+ would be a *confidently wrong* input rather than a recognisably broken one."*
1036
+
1037
+ When `available_bands` names the polarisations, they are honoured. When it does not, the **channel count**
1038
+ is used to select the most likely convention, and **the choice is recorded on the descriptor's
1039
+ `available_bands`** so the assumption appears in the trace rather than being invisible.
1040
+
1041
+ The **slot order is the sensor's own**:
1042
+
1043
+ ```python
1044
+ band_map: dict[str, str] = {source: source for source in bands}
1045
+ availability = [i < len(bands) for i in range(canonical_channels)]
1046
+ ```
1047
+
1048
+ The comment explains: *"Forcing HH/HV into a VV/VH frame would mean either relabelling the polarisations
1049
+ (a lie the mask would then contradict) or masking both off and discarding real data."* So the tensor holds
1050
+ the radar measurements that exist, and `available_bands[i]` says which polarisation slot `i` is.
1051
+
1052
+ `_known_sar(sensor)` is described as *"Best-effort polarisation layout for a named SAR sensor, or `[]`"*:
1053
+
1054
+ | Sensor key | Returns |
1055
+ |---|---|
1056
+ | `sentinel_1`, `sentinel1` | `["VV", "VH"]` |
1057
+ | `risat*` | `["VV", "VH"]` β€” **recorded as an ASSUMPTION, not a fact** |
1058
+ | `alos_palsar`, `palsar`, `alos2` | `["HH", "HV"]` |
1059
+ | anything else | `[]` β€” *"An empty return is honest: it says 'I do not know this sensor's polarisations'"* |
1060
+
1061
+ The `risat` case carries an explicit comment: *"RISAT-1's most common operational mode is dual-pol.
1062
+ Recorded as an ASSUMPTION, not a fact β€” see the module docstring. Callers with real band metadata should
1063
+ pass `available_bands` and override this."* This is an `ASSUMPTION` in the project's status vocabulary,
1064
+ and it is the reason the hidden-set behaviour on RISAT cannot be claimed as measured.
1065
+
1066
+ ### 8.5 Applying the canonical layout β€” `_apply_canonical`
1067
+
1068
+ This is *"the function the two hard rules live in"* (`specialists/optical_sar/sensor_adapter.py:428`). It
1069
+ places real bands in canonical positions and zero-fills the rest:
1070
+
1071
+ ```python
1072
+ canonical = np.zeros((n_canonical, height, width), dtype=np.float32)
1073
+ mask = np.zeros((n_canonical,), dtype=bool)
1074
+
1075
+ for source_index, source_name in enumerate(descriptor.available_bands):
1076
+ if source_index >= n_source:
1077
+ continue
1078
+ canonical_name = descriptor.band_map.get(source_name)
1079
+ if canonical_name is None or canonical_name not in order:
1080
+ continue
1081
+ channel_index = order.index(canonical_name)
1082
+ if channel_index >= n_canonical:
1083
+ continue
1084
+ canonical[channel_index] = array[source_index].astype(np.float32)
1085
+ mask[channel_index] = True
1086
+ ```
1087
+
1088
+ The docstring's key claim: *"There is no branch anywhere below that writes a value into an unavailable
1089
+ channel."* The absence of such a branch **is** the mask. If no band could be placed while the array
1090
+ carries bands, it raises `UnsupportedBandsError` rather than returning an all-zero tensor silently.
1091
+
1092
+ `_canonical_order(descriptor)` decides the slot order: for a descriptor whose `available_bands` are all
1093
+ SAR polarisations, the order is the descriptor's own list (padded to canonical width); otherwise it is
1094
+ `OPTICAL_CANONICAL`.
1095
+
1096
+ `adapt_optical` and `adapt_sar` are thin wrappers that call `_apply_canonical` with the right modality
1097
+ label. `adapt_sar`'s docstring names the plan's internal representation exactly:
1098
+ `canonical_sar[2, H, W]` + `sar_channel_mask[2]`.
1099
+
1100
+ ### 8.6 `SensorAdapterOutput`
1101
+
1102
+ The dataclass returned by `adapt_*` carries `canonical` `(C, H, W)` float32, `mask` `(C,)` bool, and the
1103
+ `descriptor`. It exposes `n_channels`, `missing_indices` (`[i for i, ok in enumerate(self.mask) if not
1104
+ ok]`), `available_bands`, and a `to_dict()` that includes `n_available` and `missing_channels`.
1105
+
1106
+ ### 8.7 `positional_fallback_descriptor`
1107
+
1108
+ Used when a raster has the right **channel count** for a modality but **no band labels**. The fallback
1109
+ order is chosen from `_FALLBACK_SAR_ORDER` (`("VV", "VH")`, `("HH", "HV")`) for 1–2 channels, or the
1110
+ canonical optical prefix for β‰₯3 channels. The fact that it was a fallback is **recorded in the `sensor`
1111
+ field** so it *"cannot be mistaken for measured metadata."* The `complement=True` flag selects the other
1112
+ convention, which lets a caller test both HH/HV and VV/VH orderings without asserting which is correct.
1113
+
1114
+ The optical-SAR specialist uses this in `_descriptor_for`
1115
+ (`specialists/optical_sar/specialist.py:398`) and appends a warning that names the assumption: *"The
1116
+ polarisation IDENTITY of each channel is NOT known from a band count alone β€” supply a sensor descriptor
1117
+ to state it."*
1118
+
1119
+ ### 8.8 Finding C-1 β€” the mask goes to the fusion head, not to CROMA
1120
+
1121
+ This is the single most important fact in this section. `core/schemas.py`'s module docstring lists it
1122
+ among the findings baked into the schemas: *"C-1 channel-availability mask is a first-class fusion input,
1123
+ never a CROMA input."*
1124
+
1125
+ `specialists/optical_sar/fusion_head.py` devotes a section to the rationale:
1126
+
1127
+ > CROMA is a masked autoencoder. Handing it an availability mask invites it to reconstruct the missing
1128
+ > channels β€” which is precisely the fabrication the sensor adapter exists to prevent: the model would
1129
+ > output plausible values for bands no sensor measured, and those values would then be treated as data.
1130
+ > The mask is therefore consumed HERE, where it can only do one thing: tell the classifier which inputs
1131
+ > to distrust.
1132
+
1133
+ Mechanically: `CROMAEncoder.encode` has a pinned signature of exactly `(self, optical, sar)` β€” the
1134
+ radiometry module records that `tests/unit/test_optical_sar_croma.py` *"asserts the absence of any mask
1135
+ parameter, because freeze finding C-1 says CROMA never receives one."* The mask instead enters at
1136
+ `assemble_fusion_input` (Β§8.9). The optical-SAR specialist's `JOINT_FEATURE_REGION` evidence item records
1137
+ the fact explicitly in its payload: `"mask_consumed_by": "fusion_head"` and `"croma_received_mask":
1138
+ False` (`specialists/optical_sar/specialist.py:816`).
1139
+
1140
+ `EvidenceType` even has a dedicated member for the mask as trust evidence:
1141
+ `AVAILABILITY_MASK = "availability_mask" # C-1: modality trust evidence` (`core/schemas.py:76`). The
1142
+ optical-SAR specialist emits one `AVAILABILITY_MASK` item per modality, carrying the sensor, canonical
1143
+ channel count, available bands, band map, the mask itself, the missing indices,
1144
+ `missing_channels_zero_filled`, `normalization` and `resolution`.
1145
+
1146
+ ### 8.9 The frozen fusion concatenation
1147
+
1148
+ `specialists/optical_sar/fusion_head.py` defines the concatenation that the mask feeds into:
1149
+
1150
+ ```
1151
+ optical_GAP (B, 768)
1152
+ SAR_GAP (B, 768)
1153
+ joint_GAP (B, 768)
1154
+ optical_mask (B, 12) <- availability, from the sensor adapter
1155
+ sar_mask (B, 2) <- availability, from the sensor adapter
1156
+ ---------
1157
+ concat (B, 2318)
1158
+ ```
1159
+
1160
+ `expected_fusion_dim(encoder_dim=768, optical_channels=12, sar_channels=2, modalities_used=3)` returns
1161
+ `3*768 + 12 + 2 = 2318`. The module recomputes the width and **refuses to build on a mismatch**, so *"a
1162
+ config edit cannot silently reshape the first Linear layer into something that trains but means
1163
+ nothing."* `core/config.py` recomputes it independently at load time from `croma.encoder_dim`,
1164
+ `croma.modalities_used`, `croma.optical_channels` and `croma.sar_channels`, and rejects a config whose
1165
+ `fusion.input_dim` disagrees (finding C-1 in the validator's own message).
1166
+
1167
+ `assemble_fusion_input` asserts the order rather than assuming it β€” *"a permutation here produces a tensor
1168
+ of exactly the right shape that trains to a worse number β€” the hardest kind of bug to notice"* β€” and
1169
+ validates that the three GAP vectors agree on batch size and that both masks agree with it.
1170
+
1171
+ `channel_dropout(features, mask, *, keep_probabilities, rng)` is the mechanism that *"teaches the head to
1172
+ trust the availability mask"*. Its load-bearing line is:
1173
+
1174
+ ```python
1175
+ keep = draws & (msk > 0.0)
1176
+ out_mask = keep.astype(np.float32)
1177
+ out = out * out_mask
1178
+ ```
1179
+
1180
+ Dropped channels are **zeroed, not renormalised** β€” *"renormalising would fabricate a scale that the real
1181
+ missing-channel case does not have"* β€” and the mask is updated in lockstep. `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md`
1182
+ Β§5 independently verified this against freeze Β§2.5 and ruled: *"the engineer's implementation SATISFIES
1183
+ Β§2.5's intent. No change entry required."* Its table records that the `& (msk > 0.0)` guard is the
1184
+ load-bearing one: *"no training transform can resurrect a band the sensor did not measure."*
1185
+
1186
+ The schedule (plan section 20) is optical at `100/80/60/40%` and SAR at `100/50%`, declared in
1187
+ `training/fusion/train.py` as `OPTICAL_DROPOUT_RATES = (1.0, 0.8, 0.6, 0.4)` and
1188
+ `SAR_DROPOUT_RATES = (1.0, 0.5)`.
1189
+
1190
+ `mask_availability_stats(optical_mask, sar_mask)` returns `optical_available_fraction`,
1191
+ `sar_available_fraction`, `optical_channels_present` and `sar_channels_present`. These feed the
1192
+ confidence system: the two modality-confidence terms in the optical-SAR specialist are exactly the
1193
+ availability fractions, which the specialist's docstring justifies: *"On the hidden set the optical side
1194
+ is Cartosat-2S with 4 bands against a canonical 12, so 8 of 12 channels are always absent. A classifier
1195
+ resting on 4 measured channels is not the same claim as one resting on 12."*
1196
+
1197
+ ### 8.10 The optical-SAR confidence components
1198
+
1199
+ `OpticalSarSpecialist._confidence_components` (`specialists/optical_sar/specialist.py:481`) implements
1200
+ plan section 26's four components with these weights:
1201
+
1202
+ ```python
1203
+ CONFIDENCE_WEIGHTS: dict[str, float] = {
1204
+ "fusion_margin": 0.40,
1205
+ "optical_confidence": 0.20,
1206
+ "sar_confidence": 0.20,
1207
+ "cross_modal_agreement": 0.20,
1208
+ }
1209
+ ```
1210
+
1211
+ `cross_modal_agreement` is `cosine(optical_GAP, SAR_GAP)`, rescaled from `[-1, 1]` to `[0, 1]` via
1212
+ `AGREEMENT_FLOOR = -1.0`. `cross_modal_agreement` in `inference.py` returns `None` on any exception or a
1213
+ shape mismatch, and the specialist treats `None` as a zero component. The composition is a weighted sum,
1214
+ **floored to 0.0** when there is no prediction or no trained head β€” *"Both are signal GAPS, not weak
1215
+ signals, so they resolve to 0.0 rather than merely discounting."* The components dict records
1216
+ `has_croma`, `trained_head`, `no_prediction` and the raw availability counts, so *"the trace states every
1217
+ reason the score is zero, not just the first one encountered."*
1218
+
1219
+ ---
1220
+
1221
+ ## 9. Pair-compatibility rules
1222
+
1223
+ Two workflows take a **pair** of assets, and both enforce compatibility before reading pixels.
1224
+
1225
+ ### 9.1 Change β€” exactly two, temporally distinct, co-registered
1226
+
1227
+ `ChangeSpecialist.validate_request` (`specialists/change/specialist.py:190`) enforces **exactly two**
1228
+ assets. The comment names both common mistakes: *"ONE asset is the common mistake: the caller treated
1229
+ this as a single-image task. THREE is the other: they attached a time series. Both are rejected with the
1230
+ same typed error."* The user-facing message is *"Change detection needs exactly two images: an earlier one
1231
+ and a later one."*
1232
+
1233
+ `_assess_pair` then runs three checks in order:
1234
+
1235
+ 1. **Temporal distinctness.** Identical `path` raises `TemporalPairError`. Identical `sha256` (when both
1236
+ carry one) also raises: *"a repeated acquisition is not a temporal pair."*
1237
+ 2. **CRS compatibility** via `compare_crs` (Β§2.5). Incompatible or reprojection-requiring pairs get a
1238
+ warning, not a refusal β€” pixel-domain change detection can still proceed, but geographic claims are
1239
+ withheld.
1240
+ 3. **Registration quality** via `measure_registration` (Β§9.3).
1241
+
1242
+ If registration is **not usable**, the specialist **suppresses spatial claims** while still computing and
1243
+ returning the change maps:
1244
+
1245
+ > The maps are still returned (the caller may want to look at them), but the REGION CLAIMS are withheld:
1246
+ > a region is an assertion about *where* something changed on the ground, and we cannot make it.
1247
+
1248
+ `ChangeSpecialist._geospatial` drops the transform in that case, returning a `GeoMetadata` with
1249
+ `is_georeferenced=False` β€” *"a transform invites the caller to place a region on a map, and we have just
1250
+ said the regions are not placeable."* `_write_artifacts` writes the map **without** georeferencing when
1251
+ suppressed, and the comment explains why copying T1's transform would be wrong: *"Copying T1's transform
1252
+ onto it would produce a file that unlocks exactly the placed-on-a-map reading the suppression exists to
1253
+ forbid."* The confidence is driven to the floor via a **minimum** against the registration factor, so
1254
+ alignment is a gate, not one vote among three.
1255
+
1256
+ The change confidence weights are:
1257
+
1258
+ ```python
1259
+ CONFIDENCE_WEIGHTS: dict[str, float] = {
1260
+ "registration_quality": 0.50,
1261
+ "mean_change_probability": 0.30,
1262
+ "component_stability": 0.20,
1263
+ }
1264
+ ```
1265
+
1266
+ with `STABILITY_HEADROOM = 0.5`, `DEFAULT_MAX_SHIFT_PX = 8`, `MIN_REGION_PIXELS = 32`.
1267
+
1268
+ ### 9.2 Optical-SAR β€” exactly two, one of each modality
1269
+
1270
+ `OpticalSarSpecialist.validate_request` (`specialists/optical_sar/specialist.py:192`) enforces **exactly
1271
+ two** assets **and** that they are one optical and one SAR. The docstring explains the case worth
1272
+ explaining:
1273
+
1274
+ > Two optical images is not a malformed request in the way that three is β€” it is a plausible mistake,
1275
+ > from an operator who uploaded a bi-temporal pair to a single-modality workflow, or who labelled a
1276
+ > four-band Cartosat scene as SAR.
1277
+ >
1278
+ > If such a request were accepted, the adapter would place the second optical image into the 2-channel SAR
1279
+ > slot, zero-fill, and produce a fused tensor of the correct shape. CROMA would run. The answer would be
1280
+ > about one modality while claiming to fuse two.
1281
+
1282
+ So the modality check is done on the **assets, before any pixels are read**. `_assess_pair` raises
1283
+ `InvalidRequestError` with a machine-readable `reason` of `"two_optical"`, `"two_sar"` or
1284
+ `"indeterminate"`, and each carries a user-facing message.
1285
+
1286
+ ### 9.3 The equal-dimensions requirement, and the 726Β²-vs-736Β² error
1287
+
1288
+ The change detector requires T1 and T2 to have **identical shapes**. The guard is in the model itself,
1289
+ `specialists/change/stanet.py:415`:
1290
+
1291
+ ```python
1292
+ if t1.shape != t2.shape:
1293
+ raise SpecialistError(
1294
+ f"T1 and T2 must have the same shape; got {tuple(t1.shape)} "
1295
+ f"and {tuple(t2.shape)}",
1296
+ specialist="change",
1297
+ )
1298
+ ```
1299
+
1300
+ This is a **hard** requirement, and it is where the bundled demo pair fails. `docs/FINAL_DELIVERY_REPORT.md`
1301
+ Β§6 records it as `DEGRADED`:
1302
+
1303
+ | Finding | Status | Detail |
1304
+ |---|---|---|
1305
+ | Change on the bundled EO pair | DEGRADED | `delta-growth-t0-1975.jpg` (726Β²) and `t1-2025.jpg` (736Β²) differ in shape; the change specialist errors (`T1 and T2 must have the same shape`). A same-shape pair, or a resize step, is needed for a clean change demo. |
1306
+
1307
+ There is **no resize step on the change path**. `load_image_array` explicitly does not resample
1308
+ (Β§7.5) β€” its docstring states *"a silent resize here would put every box in the wrong place"* β€” and the
1309
+ change specialist's `_predict` only pads for the encoder's stride-8 requirement, cropping back to the
1310
+ original size:
1311
+
1312
+ ```python
1313
+ h, w = t1.shape[-2:]
1314
+ if h % 8 or w % 8:
1315
+ ph, pw = (-h) % 8, (-w) % 8
1316
+ t1 = torch.nn.functional.pad(t1, (0, pw, 0, ph), mode="reflect")
1317
+ t2 = torch.nn.functional.pad(t2, (0, pw, 0, ph), mode="reflect")
1318
+ ```
1319
+
1320
+ That padding does **not** reconcile two different input sizes: T1 and T2 are each padded by their own
1321
+ `(-dim) % 8`, so a 726Β² and a 736Β² input become 728Β² and 736Β² respectively β€” still different β€” and the
1322
+ model raises. The reflect-pad is a *stride* accommodation, not a *pair* accommodation, and the two must
1323
+ not be conflated.
1324
+
1325
+ **Consequence for operators.** The change workflow requires a same-shape pair, and the failure is
1326
+ `IMPLEMENTED` (it raises a typed error with a readable message) rather than silent. A resize or
1327
+ co-registration step that reconciles differing dimensions is **not implemented** (Β§13). The
1328
+ `FINAL_DELIVERY_REPORT` records the demo as `DEGRADED` and prescribes *"a same-shape pair, or a resize
1329
+ step"* β€” either would fix the demo; neither is currently in the code path.
1330
+
1331
+ ### 9.4 Registration measurement
1332
+
1333
+ `measure_registration` (`specialists/change/postprocess.py:134`) uses `cv2.phaseCorrelate` on grayscale
1334
+ collapses of the two images, returning a `RegistrationQuality` with `shift_x`, `shift_y`, `response`,
1335
+ `max_shift_px`, `is_usable` and `reason`. Defaults: `max_shift_px=8`, `min_response=0.15`.
1336
+
1337
+ The module docstring is candid about the method's limit:
1338
+
1339
+ > Deliberately a TRANSLATION model. Real mis-registration includes rotation and warp, which phase
1340
+ > correlation does not recover; a small residual after a translation correction is therefore not proof of
1341
+ > good alignment. What the measurement can do honestly is flag the *bad* cases, and that is how it is
1342
+ > used: as a gate, not a certificate.
1343
+
1344
+ `RegistrationQuality.confidence_factor()` applies two independent penalties and takes the worse:
1345
+ a shift penalty (`1.0` at zero shift, `0.0` at `max_shift_px`) and a response factor (`0.0` at zero
1346
+ response, `1.0` at response β‰₯ 0.15). A pair that is offset *and* uncorrelated takes the worse of the two β€”
1347
+ *"the conservative choice."*
1348
+
1349
+ A failure to *measure* is treated as a measurement of failure β€” both mean the pair cannot be trusted
1350
+ spatially β€” but the distinction is recorded in the warning: *"A failure to MEASURE is not the same as a
1351
+ measurement of failure."*
1352
+
1353
+ Two registration-gate hypotheses were tested and both are recorded as **REJECTED** or
1354
+ **INSUFFICIENT**: `docs/CHANGE_REGISTRATION_GATE_COHERENCE_TEST.md` records a rejected hypothesis (AUC S2
1355
+ vs S3 ≀ 0.59), and `docs/CHANGE_REGISTRATION_GATE_TEXTURE_TEST.md` records a low-texture signal that was
1356
+ **CONFIRMED but INSUFFICIENT**, with option 3 **ELIMINATED**. So the gate as implemented (phase
1357
+ correlation) is the surviving method, and no additional coherence- or texture-based gate was adopted.
1358
+
1359
+ ---
1360
+
1361
+ ## 10. Large-image tiling policy
1362
+
1363
+ ### 10.1 What is declared
1364
+
1365
+ `configs/base.yaml` declares the full tiling policy:
1366
+
1367
+ ```yaml
1368
+ image:
1369
+ max_pixels: 25000000
1370
+ tile_size: 512
1371
+ tile_overlap: 128
1372
+ max_tiles: 64
1373
+ # plan section 9.1 tile policy: whole-image thumbnail first, then top-K tiles.
1374
+ # max_tiles is the hard ceiling on tiles *examined*; top_k_tiles is how many
1375
+ # are actually sent through a specialist.
1376
+ top_k_tiles: 4
1377
+ ```
1378
+
1379
+ | Key | Value | Declared meaning |
1380
+ |---|---|---|
1381
+ | `max_pixels` | 25,000,000 | Pixel budget; exceeding it is recoverable (the caller may downscale) |
1382
+ | `tile_size` | 512 | Tile edge in pixels |
1383
+ | `tile_overlap` | 128 | Overlap between adjacent tiles |
1384
+ | `max_tiles` | 64 | Hard ceiling on tiles **examined** |
1385
+ | `top_k_tiles` | 4 | How many tiles are actually sent through a specialist |
1386
+
1387
+ The comment states the intended policy: *"whole-image thumbnail first, then top-K tiles."* The change
1388
+ workflow declares a different tiling resolution, `change.tile_size: 256` with `change.tile_overlap: 0`
1389
+ β€” 256 is *"the change model's OWN training resolution"* (LEVIR-CD patches are 256Γ—256).
1390
+
1391
+ ### 10.2 What is enforced
1392
+
1393
+ `core/config.py` validates the tiling policy at load time, but only one relationship:
1394
+
1395
+ ```python
1396
+ # --- tiling policy -------------------------------------------------
1397
+ if self.get("image.top_k_tiles", 0) > self.get("image.max_tiles", 0):
1398
+ errors.append("image.top_k_tiles cannot exceed image.max_tiles")
1399
+ ```
1400
+
1401
+ So the config loader guarantees that the number of tiles sent through a specialist does not exceed the
1402
+ number examined. It does not implement tiling.
1403
+
1404
+ The pixel budget **is** enforced, but at the raster-inspection layer rather than as a tiling step:
1405
+ `inspect_raster` raises `OversizedImageError` when `width * height > max_pixels`, and the error is
1406
+ `recoverable` β€” the docstring says *"the caller may downscale."* The controller passes the configured
1407
+ budget into `_inspect_asset` during VALIDATE.
1408
+
1409
+ ### 10.3 What is not implemented β€” the honest status
1410
+
1411
+ **The tiling policy is `DECLARED (not read)`.** `image.tile_size`, `image.tile_overlap`, `image.max_tiles`
1412
+ and `image.top_k_tiles` appear only in `configs/base.yaml` and in the `core/config.py` validation above.
1413
+ No code on the serving path slices a raster into tiles, selects the top-K, or stitches results.
1414
+ `architecture/03-request-lifecycle.md` reaches the same conclusion: the tiling policy is referenced only
1415
+ in config validation and not implemented on the serving path.
1416
+
1417
+ The practical consequence: a raster within the `max_pixels` budget is handed to a specialist **whole**.
1418
+ The VLM path pins `processor_longest_edge: 512` so the processor does not upscale and split a tile, but
1419
+ that pin assumes a 512-px input β€” it is a *control against* the processor's default 2048, not a tiling
1420
+ step. `core/config.py` enforces the relationship (`processor_longest_edge` must not exceed
1421
+ `image.tile_size`), with the measured rationale recorded: the default 2048 upscales a 512 tile 4Γ— and then
1422
+ splits it into ~17 sub-images, *"the real figure is ~17x"* against the plan's estimated 4Γ—.
1423
+
1424
+ `UNKNOWN β€” not established from the available evidence` how a raster between the tile size and the pixel
1425
+ budget is handled by the VLM path in practice; the config implies the tile policy would govern it, but
1426
+ the policy is not implemented, so the behaviour is whatever the specialist does with the whole image.
1427
+
1428
+ ---
1429
+
1430
+ ## 11. Grounding resolution β€” the frozen 224 decision
1431
+
1432
+ Grounding runs at **224 px**, and this is one of the project's cleanest pre-registered results.
1433
+ `configs/base.yaml`:
1434
+
1435
+ ```yaml
1436
+ grounding:
1437
+ ...
1438
+ # RESOLVED 2026-09-16. 448 did NOT earn its cost over all 16,159 VRSBench
1439
+ # eval records: mean best IoU -0.0147, Recall@0.5 -0.0022, and every recall
1440
+ # threshold lower (-0.0699 @0.10, -0.0243 @0.25), at 1.59x the latency.
1441
+ # Paired over identical samples: mean diff -0.0147, 95% CI [-0.0160,-0.0134],
1442
+ # t = -22.63. 448 was better on 8.5% of records, worse on 20.9%.
1443
+ # Pre-registered rule and the paired test AGREE on 224.
1444
+ # Evidence: docs/PHASE7_RESOLUTION_DECISION.md
1445
+ image_size: 224
1446
+ resolution_frozen: true
1447
+ nms_iou: 0.50
1448
+ max_candidates: 20
1449
+ confidence_threshold: 0.40
1450
+ benchmark_box_scale: 100.0
1451
+ coordinate_system: normalized_0_1
1452
+ encoder_projected_dim: 512
1453
+ ```
1454
+
1455
+ ### 11.1 The pre-registered rule
1456
+
1457
+ `docs/PHASE7_RESOLUTION_DECISION.md` records the rule, **fixed before the result was seen**:
1458
+
1459
+ ```
1460
+ 448 WINS if Recall@0.5 improves by >= 0.05 absolute
1461
+ OR mean best IoU improves by >= 0.05 absolute
1462
+ 224 WINS otherwise
1463
+ INCONCLUSIVE if fewer than 30 samples were scored
1464
+ ```
1465
+
1466
+ The artifact records `rule_changed_since_preregistration: false`.
1467
+
1468
+ ### 11.2 The result
1469
+
1470
+ ```
1471
+ 224 WINS
1472
+ recall@0.5 gain 448/224 : -0.0022
1473
+ bestIoU gain 448/224 : -0.0147
1474
+ latency ratio : 1.59x
1475
+ ```
1476
+
1477
+ Neither component came close to the +0.05 margin; both were **negative**. The measured detail:
1478
+
1479
+ | Metric | 224 | 448 | Delta |
1480
+ |---|---|---|---|
1481
+ | Token grid | 7 Γ— 7 = 49 | 14 Γ— 14 = 196 | 4.0Γ— tokens |
1482
+ | Mean best IoU | **0.0972** | 0.0825 | **βˆ’0.0147** |
1483
+ | Recall@0.10 | **0.3298** | 0.2599 | **βˆ’0.0699** |
1484
+ | Recall@0.25 | **0.1187** | 0.0944 | **βˆ’0.0243** |
1485
+ | Recall@0.50 | **0.0234** | 0.0212 | **βˆ’0.0022** |
1486
+ | Latency mean | **20.0 ms** | 31.8 ms | 1.59Γ— |
1487
+ | Latency p90 | **20.9 ms** | 32.9 ms | 1.57Γ— |
1488
+ | Peak VRAM | **592.1 MB** | 599.8 MB | +7.7 MB |
1489
+
1490
+ All 16,159 / 16,159 VRSBench eval records were scored at both resolutions on a Tesla T4.
1491
+
1492
+ ### 11.3 The paired test
1493
+
1494
+ Because both resolutions scored the **same 16,159 samples**, the paired test is the stronger statistic:
1495
+
1496
+ ```
1497
+ paired samples : 16159
1498
+ mean 224 : 0.0972
1499
+ mean 448 : 0.0825
1500
+ mean paired diff : -0.0147 (95% CI -0.0160 .. -0.0134)
1501
+ t statistic : -22.63
1502
+ CI excludes zero : True
1503
+
1504
+ 448 better on : 1371/16159 ( 8.5%)
1505
+ 448 worse on : 3372/16159 (20.9%)
1506
+ identical : 11416/16159 (70.6%)
1507
+ ```
1508
+
1509
+ The paired test and the pre-registered rule **agree**, so there is *"no rule-versus-evidence disagreement
1510
+ to escalate."* The win/loss split is informative on its own: 448 wins on only 8.5% and loses on 20.9% β€”
1511
+ *"the finer grid is not merely neutral, it is actively harmful on a fifth of the corpus."*
1512
+
1513
+ The recall ladder narrows as the threshold rises (βˆ’0.0699 at 0.10, βˆ’0.0243 at 0.25, βˆ’0.0022 at 0.50),
1514
+ which the document reads as *"the signature of a method that cannot reach high IoU either way."*
1515
+
1516
+ ### 11.4 Why 448 did not help
1517
+
1518
+ The document's honest reading: the zero-shot method selects a patch by text similarity and returns that
1519
+ patch's box. At 224 a box is 1/7 of the image; at 448 it is 1/14. A finer grid is only better if the
1520
+ target is small **and** the similarity peak lands on the correct fine cell. Two things work against that:
1521
+ the peak is not sharper at 448 (the similarity field on frozen features is smooth, so the argmax moves
1522
+ around), and Recall@0.10 β€” the loosest threshold β€” degrades *most*, which means *"the fine grid is adding
1523
+ positional noise rather than positional precision."*
1524
+
1525
+ The document is careful to attribute this to the zero-shot baseline, not to RemoteCLIP: *"A **learned**
1526
+ head trained to regress boxes from these features may respond differently to resolution."*
1527
+
1528
+ ### 11.5 What the decision does not establish
1529
+
1530
+ - **Whether the zero-shot baseline is good.** It is not: mean best IoU 0.0972 and Recall@0.5 0.0234 are
1531
+ *"weak"*. It is an ablation floor for the Phase 8 head.
1532
+ - **Whether a trained head has the same resolution sensitivity.** Re-opening the question is legitimate
1533
+ **if** the head's validation curve suggests it, and would be *"a new pre-registered experiment, not a
1534
+ silent retune."*
1535
+ - **Anything about hidden ISRO/SAC imagery.** *"VRSBench is overhead optical. The hidden set is
1536
+ Cartosat-2S + RISAT, a different distribution entirely."*
1537
+
1538
+ ### 11.6 Downstream consequences of the frozen resolution
1539
+
1540
+ The frozen resolution propagates into concrete serving parameters:
1541
+
1542
+ - **The decode grid.** `build_grounding_specialist` sets `grid = resolution // 32`, so at 224 the grid is
1543
+ **7Γ—7 = 49 cells**. The head emits `(1, 49, 5)` β€” `[tx, ty, tw, th, objectness]` per cell.
1544
+ - **The projected dimension.** `grounding.encoder_projected_dim: 512` is the measured `visual.proj`
1545
+ output of RemoteCLIP ViT-B/32 (the transformer width is 768; `visual.proj` maps to 512). The
1546
+ per-cell feature is `concat([patch, text, patch*text, global_pool]) = 4 Γ— 512 = 2048`, which
1547
+ `grounding_head.feature_dim: 2048` must equal β€” enforced at load time by `core/config.py`, which
1548
+ rejects any other value, because *"a mismatch here is a SILENT shape error"* that torch raises only at
1549
+ the similarity step, *"after the patch features are already cached."*
1550
+ - **NMS.** `grounding.nms_iou: 0.50` is passed to `nms(boxes, scores, iou_threshold)`
1551
+ (`specialists/grounding/head.py:294`).
1552
+ - **The serving candidate budget.** `GroundingSpecialist` defaults `max_candidates=6`, **not** the
1553
+ config's `20`. The class docstring records why: *"The Phase 8 decision measured that a 20-box budget
1554
+ inflates mean best IoU relative to the 5.99-box baseline; 6 is the decode-matched setting."* The builder
1555
+ reads `grounding.serving_max_candidates` with a default of 6.
1556
+ - **The degenerate-box guard.** `DEGENERATE_AREA_FRACTION = 0.9` drops any zero-shot candidate covering
1557
+ more than 90% of the frame, *"the shape a failed localisation takes once it has been smoothed into a
1558
+ rectangle."* This guard applies to the **zero-shot fallback path only** β€” *"the trained-head decode is
1559
+ a regressed box with its own learned prior; filtering it would change the frozen Phase 8 benchmark
1560
+ numbers."*
1561
+ - **The decode function.** `decode_cell_relative` (`specialists/grounding/head.py:204`) turns
1562
+ `(B, N, 4)` cell-relative parameters into `(B, N, 4)` normalised xyxy, clamped to `[0, 1]` and
1563
+ order-corrected by `enforce_order` (which swaps inverted corners rather than letting IoU go silently to
1564
+ zero). The decode is *"differentiable: the output feeds the box and GIoU losses directly, so the loss is
1565
+ computed on exactly the coordinates the metric measures."*
1566
+
1567
+ ---
1568
+
1569
+ ## 12. Worked examples
1570
+
1571
+ ### 12.1 A 4-band Cartosat-2S optical scene
1572
+
1573
+ 1. **Inspect.** `inspect_raster` reads `band_count=4`, no CRS or with CRS. `infer_modality(4)` returns
1574
+ `OPTICAL` (4 is in `_OPTICAL_BAND_COUNTS`).
1575
+ 2. **Descriptor.** With no `SensorDescriptor`, `_descriptor_for` maps the 4 bands positionally onto
1576
+ `OPTICAL_CANONICAL[:4]` = `B01, B02, B03, B04` and appends a warning naming the assumption. With a
1577
+ descriptor built by `build_optical_adapter("cartosat_2s")`, the bands are `["B1","B2","B3","B4"]`
1578
+ canonicalised to `B01..B04`, and `availability = [True, True, True, True, False Γ— 8]`.
1579
+ 3. **Canonicalise.** `adapt_optical` returns a `(12, H, W)` tensor with channels 0–3 carrying the real
1580
+ bands and channels 4–11 exactly zero, plus a 12-element mask with four `True`.
1581
+ 4. **Resize.** `resize_to_canonical` bilinearly resizes to `120Γ—120`. The zero channels stay zero.
1582
+ 5. **Encode.** `encoder.encode(optical_tensor[None], sar_tensor[None])` β€” **no mask** (C-1).
1583
+ 6. **Assemble.** `assemble_fusion_input` concatenates the three 768-d GAPs with the 12-element optical
1584
+ mask and the 2-element SAR mask β†’ `(1, 2318)`.
1585
+ 7. **Head.** With a trained head, the classifier emits 19 logits β†’ softmax β†’ a label and a margin. With
1586
+ no head, `probabilities` stays `None`, `degraded=True`, and no label is invented.
1587
+ 8. **Evidence.** An `AVAILABILITY_MASK` item per modality records the sensor, the 4 real bands, the band
1588
+ map, the mask (`[true Γ— 4, false Γ— 8]`), and `missing_channels: [4,5,6,7,8,9,10,11]`. A
1589
+ `JOINT_FEATURE_REGION` item records `fusion_input_dim: 2318`, `mask_consumed_by: "fusion_head"`,
1590
+ `croma_received_mask: false`.
1591
+ 9. **Confidence.** `optical_confidence = 4/12 = 0.3333`; the SAR side contributes its own fraction; the
1592
+ fusion margin (if a head exists) is the largest weight. If no head, the raw confidence is **0.0** with
1593
+ `no_prediction: 1.0` recorded.
1594
+
1595
+ ### 12.2 A RISAT SAR scene with no band labels
1596
+
1597
+ 1. **Descriptor.** `positional_fallback_descriptor(stem, 2)` selects `("VV", "VH")` from
1598
+ `_FALLBACK_SAR_ORDER` and records the fallback in the `sensor` field. The specialist appends: *"The
1599
+ polarisation IDENTITY of each channel is NOT known from a band count alone."*
1600
+ 2. **Canonicalise.** `adapt_sar` returns `(2, H, W)` with both slots filled, mask `[True, True]`.
1601
+ 3. **Trace.** `available_bands = ["VV", "VH"]` is the **assumption**, not a measurement. A caller with
1602
+ real polarisation labels passes `available_bands=["HH","HV"]`, in which case the descriptor's own order
1603
+ becomes the slot order and the mask reflects it β€” no relabelling, no discarded data.
1604
+
1605
+ ### 12.3 A change pair with differing dimensions
1606
+
1607
+ 1. **Validate.** Exactly two assets, distinct paths/SHA-256, CRS compared.
1608
+ 2. **Register.** `measure_registration` collapses both to grayscale float32 and runs `phaseCorrelate`.
1609
+ 3. **Shape check.** If the two images differ in size, `STANetStyleChangeDetector.forward` raises
1610
+ `SpecialistError("T1 and T2 must have the same shape; got ...")` β€” the 726Β²/736Β² case (Β§9.3). No resize
1611
+ step exists to reconcile them.
1612
+ 4. **If shapes match but registration is poor:** the maps are still computed, the regions are withheld,
1613
+ the transform is dropped from the `geospatial` block, and the confidence is floored at 0.0 with
1614
+ `suppressed_by_registration: 1.0` recorded.
1615
+
1616
+ ---
1617
+
1618
+ ## 13. What is NOT implemented β€” exhaustive
1619
+
1620
+ Each item below was checked against the code. Where a capability is absent, it is named rather than
1621
+ implied; where the absence cannot be confirmed from the files read, it is marked `UNKNOWN`.
1622
+
1623
+ | Capability | Status | Evidence |
1624
+ |---|---|---|
1625
+ | **Reprojection pipeline** | **NOT IMPLEMENTED** | `geospatial/crs.py::compare_crs` *reports* `requires_reprojection`; no code performs a reprojection. `preprocessing/imagery.py` explicitly does not reproject. |
1626
+ | **Cloud masking** | **NOT IMPLEMENTED** | No cloud, cirrus, QA or `SCL` handling anywhere in `preprocessing/` or the specialists. `UNKNOWN β€” not established from the available evidence` whether any dataset adapter applies a cloud mask. |
1627
+ | **Atmospheric correction** | **NOT IMPLEMENTED** | No atmospheric, surface-reflectance or `L2A`-processing code. Rasters are consumed as delivered. |
1628
+ | **Mosaicking** | **NOT IMPLEMENTED** | No mosaicking, seamline or composite code exists. |
1629
+ | **Tiling / top-K tile selection** | **DECLARED (not read)** | `image.tile_size`, `image.tile_overlap`, `image.max_tiles`, `image.top_k_tiles` are read only by `core/config.py`'s `top_k_tiles <= max_tiles` check. No serving-path tiling. |
1630
+ | **Optical percentile 2/98 conditioning** | **DECLARED (not read)** | `optical.normalization`, `optical.lower_percentile`, `optical.upper_percentile` are read by no code (`radiometry.py` docstring; PHASE14 Β§1). |
1631
+ | **SAR dB clip βˆ’30/+5 conditioning** | **DECLARED (not read)** | `sar.representation`, `sar.clip_min_db`, `sar.clip_max_db` are read by no code. |
1632
+ | **Resampling / cropping on the display path** | **DELIBERATELY ABSENT** | `preprocessing/imagery.py` docstring: *"It does not resample, crop, or reproject."* |
1633
+ | **Resize to reconcile differing pair dimensions** | **NOT IMPLEMENTED** | The change path requires equal shapes and has no reconciling resize (Β§9.3). |
1634
+ | **Nodata filling** | **NOT IMPLEMENTED** | `quality.py` refuses any non-finite input outright; its comment records: *"Filling nodata from the raster profile belongs in the tiling work."* |
1635
+ | **Area computation for geographic CRS** | **NOT IMPLEMENTED** | `pixel_area_m2` returns `None` for a non-metric CRS rather than approximating. |
1636
+ | **GeoJSON / WKT / KML output** | **NOT IMPLEMENTED** | Coordinates are emitted as lists; no geometry serialisation exists. |
1637
+ | **Multi-band raster β†’ RGB band selection policy** | **PARTIAL** | `load_image_array` takes the first three bands as RGB (or repeats a single band); there is no configurable band combination. |
1638
+ | **Sensor adapter for sensors beyond the known list** | **REFUSED BY DESIGN** | An unknown optical sensor with no explicit band list raises `UnsupportedBandsError`; `_known_sar` returns `[]` for an unknown SAR sensor. |
1639
+
1640
+ **A note on the "not implemented" list.** None of these omissions is hidden. The two most consequential β€”
1641
+ the tiling policy and the radiometric conditioning β€” are declared in the frozen config and read by no
1642
+ code, and both the config comments and the DEV-2 ruling name that state explicitly rather than letting the
1643
+ key read as a satisfied requirement.
1644
+
1645
+ ---
1646
+
1647
+ ## 14. What is NOT RUN / OPEN / BLOCKED for this topic
1648
+
1649
+ ### NOT RUN
1650
+
1651
+ - **The optical-SAR forward pass at serving time with a trained head.** `CROMA_base.pt` is not present
1652
+ locally and `use_croma.py` must be vendored; the specialist runs sensor-only and produces no label.
1653
+ - **A system-level geospatial accuracy benchmark.** No end-to-end benchmark exists; no system-level
1654
+ accuracy is claimed anywhere.
1655
+ - **Reprojection of a genuinely mismatched-CRS pair.** The code path that would consume a reprojection is
1656
+ absent, so the case has never been exercised.
1657
+ - **Change on a bundled pair with equal dimensions.** The only bundled demo pair is 726Β² vs 736Β² and
1658
+ errors; a same-shape pair has not been run as a demo.
1659
+
1660
+ ### OPEN
1661
+
1662
+ - **The `optical.*` / `sar.*` conditioning stage.** `OPEN` β€” either it gets an implementation or the keys
1663
+ are retired (`docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§6 item 3).
1664
+ - **The arm-constructibility question.** The Β§4-registered arm B (*"percentile/dB only, no encoder-input
1665
+ stretch"*) is **not constructible** while stage 1 is unimplemented; the arm now named B corresponds to
1666
+ the registered arm C. The ruling records this as *"a human decision and deliberately not resolved."*
1667
+ - **The optical-SAR ruling.** Accuracy **0.931** with macro-F1 **0.434161**; ruling **OPEN**. Never quote
1668
+ the accuracy without the macro-F1.
1669
+ - **The grounding resolution question after a trained head.** Re-opening 224 is legitimate only as a new
1670
+ pre-registered experiment.
1671
+ - **The change-VQA metric ruling.** `OPEN` and owner-gated.
1672
+ - **Whether any call site builds a `Box` outside the two conversion functions** (Β§5.4) β€” `UNKNOWN`.
1673
+
1674
+ ### BLOCKED
1675
+
1676
+ - **A clean change demo.** Blocked on a same-shape pair or a resize step.
1677
+ - **The CROMA normalisation experiment (option c).** Blocked on arm B being non-constructible and on the
1678
+ Arm-B feature cache, which does not exist (`artifacts/optical_sar/fusion_features/` holds only the
1679
+ Arm-A cache).
1680
+ - **The gated experiment's sample size.** Cannot be specified until a variance-estimation pass is run;
1681
+ the ruling deliberately declines to invent a number.
1682
+
1683
+ ### REJECTED
1684
+
1685
+ - **448 px grounding resolution.** `REJECTED` by the pre-registered rule and the paired test.
1686
+ - **Option (b) β€” match `base.yaml` alone, accept the shift.** `REJECTED` as a standalone path
1687
+ (`docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§2.1).
1688
+ - **Reusing `preprocessing/imagery.py`'s stretch for CROMA.** `REJECTED` as `forbidden` (Β§3.3 of the
1689
+ ruling).
1690
+ - **Coherence- and texture-based registration gates.** One `REJECTED`, one `CONFIRMED but INSUFFICIENT`.
1691
+
1692
+ ---
1693
+
1694
+ ## 15. Where the evidence lives
1695
+
1696
+ **Source modules (the code this chapter describes):**
1697
+
1698
+ | Path | What it defines |
1699
+ |---|---|
1700
+ | `preprocessing/raster.py` | The validation chain, `inspect_raster`, `read_bands`, `write_raster`, `infer_modality`, `file_sha256`, `pixel_area_m2` |
1701
+ | `preprocessing/imagery.py` | `load_image_array`, `load_pil_image`, the per-scene 2/98 display stretch |
1702
+ | `preprocessing/quality.py` | The deterministic quality gate, `MIN_AUTOCORRELATION=0.10`, `NOISE_ENTROPY_BITS=7.8` |
1703
+ | `geospatial/crs.py` | `parse_crs`, `describe_crs`, `compare_crs`, `is_metric`, `CRSCompatibility` |
1704
+ | `geospatial/transform.py` | `PixelWindow`, `GeoBounds`, `to_normalized`/`to_pixel`/`to_geo`, `intersect`, `iou`, `benchmark_boxes_to_normalized`, `affine_from_metadata` |
1705
+ | `core/schemas.py` | `CoordinateSystem`, `GeoMetadata`, `SensorDescriptor`, `AssetMetadata`, `Box`, `Region`, `ChangeRegion`, `Evidence._spatial_needs_crs` |
1706
+ | `core/config.py` | Config loading, validation, `Config.hash`, `device_preference` |
1707
+ | `specialists/optical_sar/sensor_adapter.py` | `OPTICAL_CANONICAL`, `SAR_CANONICAL`, `build_optical_adapter`, `describe_sar`, `_apply_canonical`, `positional_fallback_descriptor` |
1708
+ | `specialists/optical_sar/radiometry.py` | `TRANSFORM_NAME`, `normalise_for_croma`, `resolve_use_8_bit`, the zero-channel rule |
1709
+ | `specialists/optical_sar/fusion_head.py` | `expected_fusion_dim`, `assemble_fusion_input`, `channel_dropout`, `build_fusion_head` |
1710
+ | `specialists/optical_sar/inference.py` | `run_pipeline`, `resize_to_canonical`, `cross_modal_agreement`, `build_head_from_config` |
1711
+ | `specialists/optical_sar/specialist.py` | `_assess_pair`, `_descriptor_for`, `_confidence_components`, `_geospatial` |
1712
+ | `specialists/change/postprocess.py` | `RegistrationQuality`, `measure_registration`, `to_grayscale_float`, `regions_to_schema` |
1713
+ | `specialists/change/specialist.py` | `_assess_pair`, `_predict`, `_geospatial`, the suppression path |
1714
+ | `specialists/change/stanet.py:415` | The equal-shape guard |
1715
+ | `specialists/grounding/specialist.py` | `_to_geo_boxes`, the `NORMALIZED_0_1` box construction, `DEGENERATE_AREA_FRACTION` |
1716
+ | `specialists/grounding/head.py` | `decode_cell_relative`, `enforce_order`, `nms` |
1717
+ | `specialists/grounding/inference.py` | `decode_candidates_from_features`, `ground_phrase` |
1718
+ | `evaluation/metrics/grounding.py` | `benchmark_to_normalized(box, scale=100.0)` |
1719
+
1720
+ **Configuration:**
1721
+
1722
+ - `configs/base.yaml` β€” `image.*` (tiling), `optical.*`, `sar.*`, `croma.*`, `fusion.*`, `grounding.*`,
1723
+ `grounding_head.*`, `change.*`.
1724
+
1725
+ **Documents:**
1726
+
1727
+ - `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` β€” the DEV-2 ruling, the two ordered stages, the Gate F arm
1728
+ amendment.
1729
+ - `docs/CROMA_NORMALISATION_UPSTREAM_EVIDENCE.md` β€” what is and is not established about CROMA's
1730
+ pretraining distribution.
1731
+ - `docs/PHASE7_RESOLUTION_DECISION.md` β€” the 224/448 decision in full.
1732
+ - `docs/CHANGE_REGISTRATION_GATE_COHERENCE_TEST.md`, `docs/CHANGE_REGISTRATION_GATE_TEXTURE_TEST.md` β€” the
1733
+ rejected and insufficient registration gates.
1734
+ - `docs/API_CONTRACT.md` Β§2.5 (asset allowlist), Β§3.2 (coordinate systems).
1735
+ - `docs/FINAL_DELIVERY_REPORT.md` Β§6 β€” the 726Β²/736Β² change demo recorded as DEGRADED.
1736
+ - `docs/PHASE10_CDVQA_DATA_STATUS.md` β€” the CDVQA/SECOND pair layout and temporal semantics.
1737
+ - `docs/PHASE11_CROMA_STATUS.md` β€” the reproduced CROMA forward pass.
1738
+
1739
+ **Sibling release chapters:**
1740
+
1741
+ - [`DATA_PIPELINE.md`](DATA_PIPELINE.md) β€” the end-to-end data path, caches and artefacts.
1742
+ - [`architecture/03-request-lifecycle.md`](architecture/03-request-lifecycle.md) β€” the nine controller
1743
+ states, including VALIDATE and PREPROCESS.
1744
+ - [`architecture/05-specialists.md`](architecture/05-specialists.md) β€” the six specialists.
1745
+ - [`architecture/07-configuration-freeze.md`](architecture/07-configuration-freeze.md) β€” the frozen hash
1746
+ and every config key.
1747
+ - [`DATASETS.md`](DATASETS.md) β€” the corpora and their measured shapes.