Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Geospatial β deep reference
Status tags used on every substantive claim: IMPLEMENTED Β· VERIFIED Β· MEASURED Β· ATTEMPTED
Β· NOT RUN Β· BLOCKED Β· DEFERRED Β· REJECTED Β· OPEN Β· RESOLVED Β· CLOSED.
Every statement in this document was produced by reading a file in C:/Users/anish/satquery-ai/. Where
a fact is not established from the available evidence, this document writes
UNKNOWN β not established from the available evidence. Where a transform is declared in configuration
but read by no code, that is stated explicitly rather than presented as a working feature. Where a
capability is absent, it is listed in Β§13 rather than implied.
This chapter is the reference for the spatial contract: how a raster becomes an analysed image, what georeferencing survives, what coordinate system a box is in, how a sensor's bands become CROMA's canonical channels, and which of those steps are actually implemented on the serving path. It is written to be sufficient to reconstruct the subsystem from the document alone.
Sibling documents: DATA_PIPELINE.md (the end-to-end data path, caches and
artefacts), architecture/03-request-lifecycle.md (where the
PREPROCESS and VALIDATE states sit), architecture/05-specialists.md
(the six specialists in full), architecture/07-configuration-freeze.md
(the frozen config identity and every key), DATASETS.md (the corpora and their measured
shapes).
Table of contents
- How to read this document
- The raster input contract
- CRS handling
- Affine transforms and coordinate frames
- The coordinate-system contract
- Benchmark box scale β 0β100 versus 0β1
- Band-count modality inference
- Radiometric normalisation and representation
- The CROMA sensor adapter contract
- Pair-compatibility rules
- Large-image tiling policy
- Grounding resolution β the frozen 224 decision
- Worked examples
- What is NOT implemented β exhaustive
- What is NOT RUN / OPEN / BLOCKED for this topic
- Where the evidence lives
0. How to read this document
Three distinctions are load-bearing throughout, and conflating them is the single most common way to misread this system:
- Declared versus executed.
configs/base.yamlis the authoritative registry, andcore/config.pyvalidates and hashes it at load time. But declaring a policy and executing it are different claims. The tiling policy, the optical percentile stretch and the SAR dB clip are all declared inconfigs/base.yamland are all read by no code on the serving path. This document says so at each point rather than presenting the key as a feature. - Adapter versus encoder. The CROMA sensor adapter (
specialists/optical_sar/sensor_adapter.py) performs band mapping, zero-fill and availability masking. CROMA the encoder does not receive the availability mask (finding C-1). The mask is consumed by the fusion head. These are separate contracts and Β§8.9 treats the distinction in full. - Presentation versus radiometry. There are two different percentile stretches in this codebase:
a per-scene 2/98 display stretch in
preprocessing/imagery.py(and a grayscale cousin inspecialists/change/postprocess.py), and a per-channelmean Β± 2Β·stdencoder-input stretch inspecialists/optical_sar/radiometry.py. They serve different purposes and are not interchangeable. Β§7.5 is explicit about why.
The status vocabulary is used on every substantive claim. MEASURED means a number was produced by
executing code; VERIFIED means the claim was checked against a second independent source; IMPLEMENTED
means the code exists on the serving path; DECLARED (not read) means the value lives in
configs/base.yaml and no code reads it.
1. The raster input contract
1.1 What arrives, and what is accepted
An asset reaches the analysis path through POST /v1/assets, which mints an opaque asset_id. The
gateway enforces a closed content-type allowlist of exactly five types
(docs/API_CONTRACT.md Β§2.5):
| Content type | Notes |
|---|---|
image/tiff |
The type the geospatial specialists need. A client that uploads only PNG/JPEG can serve vqa, caption and grounding, but not change or optical_sar |
image/geotiff |
Accepted; media-type parameters are ignored, so image/tiff; charset=binary is accepted |
image/png |
Accepted |
image/jpeg |
Accepted |
application/octet-stream |
Accepted, but the raster reader still has to be able to open the bytes |
Two behaviours are specified rather than defaulted: a request with no declared content type is
refused rather than defaulted (defaulting is how a PDF reaches a raster reader), and a disallowed
type gets 415. The stored filename is derived from the content type, never from the client's
filename, so a client-supplied ../../ cannot influence where bytes land (docs/API_CONTRACT.md Β§2.5).
The size cap is enforced at two layers β the gateway and the Space β and both now refuse an unparsable
or non-positive value rather than silently falling back to a default
(docs/API_CONTRACT.md Β§2.5, "Size cap").
artifact_ref on every Evidence item is always null in v1 (core/schemas.py:225): v1 exposes no
artifact-serving endpoint, so a reference a client cannot retrieve is not fabricated. The schema field's
own description states this β "Reference to an externally retrievable artifact. NEVER a filesystem path
(F-16, owner ruling 2026-09-23)" β and the change and optical-SAR specialists both emit null on every
path (Β§4.3).
1.2 The validation chain
preprocessing/raster.py is the first thing every uploaded asset meets. Its module docstring names the
chain it implements, verbatim:
file -> dimensions -> bands -> dtype -> CRS -> transform -> bounds -> nodata
-> modality -> temporal metadata
The module's three design rules are stated at the top of the file and are worth quoting because they govern every failure mode below:
- Never raise a bare exception. Every failure is a typed SatQueryError.
- Never silently drop geospatial metadata. If the source had a CRS, the returned AssetMetadata says so.
- A missing CRS degrades to non-geospatial mode; it does not abort.
1.3 inspect_raster β header-only validation
inspect_raster (preprocessing/raster.py:96) validates and describes a raster without loading its
pixel data. Its exact signature:
def inspect_raster(
path: str | Path,
*,
max_pixels: int | None = None,
explicit_modality: str | None = None,
compute_hash: bool = False,
sensor: SensorDescriptor | None = None,
) -> AssetMetadata:
The behaviour, field by field:
- The path is checked with
Path(path).exists()andis_file(); a missing path raisesRasterReadError. rasteriois imported lazily; anImportErrorraisesRasterReadErrorrather than a bare error.- Inside
rasterio.open(p), the reader extractswidth,height,count, the sorted set ofdtypes, the CRS as a string (ds.crs.to_string() if ds.crs else None), the affine transform truncated to its first six coefficients (list(ds.transform)[:6] if ds.transform else None), the bounds (list(ds.bounds)),nodata,resolution(list(ds.res)), thedriverstring, and a tiled flag read through_safe_is_tiled. _safe_is_tiled(preprocessing/raster.py:42) exists becauseDatasetReader.is_tiledis scheduled for removal; it suppresses the pending deprecation and degrades toFalserather than aborting, since the flag is informational metadata only.- Degenerate dimensions (
width <= 0 or height <= 0) raiseRasterReadError. - If
max_pixelsis given andwidth * height > max_pixels, it raisesOversizedImageErrorβ and the docstring records that this is recoverable: "the caller may downscale". - A band count
<= 0raisesUnsupportedBandsError. - Anything else rasterio raises is wrapped:
RasterReadError(f"could not open raster '{p}': {exc}").
The return is an AssetMetadata (see Β§1.4) with modality from infer_modality (Β§6) and, when
compute_hash=True, the streaming SHA-256 of the file (file_sha256, preprocessing/raster.py:82,
chunked at 1 MiB).
inspect_raster is called from core/controller.py::_inspect_asset during the VALIDATE state, with the
configured pixel budget passed through. That means the header-only nature of the inspection is a real
property of the serving path: a request never decodes a raster just to learn its shape.
1.4 AssetMetadata and GeoMetadata
GeoMetadata (core/schemas.py:120) is the geospatial detail carrier. It is extra="allow", which is
deliberate β the optical-SAR specialist adds cross-sensor keys to it (Β§4.3):
| Field | Type | Meaning |
|---|---|---|
crs |
str | None |
CRS as a string (e.g. EPSG:4326), or None |
transform |
list[float] | None |
Affine coefficients, row-major, six values |
bounds |
list[float] | None |
[minx, miny, maxx, maxy] |
width, height |
int | None |
Raster dimensions in pixels |
band_count |
int | None |
Number of bands |
dtype |
str | None |
Comma-joined distinct dtypes when the bands disagree |
nodata |
float | None |
The nodata sentinel, or None |
resolution |
list[float] | None |
(xres, yres) |
has_crs |
bool |
crs_value is not None |
is_georeferenced |
bool |
bool(crs_value and transform) |
AssetMetadata (core/schemas.py:153) is extra="forbid" and carries:
| Field | Type | Meaning |
|---|---|---|
asset_id |
str |
Minted asset_<12 hex>, opaque |
path |
str |
Resolved local path |
modality |
Modality |
From infer_modality, or declared |
sha256 |
str | None |
Present only when compute_hash=True |
geo |
GeoMetadata |
The block above |
sensor |
SensorDescriptor | None |
Present only when a sensor descriptor was supplied |
acquisition_date |
str | None |
Temporal metadata |
scene_id, dataset_id, split |
str | None |
Provenance for leakage control |
Note the distinction between has_crs and is_georeferenced: a raster can carry a CRS but no affine
transform (or vice versa), and both are recorded. Only when both are present is
is_georeferenced=True. geospatial/transform.py::to_geo requires the transform, so a raster with
has_crs=True but transform=None cannot be converted to geographic coordinates β it will raise
CoordinateError("cannot convert to geo: raster has no affine transform").
1.5 Band reading, writing, and pixel area
read_bands (preprocessing/raster.py:188) returns (bands, H, W) plus a profile that retains
crs, transform and bounds as plain values, so downstream code never loses georeferencing. The
profile is a copy of ds.profile with those three keys overwritten from the live dataset.
write_raster (preprocessing/raster.py:221) writes an array using a profile captured by read_bands,
and re-applies crs and transform explicitly in the rasterio.open(..., "w", ...) call. It is used by
evidence generation so change maps stay georeferenced β except when the pair is suppressed for poor
registration, in which case the transform is deliberately dropped (Β§9.1).
pixel_area_m2 (preprocessing/raster.py:259) returns the area of one pixel in square metres when the
CRS is metric, and None otherwise:
def pixel_area_m2(meta: AssetMetadata) -> float | None:
from geospatial.crs import is_metric
if not is_metric(meta.geo.crs):
return None
if not meta.geo.resolution or len(meta.geo.resolution) < 2:
return None
xres, yres = meta.geo.resolution[0], meta.geo.resolution[1]
return abs(xres * yres)
The docstring is explicit: "Returns None for geographic CRS or unknown resolution β callers must not guess." A degree-based CRS has no meaningful pixel area in metres, so the function refuses rather than returning a plausible-looking number.
2. CRS handling (geospatial/crs.py)
CRS comparison is kept separate from pixel arithmetic so that "can these two rasters be compared at
all?" is answerable without touching pixels (geospatial/crs.py docstring). The module depends on
pyproj.
2.1 parse_crs
def parse_crs(value: str | None) -> CRS | None:
if value is None or value == "":
return None
try:
return CRS.from_user_input(value)
except PyProjCRSError as exc:
raise MissingCRSError(f"unparseable CRS '{value}': {exc}") from exc
Absent is None; malformed raises. That distinction matters: a missing CRS is a legitimate state
(degrades to non-geospatial), while a malformed one is a defect that must surface.
2.2 describe_crs
describe_crs returns a dict for the execution trace: present, srs, name, is_geographic,
is_projected, and axis_units (the first axis's unit_name). For an absent CRS it returns
{"present": False}.
2.3 compare_crs and CRSCompatibility
CRSCompatibility (geospatial/crs.py:19) is a frozen dataclass with compatible, identical,
same_units, left, right, reason, and a derived property:
@property
def requires_reprojection(self) -> bool:
return self.compatible and not self.identical
compare_crs (geospatial/crs.py:59) decides:
- A missing CRS on either side is NOT compatible. The reason string is
"one or both rasters lack a CRS; spatial comparison is unsafe". The docstring records that this is "not fatal for non-spatial tasks". - Identical CRS (
crs_l.equals(crs_r)) returnsidentical=True, reason"identical CRS". - Same kind (both projected or both geographic) returns
compatible=True, identical=False, reason"reprojection required (<left name> -> <right name>)". - Mixed projected/geographic returns
compatible=True, identical=False, with a reason that adds "reprojection required and should be verified". The comment records the rationale: "Mixing a projected CRS with a geographic one is legal to reproject but signals a pipeline mistake."
same_units compares the first axis's unit_name on each side (defaulting to "unknown" when there is
no axis info).
Crucially: compare_crs never reprojects. It reports whether reprojection is required; the actual
transform is not implemented anywhere (Β§13).
2.4 is_metric
def is_metric(crs_value: str | None) -> bool:
crs = parse_crs(crs_value)
if crs is None:
return False
if not crs.is_projected:
return False
if not crs.axis_info:
return False
return crs.axis_info[0].unit_name in {"metre", "meter", "m"}
A geographic CRS is never metric here, and the unit check is a closed set of three spellings.
2.5 Where CRS comparison is used
The only consumer on the serving path is the change specialist's _assess_pair
(specialists/change/specialist.py:251):
compat = compare_crs(t1.geo.crs, t2.geo.crs)
if not compat.compatible:
warnings.append(
f"CRS comparison unavailable ({compat.reason}); change is "
f"reported in normalized pixel coordinates only"
)
elif compat.requires_reprojection:
warnings.append(
f"T1 and T2 use different CRS ({compat.reason}); change is "
f"valid only if the rasters already share a pixel grid"
)
The comment above it states the policy precisely: "No CRS on either side is NOT fatal for pixel-domain
change detection β the two rasters still tile to the same grid. It IS fatal for any claim about ground
coordinates, so the warning is recorded and the geospatial block is withheld downstream." The
withholding happens in ChangeSpecialist._geospatial, which drops the transform when registration is
unusable (Β§9.1).
3. Affine transforms and coordinate frames
geospatial/transform.py is where a normalised 0β1 box, a pixel box and a geographic box are converted
between. Its module docstring states the non-negotiable rules (citing docs/ARCHITECTURE_FREEZE.md
Β§2.6):
- CRS and affine transform are preserved, never silently stripped.
- Conversions are explicit about their source and target frames.
- A conversion that cannot be performed raises, rather than returning a plausible lie.
That last line is the module's governing principle, and it is why to_normalized, to_pixel and
to_geo all raise CoordinateError when they lack the context a conversion needs.
3.1 PixelWindow and GeoBounds
PixelWindow (geospatial/transform.py:29) is a frozen dataclass of [col_min, row_min, col_max, row_max]. Its __post_init__ raises on an inverted window:
if self.col_max < self.col_min or self.row_max < self.row_min:
raise CoordinateError(f"pixel window is inverted: {self.as_list()}")
It exposes width, height and area properties. GeoBounds (geospatial/transform.py:60) is the
geographic cousin: (minx, miny, maxx, maxy).
3.2 Normalised β pixel
normalized_to_pixel(box, width, height) multiplies each coordinate by the corresponding dimension. Its
docstring records a deliberate choice:
Values are NOT clipped: a box slightly outside the frame is preserved so the caller can decide whether that is a bug or a legitimate edge case.
It raises CoordinateError if len(box) != 4 or if either dimension is <= 0.
pixel_to_normalized is the inverse, dividing by the dimensions.
3.3 Pixel β geographic
pixel_to_geo(window, transform) applies the affine transform to two corners. The docstring explains the
axis convention:
rasterio/affine transforms map (col, row) -> (x, y). The y axis usually points down in pixel space, so the geographic miny comes from the BOTTOM row.
It uses the non-deprecated matmul form transform @ (col, row) and then takes min/max across the two
corners so the returned GeoBounds is ordered. It raises CoordinateError if the transform is None.
geo_to_pixel(bounds, transform) inverts the transform with ~transform, raising CoordinateError if
the inverse cannot be computed ("affine transform is not invertible: ...").
3.4 Schema-aware conversions
to_normalized, to_pixel and to_geo operate on Box objects and are schema-aware β they read
box.coordinate_system and return a Box with the target system set. The context each requires is
documented in the docstring:
pixel -> normalizedrequireswidthandheight.geo -> normalizedrequireswidth,heightand the affinetransform.
Each returns early when the box is already in the target frame, and otherwise uses
box.model_copy(update={...}) to produce a new box with the new coordinates and the new
coordinate_system.
3.5 Intersection and IoU
intersect(a, b) returns the axis-aligned intersection of two PixelWindows, or None when they are
disjoint. iou(a, b) computes intersection-over-union in [0, 1], returning 0.0 for a disjoint pair
or a zero-area union. These are used by the grounding NMS path and by change-region overlap.
3.6 affine_from_metadata
def affine_from_metadata(geo: GeoMetadata) -> Affine | None:
if not geo.transform or len(geo.transform) != 6:
return None
a, b, c, d, e, f = geo.transform
return Affine(a, b, c, d, e, f)
A GeoMetadata whose transform is missing or not exactly six values yields None; callers treat that as
"cannot georeference" rather than as an error. This is the bridge the grounding specialist uses in
_to_geo_boxes (Β§4.3).
4. The coordinate-system contract
4.1 The enum
CoordinateSystem (core/schemas.py:57) has exactly three members, and the class docstring is a warning
in itself β "Never omit this. A bare box is meaningless without it. (C-5)":
class CoordinateSystem(str, Enum):
NORMALIZED_0_1 = "normalized_0_1"
PIXEL = "pixel"
GEO = "geo"
docs/API_CONTRACT.md Β§3.2 fixes the meaning of each value:
| Value | Meaning | Rendering |
|---|---|---|
normalized_0_1 |
[0, 1], origin top-left |
Multiply by image width/height |
pixel |
Absolute pixel coordinates | Use directly |
geo |
CRS coordinates (usually EPSG:4326) | Requires a map, not a 2-D canvas |
The contract adds a hard warning: "The frontend must read the coordinate_system field on each
Box/Region and must not assume one convention. A box drawn with the wrong assumption lands in
plausible-looking wrong places." It also records that earlier drafts used shorthand (normalized,
geographic) and that those strings are wrong β they would fail schema validation, because the
models are extra="forbid" and the enum is closed.
4.2 Why it is mandatory β the schema validator
The Box model (core/schemas.py:171) defaults coordinate_system to
CoordinateSystem.NORMALIZED_0_1, and its _ordered validator rejects a box whose coordinates do not
satisfy x2 >= x1 and y2 >= y1. Region (core/schemas.py:189) and ChangeRegion
(core/schemas.py:202) both carry the same field with the same default.
The enforcement that matters is on Evidence. Evidence._spatial_needs_crs
(core/schemas.py:238) rejects a spatial evidence item that carries coordinates but no
coordinate system:
@model_validator(mode="after")
def _spatial_needs_crs(self) -> "Evidence":
spatial = {
EvidenceType.BOUNDING_BOX,
EvidenceType.MASK,
EvidenceType.CHANGE_MAP,
EvidenceType.TILE,
EvidenceType.IMAGE_CROP,
EvidenceType.JOINT_FEATURE_REGION,
}
if self.type in spatial and self.coordinates and self.coordinate_system is None:
raise ValueError(
f"evidence type '{self.type.value}' carries coordinates "
"but no coordinate_system"
)
return self
So a bare box cannot be published as evidence. The validator fires only when the item both is of a
spatial type and carries coordinates β a STATISTIC or GEOLOCATION item with no coordinates is
unaffected, and a spatial item with no coordinates (e.g. an availability-mask record) is also unaffected.
GEOLOCATION is deliberately not in the spatial set: a geolocation is a point reference, not a
rectangle, and the grounding specialist sets its coordinate_system to GEO explicitly anyway.
Note the interaction with extra="forbid": Evidence forbids extra keys, so an author cannot smuggle a
convention through an undeclared field. GeoMetadata, by contrast, is extra="allow", which is how the
optical-SAR specialist attaches cross-sensor keys (Β§4.3).
4.3 What each specialist emits
| Specialist | Boxes / regions | Evidence coordinate systems |
|---|---|---|
grounding |
Box and Region in NORMALIZED_0_1 |
BOUNDING_BOX in NORMALIZED_0_1; a GEOLOCATION item in GEO when the CRS allows; STATISTIC items with no coordinates |
change |
ChangeRegion and Region in NORMALIZED_0_1 |
CHANGE_MAP and BOUNDING_BOX in NORMALIZED_0_1; STATISTIC items with no coordinates |
optical_sar |
No regions (empty list) |
AVAILABILITY_MASK items (no coordinates); OPTICAL_VIEW / SAR_VIEW with coordinates=[0,0,1,1] in NORMALIZED_0_1; a JOINT_FEATURE_REGION in NORMALIZED_0_1; STATISTIC items |
Grounding. GroundingSpecialist.execute (specialists/grounding/specialist.py:253) builds every box
with an explicit system:
boxes = [
Box(
x1=d.box[0], y1=d.box[1], x2=d.box[2], y2=d.box[3],
score=d.score, label=request.query,
coordinate_system=CoordinateSystem.NORMALIZED_0_1,
)
for d in decoded
]
Geographic boxes are derived, never replacements: _to_geo_boxes (specialists/grounding/specialist.py:431)
converts the normalised boxes to geographic bounds only when the raster has a CRS, a transform and
dimensions, and emits a GEOLOCATION evidence item with coordinate_system=CoordinateSystem.GEO. When
the raster has no CRS/transform, the method appends a warning β
"raster has no CRS/transform; boxes are reported in normalized coordinates only and cannot be placed on a
map" β and returns an empty list. The result's boxes stay normalised; the geo boxes are additional
evidence for a consumer with a CRS. Its own comment states the rule: "Pixel and geo coordinates are
DERIVED, never replacements."
Change. regions_to_schema (specialists/change/postprocess.py:443) builds ChangeRegion objects in
the requested system. It supports NORMALIZED_0_1 and PIXEL and raises for anything else:
else:
raise ValueError(
f"regions_to_schema cannot produce {coordinate_system.value} "
f"coordinates; convert after this call"
)
The change specialist calls it with the default (NORMALIZED_0_1). Geographic conversion is a deliberate
post-step β the docstring records: "Callers that want geographic coordinates convert afterwards via
geospatial.transform, which keeps this function free of rasterio and CRS concerns."
Optical-SAR. The specialist emits no regions. Its view evidence items carry
coordinates=[0.0, 0.0, 1.0, 1.0] in NORMALIZED_0_1 β the whole frame β because the view is a
presentation of the canonical tensor, not a localisation. Its geospatial block is the optical
asset's georeferencing (chosen because it is the asset a reader can navigate by), with cross-sensor keys
added via GeoMetadata's extra="allow": optical_crs, sar_crs, sar_bounds, sar_resolution,
optical_availability, sar_availability, optical_sensor, sar_sensor, and
resolution_ratio_sar_to_optical when both resolutions are known
(specialists/optical_sar/specialist.py:597). The comment records the reason the two footprints are
carried separately rather than fused: "RISAT and Cartosat-2S are different missions with different
orbits."
5. Benchmark box scale β 0β100 versus 0β1
This is a small contract with an outsized failure mode, and the codebase treats it as such.
5.1 The two conventions
- VRSBench (the grounding benchmark) states verbatim that "all box coordinates are normalized to 0-100".
- This project stores every box normalised to 0β1.
Treating a 0β100 box as a pixel box, or as a 0β1 box, is a silent, catastrophic bug: the numbers are all in range, the schema accepts them, and the box lands in the wrong place.
5.2 The conversion functions
geospatial/transform.py:286 provides the conversion, and its docstring names the failure mode it exists
to prevent:
def benchmark_boxes_to_normalized(
boxes: Iterable[Sequence[float]], scale: float = 100.0
) -> list[list[float]]:
"""Convert benchmark boxes (e.g. VRSBench's 0-100) to our internal 0-1.
Finding from Phase 0: VRSBench states verbatim that "all box coordinates are
normalized to 0-100". Treating those numbers as pixels is a silent, catastrophic
bug β this function exists so that mistake can be made once, explicitly.
"""
if scale <= 0:
raise CoordinateError(f"box scale must be positive, got {scale}")
...
out.append([float(v) / scale for v in b])
normalized_boxes_to_benchmark is the exact inverse (multiply by scale). Both raise on a
non-positive scale and on a box that is not four values.
A second, independent conversion lives in the evaluation metrics:
evaluation/metrics/grounding.py::benchmark_to_normalized(box, scale=100.0) divides by the scale and
raises ValueError(f"scale must be positive, got {scale}") on a non-positive value. docs/API_CONTRACT.md
Β§3.2 names it explicitly as the site of the VRSBench conversion.
5.3 The declared scale
configs/base.yaml declares the conversion factor in the grounding block, with a comment that restates
the convention:
grounding:
...
# VRSBench stores boxes normalised to 0-100; we store 0-1.
benchmark_box_scale: 100.0
coordinate_system: normalized_0_1
benchmark_box_scale: 100.0 is therefore the declared value, and 100.0 is also the default
argument in both conversion functions. The two agree.
5.4 The double-application risk
Because the scale appears in three places β the config key, the geospatial.transform default and the
evaluation.metrics.grounding default β there is a real hazard that a caller converts 0β100 β 0β1 and
then a downstream stage divides by 100 again, producing a box 100Γ too small in each coordinate. The
codebase's mitigations are structural rather than documentary:
- The conversion is named for its direction.
benchmark_boxes_to_normalizedandnormalized_boxes_to_benchmarkare inverses with self-describing names, so a reader cannot easily apply the wrong one without noticing. - The default is the correct scale, not
1.0. A caller that forgets to passscalegets100.0, which is right for VRSBench and would produce an obviously wrong box (values > 1) for already-0β1 input β a loud failure, not a silent one. - The schema rejects out-of-range values indirectly. A
Boxbuilt from un-divided VRSBench coordinates would havescorein range but coordinates in[0, 100]; the_orderedvalidator would pass it (the ordering is still valid), so the schema alone does not catch this. The catch is that the metrics module's ownbenchmark_to_normalizedis the documented entry point, and the resolution experiment and the grounding evaluation both go through it.
Residual risk β stated honestly. There is no runtime assertion that a Box in
NORMALIZED_0_1 actually has coordinates within [0, 1]. Box._ordered checks only ordering. A
mis-scaled box would therefore pass schema validation. UNKNOWN β not established from the available evidence whether any production call site currently constructs a Box without going through one of the
two conversion functions; the serving path (GroundingSpecialist.execute) builds boxes from the decoder's
output, which is already in [0, 1], so the serving path is not exposed. The risk is confined to
evaluation and experiment code that reads VRSBench directly.
6. Band-count modality inference
6.1 The heuristics
preprocessing/raster.py defines two closed sets:
_OPTICAL_BAND_COUNTS = {3, 4, 8, 11, 12, 13}
_SAR_BAND_COUNTS = {1, 2}
with the comment: "Band-count heuristics for modality detection when explicit metadata is absent. These are heuristics, not ground truth β the sensor adapter is authoritative when a sensor descriptor is supplied."
6.2 infer_modality
def infer_modality(band_count: int, explicit: str | None = None) -> Modality:
if explicit:
try:
return Modality(explicit.lower())
except ValueError:
pass
if band_count in _SAR_BAND_COUNTS:
return Modality.SAR
if band_count in _OPTICAL_BAND_COUNTS:
return Modality.OPTICAL
# 12-band optical and 2-band SAR are both plausible; default to optical only
# when the count clearly favours it.
if band_count >= 4:
return Modality.OPTICAL
return Modality.UNKNOWN
The decision order is worth reading carefully:
- An explicit label wins if it parses to a
Modalitymember. An explicit label that does not parse (e.g."radar") is silently ignored and inference proceeds β theexcept ValueError: passis deliberate. {1, 2}βSAR.{3, 4, 8, 11, 12, 13}βOPTICAL.>= 4(and not already matched) βOPTICAL.- Otherwise β
UNKNOWN.
So the effective mapping is: 1β2 β SAR; 3β13 β optical (by set or by the >= 4 fallback); 0 or a
non-positive count β UNKNOWN (though inspect_raster rejects band_count <= 0 before this point).
6.3 The declared-wins rule
In the optical-SAR specialist, _assess_pair (specialists/optical_sar/specialist.py:223) uses the
asset's declared modality first, falling back to band-count inference only when it is UNKNOWN:
for asset in assets:
modality = asset.modality
if modality is Modality.UNKNOWN:
inferred = self._infer_modality(asset, warnings)
resolved.append(inferred)
else:
resolved.append(modality)
_infer_modality appends a warning that names the heuristic for what it is: "A band count is a
heuristic, not a sensor declaration β label the asset explicitly if this is wrong." The docstring above
_assess_pair states the rationale: "The declared value wins when present, because the caller may know
something the band count does not β a 2-band Cartosat stack, or a 12-band decomposition product."
6.4 Why the heuristic is not authoritative
The comment on the sets and the warning text both point at the same thing: a band count does not identify a sensor. A 2-band file could be a dual-pol SAR pair or a two-band optical stack; a 4-band file could be RGB+NIR or a four-polarisation SAR product. The adapter (Β§8) is the authoritative path, because it maps named bands onto CROMA's canonical channels and records what it could not place.
7. Radiometric normalisation and representation
This section covers a genuine specification gap and its partial resolution. It is presented in full because the gap is load-bearing and the honest state is more useful than a tidy summary.
7.1 What configuration declares
configs/base.yaml declares two radiometric transforms:
optical:
normalization: percentile
lower_percentile: 2
upper_percentile: 98
# CROMA expects exactly 12 optical channels.
canonical_channels: 12
sar:
representation: db
clip_min_db: -30
clip_max_db: 5
# CROMA expects exactly 2 SAR channels (VV, VH).
canonical_channels: 2
So the declared optical normalisation is a percentile stretch at the 2nd and 98th percentiles, and the declared SAR representation is decibels clipped to [β30, +5].
7.2 The DEV-2 ruling β two ordered stages
docs/PHASE14_CROMA_NORMALISATION_CHANGE.md records the ruling that governs this. The gap it confirmed:
CROMA's own example preprocessing is a per-channel dynamic-range stretch
(mean Β± 2Β·std β [0, 1], optionally through a uint8 round-trip), while configs/base.yaml specifies
something structurally different. The ruling's verification table records that a grep of
specialists/optical_sar/ for NORM_*, clip_min_db, lower_percentile, use_8_bit returned six
hits, all constant definitions/labels at sensor_adapter.py:125,126,210,323,547,548, and zero
application sites.
The ruling resolves this by declaring the pipeline to have two ordered stages, and it states explicitly that they are not alternatives:
| Stage | Transform | Owner | Status |
|---|---|---|---|
| Radiometric conditioning | per-scene percentile / dB-clip, per configs/base.yaml |
preprocessing / sensor adapter | specified in the ruling, must be implemented |
| Encoder-input normalisation | per-channel mean Β± 2Β·std β [0,1], then optional Γ255 uint8 β /255 |
a new function consumed by croma.py |
mandatory, immediate |
The rationale for the CROMA stage being mandatory: "it is the only stage that bounds the input range.
Without it the encoder receives unbounded input." The use_8_bit branch is the safer default because
the uint8 round-trip guarantees [0, 1], whereas the float path's clip applies to a stretch that is
unbounded by construction (mean Β± 2Β·std is not a min/max clamp). The decision on use_8_bit is
ENABLED.
7.3 The implemented stage β specialists/optical_sar/radiometry.py
The encoder-input stage is implemented, in full, at specialists/optical_sar/radiometry.py. Its
module docstring states the boundary between what is and is not implemented:
Implemented: the encoder-input stage β per-channel
mean +/- 2*std->[0, 1], the transform the CROMA authors' own README instructs users to apply.NOT implemented: the percentile / dB conditioning stage that
base.yamlnames.
The transform is quoted verbatim from the CROMA README in the docstring:
min_value = x[:, c].mean() - 2 * x[:, c].std()
max_value = x[:, c].mean() + 2 * x[:, c].std()
img = (x[:, c] - min_value) / (max_value - min_value) * 255.0
img = clip(img, 0, 255).to(uint8)
# before the forward pass:
x = x.float() / 255
The module makes two deliberate deviations from the README, both documented:
- Per sample, not per batch. The README computes
x[:, c].mean()over the whole batch, which makes one image's encoding a function of its neighbours. This module computes the window per sample, because "a serving system whose whole point is reproducibility" must not produce a different answer because another request was batched alongside it. AtN = 1the two are identical, and the README's own example runs atN = 1. - Unbiased standard deviation (
ddof=1). Upstream is torch, whose.std()defaults to the unbiased estimator; NumPy's defaults toddof=0. The module usesddof=1to match the reference implementation, and states the numerical immateriality (~1.00003 factor at 120Γ120) so nobody has to guess which convention a number came from.
The key constants:
| Constant | Value | Meaning |
|---|---|---|
TRANSFORM_NAME |
"per_channel_mean_pm_2std" |
Recorded on every trace |
WINDOW_SIGMAS |
2.0 |
The window half-width |
DEFAULT_USE_8_BIT |
True |
The ruling's decision |
ENV_USE_8_BIT |
"SATQUERY_CROMA_USE_8_BIT" |
Hash-exempt override channel |
MIN_FINITE_PIXELS |
2 |
Below this, a std is not meaningful |
STATUS_NORMALISED / STATUS_UNAVAILABLE / STATUS_DEGENERATE |
strings | Per-channel status |
normalise_for_croma(array, *, use_8_bit, modality) is pure and deterministic β the docstring
states: "no sampling, no learned statistics, no RNG, and no dependence on batch composition. Same input
-> same output." It accepts (B, C, H, W) or (C, H, W), returns a float32 array of the same shape
plus a RadiometryReport, and leaves skipped channels exactly zero.
7.4 The zero-channel rule
This is the load-bearing part, and it is the C-1 discipline applied to radiometry. An unavailable channel
is all zeros (the sensor adapter's convention). Such a channel has std = 0, so max_value - min_value == 0 and the stretch divides by zero. The module therefore skips and leaves at exactly zero any
channel with no dynamic range, and records a different status for each reason:
unavailableβ the channel is all zeros (if not np.any(channel)), i.e. the sensor did not measure it.degenerateβ the channel is present but constant (_channel_windowreturnsNonebecauseupper - lower <= 0), i.e. the sensor measured a flat band.normalisedβ the real case.
The docstring explains why the distinction matters: "'the sensor did not measure this' is not the same
fact as 'the sensor measured a flat band', and the trace must not merge them." The mechanism is the
degenerate-window test, which needs no mask β deliberately, because CROMAEncoder.encode has a
pinned signature of exactly (self, optical, sar) and finding C-1 says CROMA never receives a mask.
Inferring unavailability from the data keeps that guard intact.
A non-finite pixel (NaN/Inf) is "a defective measurement, not a measurement of zero"; it is pinned to
0.0 rather than allowed to propagate a NaN into the frozen encoder, where "one NaN would poison the
whole forward pass." The finite_pixels count in the report keeps the gap detectable.
resolve_use_8_bit(config) resolves the flag in a documented order β environment variable first, then
croma.use_8_bit when the key exists and is a bool, then DEFAULT_USE_8_BIT β and returns
(value, source) where source is "env", "config" or "default". The reason it is not simply a
config key is recorded: adding croma.use_8_bit to configs/base.yaml would move Config.hash away
from the frozen 78f1e3700da15aa1 and detach the Phase 9 benchmark. The hash-exempt channel is the same
pattern already used for SATQUERY_DEVICE.
summarise(optical_report, sar_report) bundles both modalities under one use_8_bit and asserts they
agree, raising ValueError if they do not: "upstream applies one transform to both."
7.5 The display stretches β a different transform
Two other percentile stretches exist in the codebase. Neither is the CROMA transform, and the DEV-2 ruling names reusing one of them as forbidden.
preprocessing/imagery.py::load_image_array produces a displayable (H, W, 3) uint8 array:
arr = array.astype(np.float32)
finite = arr[np.isfinite(arr)]
if finite.size:
lo, hi = np.percentile(finite, (2, 98))
if hi > lo:
arr = (arr - lo) / (hi - lo)
arr = np.clip(arr, 0.0, 1.0)
return (arr * 255).astype(np.uint8)
Its docstring is explicit about its scope: it "does not resample, crop, or reproject. Those change the pixel grid, and the grounding specialist converts normalized boxes to pixel coordinates using the ORIGINAL raster's dimensions β a silent resize here would put every box in the wrong place." The determinism claim is stated: "the same file always yields the same array, so a grounding box and a VQA answer describe identical pixels."
docs/PHASE14_CROMA_NORMALISATION_CHANGE.md Β§1.1 records the three disqualifying differences between
this stretch and CROMA's: it is per-scene (one (lo, hi) over all bands) rather than per-channel;
it returns uint8 3-band rather than float 12-channel; and its purpose is presentation, not
radiometry.
specialists/change/postprocess.py::to_grayscale_float applies a per-scene 2/98 percentile stretch
to a grayscale collapse, so that "a uint16 raster and a uint8 raster of the same scene correlate
identically." Its comment records a measured defect fixed in place: np.percentile returns float64
scalars even for float32 input, so the array is promoted; without the explicit
.astype(np.float32, copy=False) the function returns float64 and cv2.phaseCorrelate fails its own
type assertion against a CV_32F Hanning window β "Measured, not theorised."
7.6 The unimplemented stage, and its consequence
The percentile/dB conditioning stage is not implemented. The radiometry module's docstring states the consequence plainly:
Consequence: the two
optical.*and threesar.*config keys are still read by no code. That is recorded, not fixed.
docs/PHASE14_CROMA_NORMALISATION_CHANGE.md Β§6 item 3 adds: "A config key that no code reads is worse
than a missing key: it reads as a satisfied requirement." The ruling's follow-on table assigns the
implementation to the engineer, with the alternative of explicitly retiring the keys.
The status of the whole chain is therefore:
| Stage | Declared | Implemented | Read by code |
|---|---|---|---|
| Optical percentile 2/98 conditioning | β
optical.* |
β | β |
| SAR dB clip β30/+5 conditioning | β
sar.* |
β | β |
Per-channel mean Β± 2Β·std β [0,1] |
β (ruling Β§3.1) | β
radiometry.py |
β
(via training/fusion/extract.py; see below) |
Where the implemented stretch actually runs. The serving-path pipeline
(specialists/optical_sar/inference.py::run_pipeline) does not call normalise_for_croma. It
canonicalises, resizes to 120Γ120 and calls encoder.encode(...). The stretch is applied by the
extraction path (training/fusion/extract.py), which states: "This pipeline applies the DEV-2
encoder-input stretch itself (radiometry.normalise_for_croma), because the arm's use_8_bit has to be
applied at exactly one place." extract.py also refuses an encoder that would stretch a second time:
_assert_raw_encoder "refuses it by name rather than trusting the caller." So the implemented
radiometric transform is a training-time transform that feeds the cached features, and the serving
path's input scaling is whatever CROMAEncoder.encode does internally (see Β§8.9 and the CROMA module).
UNKNOWN β not established from the available evidence whether the serving path's encoder applies the
same mean Β± 2Β·std stretch internally; specialists/optical_sar/croma.py::CROMAEncoder.encode was
verified to have the pinned signature (self, optical, sar) with no mask parameter, and the extraction
module's guard implies the encoder can apply a stretch when normalize_input=True (its default), but
the exact serving configuration of that flag is not established from the files read for this chapter.
8. The CROMA sensor adapter contract
specialists/optical_sar/sensor_adapter.py is "the one place where 'some sensor's bands' becomes
'CROMA's canonical channels', and it exists so that translation is explicit, inspectable and testable."
Its module docstring states the two hard rules from the architecture plan:
"No invented missing bands." "Do not fabricate spectral bands."
The failure mode the module prevents is described with unusual clarity:
A 4-band Cartosat-2S scene maps cleanly onto canonical channels 1-4. Filling channels 5-12 with something β a copy of B3 as a fake "red-edge", a zero-order hold, an interpolation β produces a tensor whose shape is correct and whose statistics look plausible. CROMA would run. Nothing would raise. And every downstream number would be computed from channels that no sensor ever measured, presented as if they were measurements.
So: missing channels are zero-filled, and the availability mask says which ones were real.
8.1 SensorDescriptor
The schema object (core/schemas.py:138) is extra="forbid", described as the "Frozen sensor-adapter
contract (plan section 18)":
| Field | Type | Default | Meaning |
|---|---|---|---|
sensor |
str |
required | Sensor identifier, e.g. "cartosat_2s" |
available_bands |
list[str] |
[] |
The bands the sensor actually delivers |
band_map |
dict[str, str] |
{} |
Source band name β canonical band name |
normalization |
str |
"percentile" |
Identifier recorded for the trace |
availability_mask |
list[bool] |
[] |
Which canonical channels are real |
resolution |
float | None |
None |
Ground sample distance in metres |
canonical_channels |
int |
0 |
Target channel count (12 optical, 2 SAR) |
missing_channels_zero_filled |
bool |
True |
The zero-fill convention |
Why band_map is canonical-channel-keyed. The module docstring addresses a direction question: the
plan's example writes band_map sensor-keyed ({"B1": "c1", ...}), and the schema types it as
dict[str, str]. This module uses the schema's direction because the question asked at load time is
forward-looking: "canonical channel 3 β did I get a real band for it, and if so which one?"
canonical_to_sensor is exposed as the inverse for reporting.
8.2 The canonical channel orders
OPTICAL_CANONICAL: tuple[str, ...] = (
"B01", # coastal aerosol
"B02", # blue
"B03", # green
"B04", # red
"B05", # red edge 1
"B06", # red edge 2
"B07", # red edge 3
"B08", # NIR
"B8A", # narrow NIR
"B09", # water vapour
"B11", # SWIR 1
"B12", # SWIR 2
)
SAR_CANONICAL: tuple[str, ...] = ("VV", "VH")
The comment above OPTICAL_CANONICAL states the contract: "The ORDER is part of the contract: channel
index i must always mean the same physical band, or a pretrained encoder's weights would be applied to
the wrong input." This is Sentinel-2 L2A with the cirrus band removed, per CROMA's README: "Sentinel-2
must be 12 channels (remove the cirrus band if necessary)." Note that B10 (cirrus) is not in the
canonical order β that is the removal.
SAR_CANONICAL is described as "Sentinel-1 dual-pol order CROMA was pretrained with. Positional fallback
ONLY."
_BAND_ALIASES maps common spellings onto the canonical form: B1βB01, BLUEβB02, NIRβB08,
SWIR1βB11, VVPOLβVV, and so on. The comment is a caution: "Kept small and explicit: a guess here
becomes a silently mislabelled band." _canonicalise strips whitespace and underscores and uppercases
before lookup, and never invents a band β an unmapped name passes through unchanged and is later
rejected.
8.3 Building an optical adapter
build_optical_adapter(sensor, *, available_bands, canonical_channels=12, normalization="percentile", resolution=None) (specialists/optical_sar/sensor_adapter.py:205):
- When
available_bandsisNone, the sensor is looked up in_KNOWN_OPTICAL_SENSORS; an unknown sensor with no explicit band list raisesUnsupportedBandsErrorβ "refusing to guess a band arrangement." - Each band is canonicalised. Any canonical name not in
OPTICAL_CANONICALraisesUnsupportedBandsError: "do not map onto the canonical Sentinel-2 order ... supply an explicit band_map rather than guessing a position." - Duplicate bands raise.
- More bands than
canonical_channelsraises. band_mapis built as{source: canonical for source, canonical in zip(bands, canonical)}β keyed by the source name, because "that is what an array row can actually be looked up by."availabilityis[canonical_name in band_map.values() for canonical_name in OPTICAL_CANONICAL[:canonical_channels]].
The returned descriptor has missing_channels_zero_filled=True.
_KNOWN_OPTICAL_SENSORS is "deliberately tiny and explicit. Adding one is a code change, not a guess at
runtime":
| Sensor | Bands |
|---|---|
sentinel_2 |
All 12 canonical |
cartosat_2s |
["B1", "B2", "B3", "B4"] |
cartosat_2 |
["B1", "B2", "B3", "B4"] |
landsat_8 |
["B1", "B2", "B3", "B4"] |
The comment on cartosat_2s is a standing instruction: "Cartosat-2S is a 4-band VNIR imager. It is NOT a
Sentinel-2 clone and must not be treated as one: the 8 channels it lacks stay masked-off."
8.4 Describing a SAR sensor β inspect, do not assume
describe_sar(sensor, *, available_bands, canonical_channels=2, normalization="db", resolution=None)
(specialists/optical_sar/sensor_adapter.py:318) implements the plan's instruction:
"RISAT imagery may vary in acquisition/polarization characteristics... the SAR adapter must inspect actual available channels rather than assuming one fixed polarization pair."
The module docstring explains why a hardcoded ["VV", "VH"] would be wrong: "ISRO describes RISAT-1 as
a C-band SAR mission with multiple polarization configurations (single HH, single VV, dual HH+HV, dual
VV+VH, ...). A hardcoded ["VV", "VH"] would silently mislabel an HH/HV scene, and the resulting tensor
would be a confidently wrong input rather than a recognisably broken one."
When available_bands names the polarisations, they are honoured. When it does not, the channel count
is used to select the most likely convention, and the choice is recorded on the descriptor's
available_bands so the assumption appears in the trace rather than being invisible.
The slot order is the sensor's own:
band_map: dict[str, str] = {source: source for source in bands}
availability = [i < len(bands) for i in range(canonical_channels)]
The comment explains: "Forcing HH/HV into a VV/VH frame would mean either relabelling the polarisations
(a lie the mask would then contradict) or masking both off and discarding real data." So the tensor holds
the radar measurements that exist, and available_bands[i] says which polarisation slot i is.
_known_sar(sensor) is described as "Best-effort polarisation layout for a named SAR sensor, or []":
| Sensor key | Returns |
|---|---|
sentinel_1, sentinel1 |
["VV", "VH"] |
risat* |
["VV", "VH"] β recorded as an ASSUMPTION, not a fact |
alos_palsar, palsar, alos2 |
["HH", "HV"] |
| anything else | [] β "An empty return is honest: it says 'I do not know this sensor's polarisations'" |
The risat case carries an explicit comment: "RISAT-1's most common operational mode is dual-pol.
Recorded as an ASSUMPTION, not a fact β see the module docstring. Callers with real band metadata should
pass available_bands and override this." This is an ASSUMPTION in the project's status vocabulary,
and it is the reason the hidden-set behaviour on RISAT cannot be claimed as measured.
8.5 Applying the canonical layout β _apply_canonical
This is "the function the two hard rules live in" (specialists/optical_sar/sensor_adapter.py:428). It
places real bands in canonical positions and zero-fills the rest:
canonical = np.zeros((n_canonical, height, width), dtype=np.float32)
mask = np.zeros((n_canonical,), dtype=bool)
for source_index, source_name in enumerate(descriptor.available_bands):
if source_index >= n_source:
continue
canonical_name = descriptor.band_map.get(source_name)
if canonical_name is None or canonical_name not in order:
continue
channel_index = order.index(canonical_name)
if channel_index >= n_canonical:
continue
canonical[channel_index] = array[source_index].astype(np.float32)
mask[channel_index] = True
The docstring's key claim: "There is no branch anywhere below that writes a value into an unavailable
channel." The absence of such a branch is the mask. If no band could be placed while the array
carries bands, it raises UnsupportedBandsError rather than returning an all-zero tensor silently.
_canonical_order(descriptor) decides the slot order: for a descriptor whose available_bands are all
SAR polarisations, the order is the descriptor's own list (padded to canonical width); otherwise it is
OPTICAL_CANONICAL.
adapt_optical and adapt_sar are thin wrappers that call _apply_canonical with the right modality
label. adapt_sar's docstring names the plan's internal representation exactly:
canonical_sar[2, H, W] + sar_channel_mask[2].
8.6 SensorAdapterOutput
The dataclass returned by adapt_* carries canonical (C, H, W) float32, mask (C,) bool, and the
descriptor. It exposes n_channels, missing_indices ([i for i, ok in enumerate(self.mask) if not ok]), available_bands, and a to_dict() that includes n_available and missing_channels.
8.7 positional_fallback_descriptor
Used when a raster has the right channel count for a modality but no band labels. The fallback
order is chosen from _FALLBACK_SAR_ORDER (("VV", "VH"), ("HH", "HV")) for 1β2 channels, or the
canonical optical prefix for β₯3 channels. The fact that it was a fallback is recorded in the sensor
field so it "cannot be mistaken for measured metadata." The complement=True flag selects the other
convention, which lets a caller test both HH/HV and VV/VH orderings without asserting which is correct.
The optical-SAR specialist uses this in _descriptor_for
(specialists/optical_sar/specialist.py:398) and appends a warning that names the assumption: "The
polarisation IDENTITY of each channel is NOT known from a band count alone β supply a sensor descriptor
to state it."
8.8 Finding C-1 β the mask goes to the fusion head, not to CROMA
This is the single most important fact in this section. core/schemas.py's module docstring lists it
among the findings baked into the schemas: "C-1 channel-availability mask is a first-class fusion input,
never a CROMA input."
specialists/optical_sar/fusion_head.py devotes a section to the rationale:
CROMA is a masked autoencoder. Handing it an availability mask invites it to reconstruct the missing channels β which is precisely the fabrication the sensor adapter exists to prevent: the model would output plausible values for bands no sensor measured, and those values would then be treated as data. The mask is therefore consumed HERE, where it can only do one thing: tell the classifier which inputs to distrust.
Mechanically: CROMAEncoder.encode has a pinned signature of exactly (self, optical, sar) β the
radiometry module records that tests/unit/test_optical_sar_croma.py "asserts the absence of any mask
parameter, because freeze finding C-1 says CROMA never receives one." The mask instead enters at
assemble_fusion_input (Β§8.9). The optical-SAR specialist's JOINT_FEATURE_REGION evidence item records
the fact explicitly in its payload: "mask_consumed_by": "fusion_head" and "croma_received_mask": False (specialists/optical_sar/specialist.py:816).
EvidenceType even has a dedicated member for the mask as trust evidence:
AVAILABILITY_MASK = "availability_mask" # C-1: modality trust evidence (core/schemas.py:76). The
optical-SAR specialist emits one AVAILABILITY_MASK item per modality, carrying the sensor, canonical
channel count, available bands, band map, the mask itself, the missing indices,
missing_channels_zero_filled, normalization and resolution.
8.9 The frozen fusion concatenation
specialists/optical_sar/fusion_head.py defines the concatenation that the mask feeds into:
optical_GAP (B, 768)
SAR_GAP (B, 768)
joint_GAP (B, 768)
optical_mask (B, 12) <- availability, from the sensor adapter
sar_mask (B, 2) <- availability, from the sensor adapter
---------
concat (B, 2318)
expected_fusion_dim(encoder_dim=768, optical_channels=12, sar_channels=2, modalities_used=3) returns
3*768 + 12 + 2 = 2318. The module recomputes the width and refuses to build on a mismatch, so "a
config edit cannot silently reshape the first Linear layer into something that trains but means
nothing." core/config.py recomputes it independently at load time from croma.encoder_dim,
croma.modalities_used, croma.optical_channels and croma.sar_channels, and rejects a config whose
fusion.input_dim disagrees (finding C-1 in the validator's own message).
assemble_fusion_input asserts the order rather than assuming it β "a permutation here produces a tensor
of exactly the right shape that trains to a worse number β the hardest kind of bug to notice" β and
validates that the three GAP vectors agree on batch size and that both masks agree with it.
channel_dropout(features, mask, *, keep_probabilities, rng) is the mechanism that "teaches the head to
trust the availability mask". Its load-bearing line is:
keep = draws & (msk > 0.0)
out_mask = keep.astype(np.float32)
out = out * out_mask
Dropped channels are zeroed, not renormalised β "renormalising would fabricate a scale that the real
missing-channel case does not have" β and the mask is updated in lockstep. docs/PHASE14_CROMA_NORMALISATION_CHANGE.md
Β§5 independently verified this against freeze Β§2.5 and ruled: "the engineer's implementation SATISFIES
Β§2.5's intent. No change entry required." Its table records that the & (msk > 0.0) guard is the
load-bearing one: "no training transform can resurrect a band the sensor did not measure."
The schedule (plan section 20) is optical at 100/80/60/40% and SAR at 100/50%, declared in
training/fusion/train.py as OPTICAL_DROPOUT_RATES = (1.0, 0.8, 0.6, 0.4) and
SAR_DROPOUT_RATES = (1.0, 0.5).
mask_availability_stats(optical_mask, sar_mask) returns optical_available_fraction,
sar_available_fraction, optical_channels_present and sar_channels_present. These feed the
confidence system: the two modality-confidence terms in the optical-SAR specialist are exactly the
availability fractions, which the specialist's docstring justifies: "On the hidden set the optical side
is Cartosat-2S with 4 bands against a canonical 12, so 8 of 12 channels are always absent. A classifier
resting on 4 measured channels is not the same claim as one resting on 12."
8.10 The optical-SAR confidence components
OpticalSarSpecialist._confidence_components (specialists/optical_sar/specialist.py:481) implements
plan section 26's four components with these weights:
CONFIDENCE_WEIGHTS: dict[str, float] = {
"fusion_margin": 0.40,
"optical_confidence": 0.20,
"sar_confidence": 0.20,
"cross_modal_agreement": 0.20,
}
cross_modal_agreement is cosine(optical_GAP, SAR_GAP), rescaled from [-1, 1] to [0, 1] via
AGREEMENT_FLOOR = -1.0. cross_modal_agreement in inference.py returns None on any exception or a
shape mismatch, and the specialist treats None as a zero component. The composition is a weighted sum,
floored to 0.0 when there is no prediction or no trained head β "Both are signal GAPS, not weak
signals, so they resolve to 0.0 rather than merely discounting." The components dict records
has_croma, trained_head, no_prediction and the raw availability counts, so "the trace states every
reason the score is zero, not just the first one encountered."
9. Pair-compatibility rules
Two workflows take a pair of assets, and both enforce compatibility before reading pixels.
9.1 Change β exactly two, temporally distinct, co-registered
ChangeSpecialist.validate_request (specialists/change/specialist.py:190) enforces exactly two
assets. The comment names both common mistakes: "ONE asset is the common mistake: the caller treated
this as a single-image task. THREE is the other: they attached a time series. Both are rejected with the
same typed error." The user-facing message is "Change detection needs exactly two images: an earlier one
and a later one."
_assess_pair then runs three checks in order:
- Temporal distinctness. Identical
pathraisesTemporalPairError. Identicalsha256(when both carry one) also raises: "a repeated acquisition is not a temporal pair." - CRS compatibility via
compare_crs(Β§2.5). Incompatible or reprojection-requiring pairs get a warning, not a refusal β pixel-domain change detection can still proceed, but geographic claims are withheld. - Registration quality via
measure_registration(Β§9.3).
If registration is not usable, the specialist suppresses spatial claims while still computing and returning the change maps:
The maps are still returned (the caller may want to look at them), but the REGION CLAIMS are withheld: a region is an assertion about where something changed on the ground, and we cannot make it.
ChangeSpecialist._geospatial drops the transform in that case, returning a GeoMetadata with
is_georeferenced=False β "a transform invites the caller to place a region on a map, and we have just
said the regions are not placeable." _write_artifacts writes the map without georeferencing when
suppressed, and the comment explains why copying T1's transform would be wrong: "Copying T1's transform
onto it would produce a file that unlocks exactly the placed-on-a-map reading the suppression exists to
forbid." The confidence is driven to the floor via a minimum against the registration factor, so
alignment is a gate, not one vote among three.
The change confidence weights are:
CONFIDENCE_WEIGHTS: dict[str, float] = {
"registration_quality": 0.50,
"mean_change_probability": 0.30,
"component_stability": 0.20,
}
with STABILITY_HEADROOM = 0.5, DEFAULT_MAX_SHIFT_PX = 8, MIN_REGION_PIXELS = 32.
9.2 Optical-SAR β exactly two, one of each modality
OpticalSarSpecialist.validate_request (specialists/optical_sar/specialist.py:192) enforces exactly
two assets and that they are one optical and one SAR. The docstring explains the case worth
explaining:
Two optical images is not a malformed request in the way that three is β it is a plausible mistake, from an operator who uploaded a bi-temporal pair to a single-modality workflow, or who labelled a four-band Cartosat scene as SAR.
If such a request were accepted, the adapter would place the second optical image into the 2-channel SAR slot, zero-fill, and produce a fused tensor of the correct shape. CROMA would run. The answer would be about one modality while claiming to fuse two.
So the modality check is done on the assets, before any pixels are read. _assess_pair raises
InvalidRequestError with a machine-readable reason of "two_optical", "two_sar" or
"indeterminate", and each carries a user-facing message.
9.3 The equal-dimensions requirement, and the 726Β²-vs-736Β² error
The change detector requires T1 and T2 to have identical shapes. The guard is in the model itself,
specialists/change/stanet.py:415:
if t1.shape != t2.shape:
raise SpecialistError(
f"T1 and T2 must have the same shape; got {tuple(t1.shape)} "
f"and {tuple(t2.shape)}",
specialist="change",
)
This is a hard requirement, and it is where the bundled demo pair fails. docs/FINAL_DELIVERY_REPORT.md
Β§6 records it as DEGRADED:
| Finding | Status | Detail |
|---|---|---|
| Change on the bundled EO pair | DEGRADED | delta-growth-t0-1975.jpg (726Β²) and t1-2025.jpg (736Β²) differ in shape; the change specialist errors (T1 and T2 must have the same shape). A same-shape pair, or a resize step, is needed for a clean change demo. |
There is no resize step on the change path. load_image_array explicitly does not resample
(Β§7.5) β its docstring states "a silent resize here would put every box in the wrong place" β and the
change specialist's _predict only pads for the encoder's stride-8 requirement, cropping back to the
original size:
h, w = t1.shape[-2:]
if h % 8 or w % 8:
ph, pw = (-h) % 8, (-w) % 8
t1 = torch.nn.functional.pad(t1, (0, pw, 0, ph), mode="reflect")
t2 = torch.nn.functional.pad(t2, (0, pw, 0, ph), mode="reflect")
That padding does not reconcile two different input sizes: T1 and T2 are each padded by their own
(-dim) % 8, so a 726Β² and a 736Β² input become 728Β² and 736Β² respectively β still different β and the
model raises. The reflect-pad is a stride accommodation, not a pair accommodation, and the two must
not be conflated.
Consequence for operators. The change workflow requires a same-shape pair, and the failure is
IMPLEMENTED (it raises a typed error with a readable message) rather than silent. A resize or
co-registration step that reconciles differing dimensions is not implemented (Β§13). The
FINAL_DELIVERY_REPORT records the demo as DEGRADED and prescribes "a same-shape pair, or a resize
step" β either would fix the demo; neither is currently in the code path.
9.4 Registration measurement
measure_registration (specialists/change/postprocess.py:134) uses cv2.phaseCorrelate on grayscale
collapses of the two images, returning a RegistrationQuality with shift_x, shift_y, response,
max_shift_px, is_usable and reason. Defaults: max_shift_px=8, min_response=0.15.
The module docstring is candid about the method's limit:
Deliberately a TRANSLATION model. Real mis-registration includes rotation and warp, which phase correlation does not recover; a small residual after a translation correction is therefore not proof of good alignment. What the measurement can do honestly is flag the bad cases, and that is how it is used: as a gate, not a certificate.
RegistrationQuality.confidence_factor() applies two independent penalties and takes the worse:
a shift penalty (1.0 at zero shift, 0.0 at max_shift_px) and a response factor (0.0 at zero
response, 1.0 at response β₯ 0.15). A pair that is offset and uncorrelated takes the worse of the two β
"the conservative choice."
A failure to measure is treated as a measurement of failure β both mean the pair cannot be trusted spatially β but the distinction is recorded in the warning: "A failure to MEASURE is not the same as a measurement of failure."
Two registration-gate hypotheses were tested and both are recorded as REJECTED or
INSUFFICIENT: docs/CHANGE_REGISTRATION_GATE_COHERENCE_TEST.md records a rejected hypothesis (AUC S2
vs S3 β€ 0.59), and docs/CHANGE_REGISTRATION_GATE_TEXTURE_TEST.md records a low-texture signal that was
CONFIRMED but INSUFFICIENT, with option 3 ELIMINATED. So the gate as implemented (phase
correlation) is the surviving method, and no additional coherence- or texture-based gate was adopted.
10. Large-image tiling policy
10.1 What is declared
configs/base.yaml declares the full tiling policy:
image:
max_pixels: 25000000
tile_size: 512
tile_overlap: 128
max_tiles: 64
# plan section 9.1 tile policy: whole-image thumbnail first, then top-K tiles.
# max_tiles is the hard ceiling on tiles *examined*; top_k_tiles is how many
# are actually sent through a specialist.
top_k_tiles: 4
| Key | Value | Declared meaning |
|---|---|---|
max_pixels |
25,000,000 | Pixel budget; exceeding it is recoverable (the caller may downscale) |
tile_size |
512 | Tile edge in pixels |
tile_overlap |
128 | Overlap between adjacent tiles |
max_tiles |
64 | Hard ceiling on tiles examined |
top_k_tiles |
4 | How many tiles are actually sent through a specialist |
The comment states the intended policy: "whole-image thumbnail first, then top-K tiles." The change
workflow declares a different tiling resolution, change.tile_size: 256 with change.tile_overlap: 0
β 256 is "the change model's OWN training resolution" (LEVIR-CD patches are 256Γ256).
10.2 What is enforced
core/config.py validates the tiling policy at load time, but only one relationship:
# --- tiling policy -------------------------------------------------
if self.get("image.top_k_tiles", 0) > self.get("image.max_tiles", 0):
errors.append("image.top_k_tiles cannot exceed image.max_tiles")
So the config loader guarantees that the number of tiles sent through a specialist does not exceed the number examined. It does not implement tiling.
The pixel budget is enforced, but at the raster-inspection layer rather than as a tiling step:
inspect_raster raises OversizedImageError when width * height > max_pixels, and the error is
recoverable β the docstring says "the caller may downscale." The controller passes the configured
budget into _inspect_asset during VALIDATE.
10.3 What is not implemented β the honest status
The tiling policy is DECLARED (not read). image.tile_size, image.tile_overlap, image.max_tiles
and image.top_k_tiles appear only in configs/base.yaml and in the core/config.py validation above.
No code on the serving path slices a raster into tiles, selects the top-K, or stitches results.
architecture/03-request-lifecycle.md reaches the same conclusion: the tiling policy is referenced only
in config validation and not implemented on the serving path.
The practical consequence: a raster within the max_pixels budget is handed to a specialist whole.
The VLM path pins processor_longest_edge: 512 so the processor does not upscale and split a tile, but
that pin assumes a 512-px input β it is a control against the processor's default 2048, not a tiling
step. core/config.py enforces the relationship (processor_longest_edge must not exceed
image.tile_size), with the measured rationale recorded: the default 2048 upscales a 512 tile 4Γ and then
splits it into ~17 sub-images, "the real figure is ~17x" against the plan's estimated 4Γ.
UNKNOWN β not established from the available evidence how a raster between the tile size and the pixel
budget is handled by the VLM path in practice; the config implies the tile policy would govern it, but
the policy is not implemented, so the behaviour is whatever the specialist does with the whole image.
11. Grounding resolution β the frozen 224 decision
Grounding runs at 224 px, and this is one of the project's cleanest pre-registered results.
configs/base.yaml:
grounding:
...
# RESOLVED 2026-09-16. 448 did NOT earn its cost over all 16,159 VRSBench
# eval records: mean best IoU -0.0147, Recall@0.5 -0.0022, and every recall
# threshold lower (-0.0699 @0.10, -0.0243 @0.25), at 1.59x the latency.
# Paired over identical samples: mean diff -0.0147, 95% CI [-0.0160,-0.0134],
# t = -22.63. 448 was better on 8.5% of records, worse on 20.9%.
# Pre-registered rule and the paired test AGREE on 224.
# Evidence: docs/PHASE7_RESOLUTION_DECISION.md
image_size: 224
resolution_frozen: true
nms_iou: 0.50
max_candidates: 20
confidence_threshold: 0.40
benchmark_box_scale: 100.0
coordinate_system: normalized_0_1
encoder_projected_dim: 512
11.1 The pre-registered rule
docs/PHASE7_RESOLUTION_DECISION.md records the rule, fixed before the result was seen:
448 WINS if Recall@0.5 improves by >= 0.05 absolute
OR mean best IoU improves by >= 0.05 absolute
224 WINS otherwise
INCONCLUSIVE if fewer than 30 samples were scored
The artifact records rule_changed_since_preregistration: false.
11.2 The result
224 WINS
recall@0.5 gain 448/224 : -0.0022
bestIoU gain 448/224 : -0.0147
latency ratio : 1.59x
Neither component came close to the +0.05 margin; both were negative. The measured detail:
| Metric | 224 | 448 | Delta |
|---|---|---|---|
| Token grid | 7 Γ 7 = 49 | 14 Γ 14 = 196 | 4.0Γ tokens |
| Mean best IoU | 0.0972 | 0.0825 | β0.0147 |
| Recall@0.10 | 0.3298 | 0.2599 | β0.0699 |
| Recall@0.25 | 0.1187 | 0.0944 | β0.0243 |
| Recall@0.50 | 0.0234 | 0.0212 | β0.0022 |
| Latency mean | 20.0 ms | 31.8 ms | 1.59Γ |
| Latency p90 | 20.9 ms | 32.9 ms | 1.57Γ |
| Peak VRAM | 592.1 MB | 599.8 MB | +7.7 MB |
All 16,159 / 16,159 VRSBench eval records were scored at both resolutions on a Tesla T4.
11.3 The paired test
Because both resolutions scored the same 16,159 samples, the paired test is the stronger statistic:
paired samples : 16159
mean 224 : 0.0972
mean 448 : 0.0825
mean paired diff : -0.0147 (95% CI -0.0160 .. -0.0134)
t statistic : -22.63
CI excludes zero : True
448 better on : 1371/16159 ( 8.5%)
448 worse on : 3372/16159 (20.9%)
identical : 11416/16159 (70.6%)
The paired test and the pre-registered rule agree, so there is "no rule-versus-evidence disagreement to escalate." The win/loss split is informative on its own: 448 wins on only 8.5% and loses on 20.9% β "the finer grid is not merely neutral, it is actively harmful on a fifth of the corpus."
The recall ladder narrows as the threshold rises (β0.0699 at 0.10, β0.0243 at 0.25, β0.0022 at 0.50), which the document reads as "the signature of a method that cannot reach high IoU either way."
11.4 Why 448 did not help
The document's honest reading: the zero-shot method selects a patch by text similarity and returns that patch's box. At 224 a box is 1/7 of the image; at 448 it is 1/14. A finer grid is only better if the target is small and the similarity peak lands on the correct fine cell. Two things work against that: the peak is not sharper at 448 (the similarity field on frozen features is smooth, so the argmax moves around), and Recall@0.10 β the loosest threshold β degrades most, which means "the fine grid is adding positional noise rather than positional precision."
The document is careful to attribute this to the zero-shot baseline, not to RemoteCLIP: "A learned head trained to regress boxes from these features may respond differently to resolution."
11.5 What the decision does not establish
- Whether the zero-shot baseline is good. It is not: mean best IoU 0.0972 and Recall@0.5 0.0234 are "weak". It is an ablation floor for the Phase 8 head.
- Whether a trained head has the same resolution sensitivity. Re-opening the question is legitimate if the head's validation curve suggests it, and would be "a new pre-registered experiment, not a silent retune."
- Anything about hidden ISRO/SAC imagery. "VRSBench is overhead optical. The hidden set is Cartosat-2S + RISAT, a different distribution entirely."
11.6 Downstream consequences of the frozen resolution
The frozen resolution propagates into concrete serving parameters:
- The decode grid.
build_grounding_specialistsetsgrid = resolution // 32, so at 224 the grid is 7Γ7 = 49 cells. The head emits(1, 49, 5)β[tx, ty, tw, th, objectness]per cell. - The projected dimension.
grounding.encoder_projected_dim: 512is the measuredvisual.projoutput of RemoteCLIP ViT-B/32 (the transformer width is 768;visual.projmaps to 512). The per-cell feature isconcat([patch, text, patch*text, global_pool]) = 4 Γ 512 = 2048, whichgrounding_head.feature_dim: 2048must equal β enforced at load time bycore/config.py, which rejects any other value, because "a mismatch here is a SILENT shape error" that torch raises only at the similarity step, "after the patch features are already cached." - NMS.
grounding.nms_iou: 0.50is passed tonms(boxes, scores, iou_threshold)(specialists/grounding/head.py:294). - The serving candidate budget.
GroundingSpecialistdefaultsmax_candidates=6, not the config's20. The class docstring records why: "The Phase 8 decision measured that a 20-box budget inflates mean best IoU relative to the 5.99-box baseline; 6 is the decode-matched setting." The builder readsgrounding.serving_max_candidateswith a default of 6. - The degenerate-box guard.
DEGENERATE_AREA_FRACTION = 0.9drops any zero-shot candidate covering more than 90% of the frame, "the shape a failed localisation takes once it has been smoothed into a rectangle." This guard applies to the zero-shot fallback path only β "the trained-head decode is a regressed box with its own learned prior; filtering it would change the frozen Phase 8 benchmark numbers." - The decode function.
decode_cell_relative(specialists/grounding/head.py:204) turns(B, N, 4)cell-relative parameters into(B, N, 4)normalised xyxy, clamped to[0, 1]and order-corrected byenforce_order(which swaps inverted corners rather than letting IoU go silently to zero). The decode is "differentiable: the output feeds the box and GIoU losses directly, so the loss is computed on exactly the coordinates the metric measures."
12. Worked examples
12.1 A 4-band Cartosat-2S optical scene
- Inspect.
inspect_rasterreadsband_count=4, no CRS or with CRS.infer_modality(4)returnsOPTICAL(4 is in_OPTICAL_BAND_COUNTS). - Descriptor. With no
SensorDescriptor,_descriptor_formaps the 4 bands positionally ontoOPTICAL_CANONICAL[:4]=B01, B02, B03, B04and appends a warning naming the assumption. With a descriptor built bybuild_optical_adapter("cartosat_2s"), the bands are["B1","B2","B3","B4"]canonicalised toB01..B04, andavailability = [True, True, True, True, False Γ 8]. - Canonicalise.
adapt_opticalreturns a(12, H, W)tensor with channels 0β3 carrying the real bands and channels 4β11 exactly zero, plus a 12-element mask with fourTrue. - Resize.
resize_to_canonicalbilinearly resizes to120Γ120. The zero channels stay zero. - Encode.
encoder.encode(optical_tensor[None], sar_tensor[None])β no mask (C-1). - Assemble.
assemble_fusion_inputconcatenates the three 768-d GAPs with the 12-element optical mask and the 2-element SAR mask β(1, 2318). - Head. With a trained head, the classifier emits 19 logits β softmax β a label and a margin. With
no head,
probabilitiesstaysNone,degraded=True, and no label is invented. - Evidence. An
AVAILABILITY_MASKitem per modality records the sensor, the 4 real bands, the band map, the mask ([true Γ 4, false Γ 8]), andmissing_channels: [4,5,6,7,8,9,10,11]. AJOINT_FEATURE_REGIONitem recordsfusion_input_dim: 2318,mask_consumed_by: "fusion_head",croma_received_mask: false. - Confidence.
optical_confidence = 4/12 = 0.3333; the SAR side contributes its own fraction; the fusion margin (if a head exists) is the largest weight. If no head, the raw confidence is 0.0 withno_prediction: 1.0recorded.
12.2 A RISAT SAR scene with no band labels
- Descriptor.
positional_fallback_descriptor(stem, 2)selects("VV", "VH")from_FALLBACK_SAR_ORDERand records the fallback in thesensorfield. The specialist appends: "The polarisation IDENTITY of each channel is NOT known from a band count alone." - Canonicalise.
adapt_sarreturns(2, H, W)with both slots filled, mask[True, True]. - Trace.
available_bands = ["VV", "VH"]is the assumption, not a measurement. A caller with real polarisation labels passesavailable_bands=["HH","HV"], in which case the descriptor's own order becomes the slot order and the mask reflects it β no relabelling, no discarded data.
12.3 A change pair with differing dimensions
- Validate. Exactly two assets, distinct paths/SHA-256, CRS compared.
- Register.
measure_registrationcollapses both to grayscale float32 and runsphaseCorrelate. - Shape check. If the two images differ in size,
STANetStyleChangeDetector.forwardraisesSpecialistError("T1 and T2 must have the same shape; got ...")β the 726Β²/736Β² case (Β§9.3). No resize step exists to reconcile them. - If shapes match but registration is poor: the maps are still computed, the regions are withheld,
the transform is dropped from the
geospatialblock, and the confidence is floored at 0.0 withsuppressed_by_registration: 1.0recorded.
13. What is NOT implemented β exhaustive
Each item below was checked against the code. Where a capability is absent, it is named rather than
implied; where the absence cannot be confirmed from the files read, it is marked UNKNOWN.
| Capability | Status | Evidence |
|---|---|---|
| Reprojection pipeline | NOT IMPLEMENTED | geospatial/crs.py::compare_crs reports requires_reprojection; no code performs a reprojection. preprocessing/imagery.py explicitly does not reproject. |
| Cloud masking | NOT IMPLEMENTED | No cloud, cirrus, QA or SCL handling anywhere in preprocessing/ or the specialists. UNKNOWN β not established from the available evidence whether any dataset adapter applies a cloud mask. |
| Atmospheric correction | NOT IMPLEMENTED | No atmospheric, surface-reflectance or L2A-processing code. Rasters are consumed as delivered. |
| Mosaicking | NOT IMPLEMENTED | No mosaicking, seamline or composite code exists. |
| Tiling / top-K tile selection | DECLARED (not read) | image.tile_size, image.tile_overlap, image.max_tiles, image.top_k_tiles are read only by core/config.py's top_k_tiles <= max_tiles check. No serving-path tiling. |
| Optical percentile 2/98 conditioning | DECLARED (not read) | optical.normalization, optical.lower_percentile, optical.upper_percentile are read by no code (radiometry.py docstring; PHASE14 Β§1). |
| SAR dB clip β30/+5 conditioning | DECLARED (not read) | sar.representation, sar.clip_min_db, sar.clip_max_db are read by no code. |
| Resampling / cropping on the display path | DELIBERATELY ABSENT | preprocessing/imagery.py docstring: "It does not resample, crop, or reproject." |
| Resize to reconcile differing pair dimensions | NOT IMPLEMENTED | The change path requires equal shapes and has no reconciling resize (Β§9.3). |
| Nodata filling | NOT IMPLEMENTED | quality.py refuses any non-finite input outright; its comment records: "Filling nodata from the raster profile belongs in the tiling work." |
| Area computation for geographic CRS | NOT IMPLEMENTED | pixel_area_m2 returns None for a non-metric CRS rather than approximating. |
| GeoJSON / WKT / KML output | NOT IMPLEMENTED | Coordinates are emitted as lists; no geometry serialisation exists. |
| Multi-band raster β RGB band selection policy | PARTIAL | load_image_array takes the first three bands as RGB (or repeats a single band); there is no configurable band combination. |
| Sensor adapter for sensors beyond the known list | REFUSED BY DESIGN | An unknown optical sensor with no explicit band list raises UnsupportedBandsError; _known_sar returns [] for an unknown SAR sensor. |
A note on the "not implemented" list. None of these omissions is hidden. The two most consequential β the tiling policy and the radiometric conditioning β are declared in the frozen config and read by no code, and both the config comments and the DEV-2 ruling name that state explicitly rather than letting the key read as a satisfied requirement.
14. What is NOT RUN / OPEN / BLOCKED for this topic
NOT RUN
- The optical-SAR forward pass at serving time with a trained head.
CROMA_base.ptis not present locally anduse_croma.pymust be vendored; the specialist runs sensor-only and produces no label. - A system-level geospatial accuracy benchmark. No end-to-end benchmark exists; no system-level accuracy is claimed anywhere.
- Reprojection of a genuinely mismatched-CRS pair. The code path that would consume a reprojection is absent, so the case has never been exercised.
- Change on a bundled pair with equal dimensions. The only bundled demo pair is 726Β² vs 736Β² and errors; a same-shape pair has not been run as a demo.
OPEN
- The
optical.*/sar.*conditioning stage.OPENβ either it gets an implementation or the keys are retired (docs/PHASE14_CROMA_NORMALISATION_CHANGE.mdΒ§6 item 3). - The arm-constructibility question. The Β§4-registered arm B ("percentile/dB only, no encoder-input stretch") is not constructible while stage 1 is unimplemented; the arm now named B corresponds to the registered arm C. The ruling records this as "a human decision and deliberately not resolved."
- The optical-SAR ruling. Accuracy 0.931 with macro-F1 0.434161; ruling OPEN. Never quote the accuracy without the macro-F1.
- The grounding resolution question after a trained head. Re-opening 224 is legitimate only as a new pre-registered experiment.
- The change-VQA metric ruling.
OPENand owner-gated. - Whether any call site builds a
Boxoutside the two conversion functions (Β§5.4) βUNKNOWN.
BLOCKED
- A clean change demo. Blocked on a same-shape pair or a resize step.
- The CROMA normalisation experiment (option c). Blocked on arm B being non-constructible and on the
Arm-B feature cache, which does not exist (
artifacts/optical_sar/fusion_features/holds only the Arm-A cache). - The gated experiment's sample size. Cannot be specified until a variance-estimation pass is run; the ruling deliberately declines to invent a number.
REJECTED
- 448 px grounding resolution.
REJECTEDby the pre-registered rule and the paired test. - Option (b) β match
base.yamlalone, accept the shift.REJECTEDas a standalone path (docs/PHASE14_CROMA_NORMALISATION_CHANGE.mdΒ§2.1). - Reusing
preprocessing/imagery.py's stretch for CROMA.REJECTEDasforbidden(Β§3.3 of the ruling). - Coherence- and texture-based registration gates. One
REJECTED, oneCONFIRMED but INSUFFICIENT.
15. Where the evidence lives
Source modules (the code this chapter describes):
| Path | What it defines |
|---|---|
preprocessing/raster.py |
The validation chain, inspect_raster, read_bands, write_raster, infer_modality, file_sha256, pixel_area_m2 |
preprocessing/imagery.py |
load_image_array, load_pil_image, the per-scene 2/98 display stretch |
preprocessing/quality.py |
The deterministic quality gate, MIN_AUTOCORRELATION=0.10, NOISE_ENTROPY_BITS=7.8 |
geospatial/crs.py |
parse_crs, describe_crs, compare_crs, is_metric, CRSCompatibility |
geospatial/transform.py |
PixelWindow, GeoBounds, to_normalized/to_pixel/to_geo, intersect, iou, benchmark_boxes_to_normalized, affine_from_metadata |
core/schemas.py |
CoordinateSystem, GeoMetadata, SensorDescriptor, AssetMetadata, Box, Region, ChangeRegion, Evidence._spatial_needs_crs |
core/config.py |
Config loading, validation, Config.hash, device_preference |
specialists/optical_sar/sensor_adapter.py |
OPTICAL_CANONICAL, SAR_CANONICAL, build_optical_adapter, describe_sar, _apply_canonical, positional_fallback_descriptor |
specialists/optical_sar/radiometry.py |
TRANSFORM_NAME, normalise_for_croma, resolve_use_8_bit, the zero-channel rule |
specialists/optical_sar/fusion_head.py |
expected_fusion_dim, assemble_fusion_input, channel_dropout, build_fusion_head |
specialists/optical_sar/inference.py |
run_pipeline, resize_to_canonical, cross_modal_agreement, build_head_from_config |
specialists/optical_sar/specialist.py |
_assess_pair, _descriptor_for, _confidence_components, _geospatial |
specialists/change/postprocess.py |
RegistrationQuality, measure_registration, to_grayscale_float, regions_to_schema |
specialists/change/specialist.py |
_assess_pair, _predict, _geospatial, the suppression path |
specialists/change/stanet.py:415 |
The equal-shape guard |
specialists/grounding/specialist.py |
_to_geo_boxes, the NORMALIZED_0_1 box construction, DEGENERATE_AREA_FRACTION |
specialists/grounding/head.py |
decode_cell_relative, enforce_order, nms |
specialists/grounding/inference.py |
decode_candidates_from_features, ground_phrase |
evaluation/metrics/grounding.py |
benchmark_to_normalized(box, scale=100.0) |
Configuration:
configs/base.yamlβimage.*(tiling),optical.*,sar.*,croma.*,fusion.*,grounding.*,grounding_head.*,change.*.
Documents:
docs/PHASE14_CROMA_NORMALISATION_CHANGE.mdβ the DEV-2 ruling, the two ordered stages, the Gate F arm amendment.docs/CROMA_NORMALISATION_UPSTREAM_EVIDENCE.mdβ what is and is not established about CROMA's pretraining distribution.docs/PHASE7_RESOLUTION_DECISION.mdβ the 224/448 decision in full.docs/CHANGE_REGISTRATION_GATE_COHERENCE_TEST.md,docs/CHANGE_REGISTRATION_GATE_TEXTURE_TEST.mdβ the rejected and insufficient registration gates.docs/API_CONTRACT.mdΒ§2.5 (asset allowlist), Β§3.2 (coordinate systems).docs/FINAL_DELIVERY_REPORT.mdΒ§6 β the 726Β²/736Β² change demo recorded as DEGRADED.docs/PHASE10_CDVQA_DATA_STATUS.mdβ the CDVQA/SECOND pair layout and temporal semantics.docs/PHASE11_CROMA_STATUS.mdβ the reproduced CROMA forward pass.
Sibling release chapters:
DATA_PIPELINE.mdβ the end-to-end data path, caches and artefacts.architecture/03-request-lifecycle.mdβ the nine controller states, including VALIDATE and PREPROCESS.architecture/05-specialists.mdβ the six specialists.architecture/07-configuration-freeze.mdβ the frozen hash and every config key.DATASETS.mdβ the corpora and their measured shapes.