Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
File size: 92,888 Bytes
cf51713 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 1462 1463 1464 1465 1466 1467 1468 1469 1470 1471 1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 1518 1519 1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 1534 1535 1536 1537 1538 1539 1540 1541 1542 1543 1544 1545 1546 1547 1548 1549 1550 1551 1552 1553 1554 1555 1556 1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 | # Geospatial β deep reference
**Status tags used on every substantive claim:** `IMPLEMENTED` Β· `VERIFIED` Β· `MEASURED` Β· `ATTEMPTED`
Β· `NOT RUN` Β· `BLOCKED` Β· `DEFERRED` Β· `REJECTED` Β· `OPEN` Β· `RESOLVED` Β· `CLOSED`.
Every statement in this document was produced by reading a file in `C:/Users/anish/satquery-ai/`. Where
a fact is not established from the available evidence, this document writes
`UNKNOWN β not established from the available evidence`. Where a transform is *declared* in configuration
but read by no code, that is stated explicitly rather than presented as a working feature. Where a
capability is absent, it is listed in [Β§13](#13-what-is-not-implemented-exhaustive) rather than implied.
This chapter is the reference for the **spatial contract**: how a raster becomes an analysed image, what
georeferencing survives, what coordinate system a box is in, how a sensor's bands become CROMA's canonical
channels, and which of those steps are actually implemented on the serving path. It is written to be
sufficient to reconstruct the subsystem from the document alone.
**Sibling documents:** [`DATA_PIPELINE.md`](DATA_PIPELINE.md) (the end-to-end data path, caches and
artefacts), [`architecture/03-request-lifecycle.md`](architecture/03-request-lifecycle.md) (where the
PREPROCESS and VALIDATE states sit), [`architecture/05-specialists.md`](architecture/05-specialists.md)
(the six specialists in full), [`architecture/07-configuration-freeze.md`](architecture/07-configuration-freeze.md)
(the frozen config identity and every key), [`DATASETS.md`](DATASETS.md) (the corpora and their measured
shapes).
---
## Table of contents
0. [How to read this document](#0-how-to-read-this-document)
1. [The raster input contract](#1-the-raster-input-contract)
2. [CRS handling](#2-crs-handling-geospatialcrs-py)
3. [Affine transforms and coordinate frames](#3-affine-transforms-and-coordinate-frames)
4. [The coordinate-system contract](#4-the-coordinate-system-contract)
5. [Benchmark box scale β 0β100 versus 0β1](#5-benchmark-box-scale--0100-versus-01)
6. [Band-count modality inference](#6-band-count-modality-inference)
7. [Radiometric normalisation and representation](#7-radiometric-normalisation-and-representation)
8. [The CROMA sensor adapter contract](#8-the-croma-sensor-adapter-contract)
9. [Pair-compatibility rules](#9-pair-compatibility-rules)
10. [Large-image tiling policy](#10-large-image-tiling-policy)
11. [Grounding resolution β the frozen 224 decision](#11-grounding-resolution--the-frozen-224-decision)
12. [Worked examples](#12-worked-examples)
13. [What is NOT implemented β exhaustive](#13-what-is-not-implemented-exhaustive)
14. [What is NOT RUN / OPEN / BLOCKED for this topic](#14-what-is-not-run--open--blocked-for-this-topic)
15. [Where the evidence lives](#15-where-the-evidence-lives)
---
## 0. How to read this document
Three distinctions are load-bearing throughout, and conflating them is the single most common way to
misread this system:
1. **Declared versus executed.** `configs/base.yaml` is the authoritative registry, and `core/config.py`
validates and hashes it at load time. But *declaring* a policy and *executing* it are different
claims. The tiling policy, the optical percentile stretch and the SAR dB clip are all declared in
`configs/base.yaml` and are all read by no code on the serving path. This document says so at each
point rather than presenting the key as a feature.
2. **Adapter versus encoder.** The CROMA sensor adapter (`specialists/optical_sar/sensor_adapter.py`)
performs band mapping, zero-fill and availability masking. CROMA the encoder does **not** receive the
availability mask (finding C-1). The mask is consumed by the fusion head. These are separate contracts
and Β§8.9 treats the distinction in full.
3. **Presentation versus radiometry.** There are two different percentile stretches in this codebase:
a per-*scene* 2/98 display stretch in `preprocessing/imagery.py` (and a grayscale cousin in
`specialists/change/postprocess.py`), and a per-*channel* `mean Β± 2Β·std` encoder-input stretch in
`specialists/optical_sar/radiometry.py`. They serve different purposes and are **not**
interchangeable. Β§7.5 is explicit about why.
The status vocabulary is used on every substantive claim. `MEASURED` means a number was produced by
executing code; `VERIFIED` means the claim was checked against a second independent source; `IMPLEMENTED`
means the code exists on the serving path; `DECLARED (not read)` means the value lives in
`configs/base.yaml` and no code reads it.
---
## 1. The raster input contract
### 1.1 What arrives, and what is accepted
An asset reaches the analysis path through `POST /v1/assets`, which mints an opaque `asset_id`. The
gateway enforces a **closed content-type allowlist of exactly five types**
(`docs/API_CONTRACT.md` Β§2.5):
| Content type | Notes |
|---|---|
| `image/tiff` | **The type the geospatial specialists need.** A client that uploads only PNG/JPEG can serve `vqa`, `caption` and `grounding`, but not `change` or `optical_sar` |
| `image/geotiff` | Accepted; media-type parameters are ignored, so `image/tiff; charset=binary` is accepted |
| `image/png` | Accepted |
| `image/jpeg` | Accepted |
| `application/octet-stream` | Accepted, but the raster reader still has to be able to open the bytes |
Two behaviours are specified rather than defaulted: a request with **no** declared content type is
**refused** rather than defaulted (defaulting is how a PDF reaches a raster reader), and a disallowed
type gets `415`. The stored filename is derived from the **content type**, never from the client's
filename, so a client-supplied `../../` cannot influence where bytes land (`docs/API_CONTRACT.md` Β§2.5).
The size cap is enforced at two layers β the gateway and the Space β and both now refuse an unparsable
or non-positive value rather than silently falling back to a default
(`docs/API_CONTRACT.md` Β§2.5, "Size cap").
`artifact_ref` on every `Evidence` item is **always `null` in v1** (`core/schemas.py:225`): v1 exposes no
artifact-serving endpoint, so a reference a client cannot retrieve is not fabricated. The schema field's
own description states this β *"Reference to an externally retrievable artifact. NEVER a filesystem path
(F-16, owner ruling 2026-09-23)"* β and the change and optical-SAR specialists both emit `null` on every
path (Β§4.3).
### 1.2 The validation chain
`preprocessing/raster.py` is the first thing every uploaded asset meets. Its module docstring names the
chain it implements, verbatim:
```
file -> dimensions -> bands -> dtype -> CRS -> transform -> bounds -> nodata
-> modality -> temporal metadata
```
The module's three design rules are stated at the top of the file and are worth quoting because they
govern every failure mode below:
> * Never raise a bare exception. Every failure is a typed SatQueryError.
> * Never silently drop geospatial metadata. If the source had a CRS, the returned AssetMetadata says so.
> * A missing CRS degrades to non-geospatial mode; it does not abort.
### 1.3 `inspect_raster` β header-only validation
`inspect_raster` (`preprocessing/raster.py:96`) validates and describes a raster **without loading its
pixel data**. Its exact signature:
```python
def inspect_raster(
path: str | Path,
*,
max_pixels: int | None = None,
explicit_modality: str | None = None,
compute_hash: bool = False,
sensor: SensorDescriptor | None = None,
) -> AssetMetadata:
```
The behaviour, field by field:
- The path is checked with `Path(path).exists()` and `is_file()`; a missing path raises `RasterReadError`.
- `rasterio` is imported lazily; an `ImportError` raises `RasterReadError` rather than a bare error.
- Inside `rasterio.open(p)`, the reader extracts `width`, `height`, `count`, the sorted set of `dtypes`,
the CRS as a string (`ds.crs.to_string() if ds.crs else None`), the affine transform truncated to its
first six coefficients (`list(ds.transform)[:6] if ds.transform else None`), the bounds
(`list(ds.bounds)`), `nodata`, `resolution` (`list(ds.res)`), the `driver` string, and a tiled flag
read through `_safe_is_tiled`.
- `_safe_is_tiled` (`preprocessing/raster.py:42`) exists because `DatasetReader.is_tiled` is scheduled for
removal; it suppresses the pending deprecation and **degrades to `False`** rather than aborting, since
the flag is informational metadata only.
- Degenerate dimensions (`width <= 0 or height <= 0`) raise `RasterReadError`.
- If `max_pixels` is given and `width * height > max_pixels`, it raises `OversizedImageError` β and the
docstring records that this is **recoverable**: *"the caller may downscale"*.
- A band count `<= 0` raises `UnsupportedBandsError`.
- Anything else rasterio raises is wrapped: `RasterReadError(f"could not open raster '{p}': {exc}")`.
The return is an `AssetMetadata` (see Β§1.4) with `modality` from `infer_modality` (Β§6) and, when
`compute_hash=True`, the streaming SHA-256 of the file (`file_sha256`, `preprocessing/raster.py:82`,
chunked at 1 MiB).
`inspect_raster` is called from `core/controller.py::_inspect_asset` during the VALIDATE state, with the
configured pixel budget passed through. That means the **header-only** nature of the inspection is a real
property of the serving path: a request never decodes a raster just to learn its shape.
### 1.4 `AssetMetadata` and `GeoMetadata`
`GeoMetadata` (`core/schemas.py:120`) is the geospatial detail carrier. It is `extra="allow"`, which is
deliberate β the optical-SAR specialist adds cross-sensor keys to it (Β§4.3):
| Field | Type | Meaning |
|---|---|---|
| `crs` | `str \| None` | CRS as a string (e.g. `EPSG:4326`), or `None` |
| `transform` | `list[float] \| None` | Affine coefficients, row-major, six values |
| `bounds` | `list[float] \| None` | `[minx, miny, maxx, maxy]` |
| `width`, `height` | `int \| None` | Raster dimensions in pixels |
| `band_count` | `int \| None` | Number of bands |
| `dtype` | `str \| None` | Comma-joined distinct dtypes when the bands disagree |
| `nodata` | `float \| None` | The nodata sentinel, or `None` |
| `resolution` | `list[float] \| None` | `(xres, yres)` |
| `has_crs` | `bool` | `crs_value is not None` |
| `is_georeferenced` | `bool` | `bool(crs_value and transform)` |
`AssetMetadata` (`core/schemas.py:153`) is `extra="forbid"` and carries:
| Field | Type | Meaning |
|---|---|---|
| `asset_id` | `str` | Minted `asset_<12 hex>`, opaque |
| `path` | `str` | Resolved local path |
| `modality` | `Modality` | From `infer_modality`, or declared |
| `sha256` | `str \| None` | Present only when `compute_hash=True` |
| `geo` | `GeoMetadata` | The block above |
| `sensor` | `SensorDescriptor \| None` | Present only when a sensor descriptor was supplied |
| `acquisition_date` | `str \| None` | Temporal metadata |
| `scene_id`, `dataset_id`, `split` | `str \| None` | Provenance for leakage control |
Note the distinction between `has_crs` and `is_georeferenced`: a raster can carry a CRS but no affine
transform (or vice versa), and both are recorded. Only when **both** are present is
`is_georeferenced=True`. `geospatial/transform.py::to_geo` requires the transform, so a raster with
`has_crs=True` but `transform=None` cannot be converted to geographic coordinates β it will raise
`CoordinateError("cannot convert to geo: raster has no affine transform")`.
### 1.5 Band reading, writing, and pixel area
`read_bands` (`preprocessing/raster.py:188`) returns `(bands, H, W)` plus a profile that **retains**
`crs`, `transform` and `bounds` as plain values, so downstream code never loses georeferencing. The
profile is a copy of `ds.profile` with those three keys overwritten from the live dataset.
`write_raster` (`preprocessing/raster.py:221`) writes an array using a profile captured by `read_bands`,
and re-applies `crs` and `transform` explicitly in the `rasterio.open(..., "w", ...)` call. It is used by
evidence generation so change maps stay georeferenced β **except** when the pair is suppressed for poor
registration, in which case the transform is deliberately dropped (Β§9.1).
`pixel_area_m2` (`preprocessing/raster.py:259`) returns the area of one pixel in square metres **when the
CRS is metric**, and `None` otherwise:
```python
def pixel_area_m2(meta: AssetMetadata) -> float | None:
from geospatial.crs import is_metric
if not is_metric(meta.geo.crs):
return None
if not meta.geo.resolution or len(meta.geo.resolution) < 2:
return None
xres, yres = meta.geo.resolution[0], meta.geo.resolution[1]
return abs(xres * yres)
```
The docstring is explicit: *"Returns None for geographic CRS or unknown resolution β callers must not
guess."* A degree-based CRS has no meaningful pixel area in metres, so the function refuses rather than
returning a plausible-looking number.
---
## 2. CRS handling (`geospatial/crs.py`)
CRS comparison is kept **separate** from pixel arithmetic so that *"can these two rasters be compared at
all?"* is answerable without touching pixels (`geospatial/crs.py` docstring). The module depends on
`pyproj`.
### 2.1 `parse_crs`
```python
def parse_crs(value: str | None) -> CRS | None:
if value is None or value == "":
return None
try:
return CRS.from_user_input(value)
except PyProjCRSError as exc:
raise MissingCRSError(f"unparseable CRS '{value}': {exc}") from exc
```
Absent is `None`; malformed raises. That distinction matters: a missing CRS is a legitimate state
(degrades to non-geospatial), while a *malformed* one is a defect that must surface.
### 2.2 `describe_crs`
`describe_crs` returns a dict for the execution trace: `present`, `srs`, `name`, `is_geographic`,
`is_projected`, and `axis_units` (the first axis's `unit_name`). For an absent CRS it returns
`{"present": False}`.
### 2.3 `compare_crs` and `CRSCompatibility`
`CRSCompatibility` (`geospatial/crs.py:19`) is a frozen dataclass with `compatible`, `identical`,
`same_units`, `left`, `right`, `reason`, and a derived property:
```python
@property
def requires_reprojection(self) -> bool:
return self.compatible and not self.identical
```
`compare_crs` (`geospatial/crs.py:59`) decides:
- **A missing CRS on either side is NOT compatible.** The reason string is
`"one or both rasters lack a CRS; spatial comparison is unsafe"`. The docstring records that this is
*"not fatal for non-spatial tasks"*.
- **Identical CRS** (`crs_l.equals(crs_r)`) returns `identical=True`, reason `"identical CRS"`.
- **Same kind** (both projected or both geographic) returns `compatible=True, identical=False`, reason
`"reprojection required (<left name> -> <right name>)"`.
- **Mixed projected/geographic** returns `compatible=True, identical=False`, with a reason that adds
*"reprojection required and should be verified"*. The comment records the rationale: *"Mixing a
projected CRS with a geographic one is legal to reproject but signals a pipeline mistake."*
`same_units` compares the first axis's `unit_name` on each side (defaulting to `"unknown"` when there is
no axis info).
**Crucially: `compare_crs` never reprojects.** It reports whether reprojection is *required*; the actual
transform is not implemented anywhere (Β§13).
### 2.4 `is_metric`
```python
def is_metric(crs_value: str | None) -> bool:
crs = parse_crs(crs_value)
if crs is None:
return False
if not crs.is_projected:
return False
if not crs.axis_info:
return False
return crs.axis_info[0].unit_name in {"metre", "meter", "m"}
```
A geographic CRS is never metric here, and the unit check is a closed set of three spellings.
### 2.5 Where CRS comparison is used
The only consumer on the serving path is the change specialist's `_assess_pair`
(`specialists/change/specialist.py:251`):
```python
compat = compare_crs(t1.geo.crs, t2.geo.crs)
if not compat.compatible:
warnings.append(
f"CRS comparison unavailable ({compat.reason}); change is "
f"reported in normalized pixel coordinates only"
)
elif compat.requires_reprojection:
warnings.append(
f"T1 and T2 use different CRS ({compat.reason}); change is "
f"valid only if the rasters already share a pixel grid"
)
```
The comment above it states the policy precisely: *"No CRS on either side is NOT fatal for pixel-domain
change detection β the two rasters still tile to the same grid. It IS fatal for any claim about ground
coordinates, so the warning is recorded and the geospatial block is withheld downstream."* The
withholding happens in `ChangeSpecialist._geospatial`, which drops the transform when registration is
unusable (Β§9.1).
---
## 3. Affine transforms and coordinate frames
`geospatial/transform.py` is where a normalised 0β1 box, a pixel box and a geographic box are converted
between. Its module docstring states the non-negotiable rules (citing `docs/ARCHITECTURE_FREEZE.md`
Β§2.6):
> * CRS and affine transform are preserved, never silently stripped.
> * Conversions are explicit about their source and target frames.
> * A conversion that cannot be performed raises, rather than returning a plausible lie.
That last line is the module's governing principle, and it is why `to_normalized`, `to_pixel` and
`to_geo` all raise `CoordinateError` when they lack the context a conversion needs.
### 3.1 `PixelWindow` and `GeoBounds`
`PixelWindow` (`geospatial/transform.py:29`) is a frozen dataclass of `[col_min, row_min, col_max,
row_max]`. Its `__post_init__` **raises** on an inverted window:
```python
if self.col_max < self.col_min or self.row_max < self.row_min:
raise CoordinateError(f"pixel window is inverted: {self.as_list()}")
```
It exposes `width`, `height` and `area` properties. `GeoBounds` (`geospatial/transform.py:60`) is the
geographic cousin: `(minx, miny, maxx, maxy)`.
### 3.2 Normalised β pixel
`normalized_to_pixel(box, width, height)` multiplies each coordinate by the corresponding dimension. Its
docstring records a deliberate choice:
> Values are NOT clipped: a box slightly outside the frame is preserved so the caller can decide whether
> that is a bug or a legitimate edge case.
It raises `CoordinateError` if `len(box) != 4` or if either dimension is `<= 0`.
`pixel_to_normalized` is the inverse, dividing by the dimensions.
### 3.3 Pixel β geographic
`pixel_to_geo(window, transform)` applies the affine transform to two corners. The docstring explains the
axis convention:
> rasterio/affine transforms map (col, row) -> (x, y). The y axis usually points down in pixel space, so
> the geographic miny comes from the BOTTOM row.
It uses the non-deprecated matmul form `transform @ (col, row)` and then takes `min`/`max` across the two
corners so the returned `GeoBounds` is ordered. It raises `CoordinateError` if the transform is `None`.
`geo_to_pixel(bounds, transform)` inverts the transform with `~transform`, raising `CoordinateError` if
the inverse cannot be computed (`"affine transform is not invertible: ..."`).
### 3.4 Schema-aware conversions
`to_normalized`, `to_pixel` and `to_geo` operate on `Box` objects and are schema-aware β they read
`box.coordinate_system` and return a `Box` with the target system set. The context each requires is
documented in the docstring:
- `pixel -> normalized` requires `width` and `height`.
- `geo -> normalized` requires `width`, `height` **and** the affine `transform`.
Each returns early when the box is already in the target frame, and otherwise uses
`box.model_copy(update={...})` to produce a new box with the new coordinates and the new
`coordinate_system`.
### 3.5 Intersection and IoU
`intersect(a, b)` returns the axis-aligned intersection of two `PixelWindow`s, or `None` when they are
disjoint. `iou(a, b)` computes intersection-over-union in `[0, 1]`, returning `0.0` for a disjoint pair
or a zero-area union. These are used by the grounding NMS path and by change-region overlap.
### 3.6 `affine_from_metadata`
```python
def affine_from_metadata(geo: GeoMetadata) -> Affine | None:
if not geo.transform or len(geo.transform) != 6:
return None
a, b, c, d, e, f = geo.transform
return Affine(a, b, c, d, e, f)
```
A `GeoMetadata` whose transform is missing or not exactly six values yields `None`; callers treat that as
"cannot georeference" rather than as an error. This is the bridge the grounding specialist uses in
`_to_geo_boxes` (Β§4.3).
---
## 4. The coordinate-system contract
### 4.1 The enum
`CoordinateSystem` (`core/schemas.py:57`) has exactly three members, and the class docstring is a warning
in itself β *"Never omit this. A bare box is meaningless without it. (C-5)"*:
```python
class CoordinateSystem(str, Enum):
NORMALIZED_0_1 = "normalized_0_1"
PIXEL = "pixel"
GEO = "geo"
```
`docs/API_CONTRACT.md` Β§3.2 fixes the meaning of each value:
| Value | Meaning | Rendering |
|---|---|---|
| `normalized_0_1` | `[0, 1]`, origin **top-left** | Multiply by image width/height |
| `pixel` | Absolute pixel coordinates | Use directly |
| `geo` | CRS coordinates (usually EPSG:4326) | Requires a map, not a 2-D canvas |
The contract adds a hard warning: *"The frontend must read the `coordinate_system` field on each
`Box`/`Region` and must not assume one convention. A box drawn with the wrong assumption lands in
plausible-looking wrong places."* It also records that earlier drafts used shorthand (`normalized`,
`geographic`) and that those strings are **wrong** β they would fail schema validation, because the
models are `extra="forbid"` and the enum is closed.
### 4.2 Why it is mandatory β the schema validator
The `Box` model (`core/schemas.py:171`) defaults `coordinate_system` to
`CoordinateSystem.NORMALIZED_0_1`, and its `_ordered` validator rejects a box whose coordinates do not
satisfy `x2 >= x1` and `y2 >= y1`. `Region` (`core/schemas.py:189`) and `ChangeRegion`
(`core/schemas.py:202`) both carry the same field with the same default.
The enforcement that matters is on `Evidence`. `Evidence._spatial_needs_crs`
(`core/schemas.py:238`) **rejects a spatial evidence item that carries coordinates but no
coordinate system**:
```python
@model_validator(mode="after")
def _spatial_needs_crs(self) -> "Evidence":
spatial = {
EvidenceType.BOUNDING_BOX,
EvidenceType.MASK,
EvidenceType.CHANGE_MAP,
EvidenceType.TILE,
EvidenceType.IMAGE_CROP,
EvidenceType.JOINT_FEATURE_REGION,
}
if self.type in spatial and self.coordinates and self.coordinate_system is None:
raise ValueError(
f"evidence type '{self.type.value}' carries coordinates "
"but no coordinate_system"
)
return self
```
So a bare box cannot be published as evidence. The validator fires only when the item both is of a
spatial type **and** carries coordinates β a `STATISTIC` or `GEOLOCATION` item with no coordinates is
unaffected, and a spatial item with no coordinates (e.g. an availability-mask record) is also unaffected.
`GEOLOCATION` is deliberately **not** in the spatial set: a geolocation is a point reference, not a
rectangle, and the grounding specialist sets its `coordinate_system` to `GEO` explicitly anyway.
Note the interaction with `extra="forbid"`: `Evidence` forbids extra keys, so an author cannot smuggle a
convention through an undeclared field. `GeoMetadata`, by contrast, is `extra="allow"`, which is how the
optical-SAR specialist attaches cross-sensor keys (Β§4.3).
### 4.3 What each specialist emits
| Specialist | Boxes / regions | Evidence coordinate systems |
|---|---|---|
| `grounding` | `Box` and `Region` in `NORMALIZED_0_1` | `BOUNDING_BOX` in `NORMALIZED_0_1`; a `GEOLOCATION` item in `GEO` when the CRS allows; `STATISTIC` items with no coordinates |
| `change` | `ChangeRegion` and `Region` in `NORMALIZED_0_1` | `CHANGE_MAP` and `BOUNDING_BOX` in `NORMALIZED_0_1`; `STATISTIC` items with no coordinates |
| `optical_sar` | No `regions` (empty list) | `AVAILABILITY_MASK` items (no coordinates); `OPTICAL_VIEW` / `SAR_VIEW` with `coordinates=[0,0,1,1]` in `NORMALIZED_0_1`; a `JOINT_FEATURE_REGION` in `NORMALIZED_0_1`; `STATISTIC` items |
**Grounding.** `GroundingSpecialist.execute` (`specialists/grounding/specialist.py:253`) builds every box
with an explicit system:
```python
boxes = [
Box(
x1=d.box[0], y1=d.box[1], x2=d.box[2], y2=d.box[3],
score=d.score, label=request.query,
coordinate_system=CoordinateSystem.NORMALIZED_0_1,
)
for d in decoded
]
```
Geographic boxes are **derived, never replacements**: `_to_geo_boxes` (`specialists/grounding/specialist.py:431`)
converts the normalised boxes to geographic bounds *only* when the raster has a CRS, a transform and
dimensions, and emits a `GEOLOCATION` evidence item with `coordinate_system=CoordinateSystem.GEO`. When
the raster has no CRS/transform, the method appends a warning β
*"raster has no CRS/transform; boxes are reported in normalized coordinates only and cannot be placed on a
map"* β and returns an empty list. The result's boxes stay normalised; the geo boxes are additional
evidence for a consumer with a CRS. Its own comment states the rule: *"Pixel and geo coordinates are
DERIVED, never replacements."*
**Change.** `regions_to_schema` (`specialists/change/postprocess.py:443`) builds `ChangeRegion` objects in
the requested system. It supports `NORMALIZED_0_1` and `PIXEL` and **raises** for anything else:
```python
else:
raise ValueError(
f"regions_to_schema cannot produce {coordinate_system.value} "
f"coordinates; convert after this call"
)
```
The change specialist calls it with the default (`NORMALIZED_0_1`). Geographic conversion is a deliberate
post-step β the docstring records: *"Callers that want geographic coordinates convert afterwards via
`geospatial.transform`, which keeps this function free of rasterio and CRS concerns."*
**Optical-SAR.** The specialist emits no `regions`. Its view evidence items carry
`coordinates=[0.0, 0.0, 1.0, 1.0]` in `NORMALIZED_0_1` β the whole frame β because the view is a
presentation of the canonical tensor, not a localisation. Its `geospatial` block is the **optical**
asset's georeferencing (chosen because it is the asset a reader can navigate by), with cross-sensor keys
added via `GeoMetadata`'s `extra="allow"`: `optical_crs`, `sar_crs`, `sar_bounds`, `sar_resolution`,
`optical_availability`, `sar_availability`, `optical_sensor`, `sar_sensor`, and
`resolution_ratio_sar_to_optical` when both resolutions are known
(`specialists/optical_sar/specialist.py:597`). The comment records the reason the two footprints are
carried separately rather than fused: *"RISAT and Cartosat-2S are different missions with different
orbits."*
---
## 5. Benchmark box scale β 0β100 versus 0β1
This is a small contract with an outsized failure mode, and the codebase treats it as such.
### 5.1 The two conventions
- **VRSBench** (the grounding benchmark) states verbatim that *"all box coordinates are normalized to
0-100"*.
- **This project** stores every box normalised to **0β1**.
Treating a 0β100 box as a pixel box, or as a 0β1 box, is a silent, catastrophic bug: the numbers are all
in range, the schema accepts them, and the box lands in the wrong place.
### 5.2 The conversion functions
`geospatial/transform.py:286` provides the conversion, and its docstring names the failure mode it exists
to prevent:
```python
def benchmark_boxes_to_normalized(
boxes: Iterable[Sequence[float]], scale: float = 100.0
) -> list[list[float]]:
"""Convert benchmark boxes (e.g. VRSBench's 0-100) to our internal 0-1.
Finding from Phase 0: VRSBench states verbatim that "all box coordinates are
normalized to 0-100". Treating those numbers as pixels is a silent, catastrophic
bug β this function exists so that mistake can be made once, explicitly.
"""
if scale <= 0:
raise CoordinateError(f"box scale must be positive, got {scale}")
...
out.append([float(v) / scale for v in b])
```
`normalized_boxes_to_benchmark` is the exact inverse (multiply by `scale`). Both raise on a
non-positive scale and on a box that is not four values.
A second, independent conversion lives in the evaluation metrics:
`evaluation/metrics/grounding.py::benchmark_to_normalized(box, scale=100.0)` divides by the scale and
raises `ValueError(f"scale must be positive, got {scale}")` on a non-positive value. `docs/API_CONTRACT.md`
Β§3.2 names it explicitly as the site of the VRSBench conversion.
### 5.3 The declared scale
`configs/base.yaml` declares the conversion factor in the grounding block, with a comment that restates
the convention:
```yaml
grounding:
...
# VRSBench stores boxes normalised to 0-100; we store 0-1.
benchmark_box_scale: 100.0
coordinate_system: normalized_0_1
```
`benchmark_box_scale: 100.0` is therefore the **declared** value, and `100.0` is also the **default
argument** in both conversion functions. The two agree.
### 5.4 The double-application risk
Because the scale appears in three places β the config key, the `geospatial.transform` default and the
`evaluation.metrics.grounding` default β there is a real hazard that a caller converts 0β100 β 0β1 and
then a downstream stage divides by 100 again, producing a box 100Γ too small in each coordinate. The
codebase's mitigations are structural rather than documentary:
1. **The conversion is named for its direction.** `benchmark_boxes_to_normalized` and
`normalized_boxes_to_benchmark` are inverses with self-describing names, so a reader cannot easily
apply the wrong one without noticing.
2. **The default is the correct scale, not `1.0`.** A caller that forgets to pass `scale` gets `100.0`,
which is right for VRSBench and would produce an obviously wrong box (values > 1) for already-0β1
input β a loud failure, not a silent one.
3. **The schema rejects out-of-range values indirectly.** A `Box` built from un-divided VRSBench
coordinates would have `score` in range but coordinates in `[0, 100]`; the `_ordered` validator would
pass it (the ordering is still valid), so the schema alone does **not** catch this. The catch is that
the metrics module's own `benchmark_to_normalized` is the documented entry point, and the resolution
experiment and the grounding evaluation both go through it.
**Residual risk β stated honestly.** There is no runtime assertion that a `Box` in
`NORMALIZED_0_1` actually has coordinates within `[0, 1]`. `Box._ordered` checks only ordering. A
mis-scaled box would therefore pass schema validation. `UNKNOWN β not established from the available
evidence` whether any production call site currently constructs a `Box` without going through one of the
two conversion functions; the serving path (`GroundingSpecialist.execute`) builds boxes from the decoder's
output, which is already in `[0, 1]`, so the serving path is not exposed. The risk is confined to
evaluation and experiment code that reads VRSBench directly.
---
## 6. Band-count modality inference
### 6.1 The heuristics
`preprocessing/raster.py` defines two closed sets:
```python
_OPTICAL_BAND_COUNTS = {3, 4, 8, 11, 12, 13}
_SAR_BAND_COUNTS = {1, 2}
```
with the comment: *"Band-count heuristics for modality detection when explicit metadata is absent. These
are heuristics, not ground truth β the sensor adapter is authoritative when a sensor descriptor is
supplied."*
### 6.2 `infer_modality`
```python
def infer_modality(band_count: int, explicit: str | None = None) -> Modality:
if explicit:
try:
return Modality(explicit.lower())
except ValueError:
pass
if band_count in _SAR_BAND_COUNTS:
return Modality.SAR
if band_count in _OPTICAL_BAND_COUNTS:
return Modality.OPTICAL
# 12-band optical and 2-band SAR are both plausible; default to optical only
# when the count clearly favours it.
if band_count >= 4:
return Modality.OPTICAL
return Modality.UNKNOWN
```
The decision order is worth reading carefully:
1. An **explicit** label wins if it parses to a `Modality` member. An explicit label that does *not* parse
(e.g. `"radar"`) is silently ignored and inference proceeds β the `except ValueError: pass` is
deliberate.
2. `{1, 2}` β `SAR`.
3. `{3, 4, 8, 11, 12, 13}` β `OPTICAL`.
4. `>= 4` (and not already matched) β `OPTICAL`.
5. Otherwise β `UNKNOWN`.
So the effective mapping is: 1β2 β SAR; 3β13 β optical (by set or by the `>= 4` fallback); 0 or a
non-positive count β `UNKNOWN` (though `inspect_raster` rejects `band_count <= 0` before this point).
### 6.3 The declared-wins rule
In the optical-SAR specialist, `_assess_pair` (`specialists/optical_sar/specialist.py:223`) uses the
asset's **declared** modality first, falling back to band-count inference only when it is `UNKNOWN`:
```python
for asset in assets:
modality = asset.modality
if modality is Modality.UNKNOWN:
inferred = self._infer_modality(asset, warnings)
resolved.append(inferred)
else:
resolved.append(modality)
```
`_infer_modality` appends a warning that names the heuristic for what it is: *"A band count is a
heuristic, not a sensor declaration β label the asset explicitly if this is wrong."* The docstring above
`_assess_pair` states the rationale: *"The declared value wins when present, because the caller may know
something the band count does not β a 2-band Cartosat stack, or a 12-band decomposition product."*
### 6.4 Why the heuristic is not authoritative
The comment on the sets and the warning text both point at the same thing: a band count does not identify
a *sensor*. A 2-band file could be a dual-pol SAR pair or a two-band optical stack; a 4-band file could be
RGB+NIR or a four-polarisation SAR product. The adapter (Β§8) is the authoritative path, because it maps
named bands onto CROMA's canonical channels and records what it could not place.
---
## 7. Radiometric normalisation and representation
This section covers a genuine specification gap and its partial resolution. It is presented in full
because the gap is load-bearing and the honest state is more useful than a tidy summary.
### 7.1 What configuration declares
`configs/base.yaml` declares two radiometric transforms:
```yaml
optical:
normalization: percentile
lower_percentile: 2
upper_percentile: 98
# CROMA expects exactly 12 optical channels.
canonical_channels: 12
sar:
representation: db
clip_min_db: -30
clip_max_db: 5
# CROMA expects exactly 2 SAR channels (VV, VH).
canonical_channels: 2
```
So the declared optical normalisation is a **percentile stretch at the 2nd and 98th percentiles**, and the
declared SAR representation is **decibels clipped to [β30, +5]**.
### 7.2 The DEV-2 ruling β two ordered stages
`docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` records the ruling that governs this. The gap it confirmed:
CROMA's own example preprocessing is a **per-channel** dynamic-range stretch
(`mean Β± 2Β·std β [0, 1]`, optionally through a uint8 round-trip), while `configs/base.yaml` specifies
something structurally different. The ruling's verification table records that a grep of
`specialists/optical_sar/` for `NORM_*`, `clip_min_db`, `lower_percentile`, `use_8_bit` returned **six
hits, all constant definitions/labels** at `sensor_adapter.py:125,126,210,323,547,548`, and **zero
application sites**.
The ruling resolves this by declaring the pipeline to have **two ordered stages**, and it states
explicitly that they are *not alternatives*:
| Stage | Transform | Owner | Status |
|---|---|---|---|
| **Radiometric conditioning** | per-scene percentile / dB-clip, per `configs/base.yaml` | preprocessing / sensor adapter | **specified in the ruling, must be implemented** |
| **Encoder-input normalisation** | per-channel `mean Β± 2Β·std` β `[0,1]`, then optional `Γ255` uint8 β `/255` | a new function consumed by `croma.py` | **mandatory, immediate** |
The rationale for the CROMA stage being mandatory: *"it is the only stage that bounds the input range.
Without it the encoder receives unbounded input."* The `use_8_bit` branch is the safer default because
the uint8 round-trip **guarantees** `[0, 1]`, whereas the float path's clip applies to a stretch that is
unbounded by construction (`mean Β± 2Β·std` is not a min/max clamp). The decision on `use_8_bit` is
**ENABLED**.
### 7.3 The implemented stage β `specialists/optical_sar/radiometry.py`
The **encoder-input** stage is implemented, in full, at `specialists/optical_sar/radiometry.py`. Its
module docstring states the boundary between what is and is not implemented:
> Implemented: the **encoder-input** stage β per-channel `mean +/- 2*std` -> `[0, 1]`, the transform the
> CROMA authors' own README instructs users to apply.
>
> NOT implemented: the percentile / dB **conditioning** stage that `base.yaml` names.
The transform is quoted verbatim from the CROMA README in the docstring:
```
min_value = x[:, c].mean() - 2 * x[:, c].std()
max_value = x[:, c].mean() + 2 * x[:, c].std()
img = (x[:, c] - min_value) / (max_value - min_value) * 255.0
img = clip(img, 0, 255).to(uint8)
# before the forward pass:
x = x.float() / 255
```
The module makes two **deliberate deviations** from the README, both documented:
1. **Per sample, not per batch.** The README computes `x[:, c].mean()` over the whole batch, which makes
one image's encoding a function of its neighbours. This module computes the window **per sample**,
because *"a serving system whose whole point is reproducibility"* must not produce a different answer
because another request was batched alongside it. At `N = 1` the two are identical, and the README's
own example runs at `N = 1`.
2. **Unbiased standard deviation (`ddof=1`).** Upstream is torch, whose `.std()` defaults to the unbiased
estimator; NumPy's defaults to `ddof=0`. The module uses `ddof=1` to match the reference
implementation, and states the numerical immateriality (~1.00003 factor at 120Γ120) so nobody has to
guess which convention a number came from.
The key constants:
| Constant | Value | Meaning |
|---|---|---|
| `TRANSFORM_NAME` | `"per_channel_mean_pm_2std"` | Recorded on every trace |
| `WINDOW_SIGMAS` | `2.0` | The window half-width |
| `DEFAULT_USE_8_BIT` | `True` | The ruling's decision |
| `ENV_USE_8_BIT` | `"SATQUERY_CROMA_USE_8_BIT"` | Hash-exempt override channel |
| `MIN_FINITE_PIXELS` | `2` | Below this, a std is not meaningful |
| `STATUS_NORMALISED` / `STATUS_UNAVAILABLE` / `STATUS_DEGENERATE` | strings | Per-channel status |
`normalise_for_croma(array, *, use_8_bit, modality)` is **pure and deterministic** β the docstring
states: *"no sampling, no learned statistics, no RNG, and no dependence on batch composition. Same input
-> same output."* It accepts `(B, C, H, W)` or `(C, H, W)`, returns a `float32` array of the same shape
plus a `RadiometryReport`, and leaves skipped channels **exactly zero**.
### 7.4 The zero-channel rule
This is the load-bearing part, and it is the C-1 discipline applied to radiometry. An unavailable channel
is all zeros (the sensor adapter's convention). Such a channel has `std = 0`, so `max_value - min_value
== 0` and the stretch divides by zero. The module therefore **skips and leaves at exactly zero** any
channel with no dynamic range, and records a **different status** for each reason:
- `unavailable` β the channel is all zeros (`if not np.any(channel)`), i.e. the sensor did not measure it.
- `degenerate` β the channel is present but constant (`_channel_window` returns `None` because
`upper - lower <= 0`), i.e. the sensor measured a flat band.
- `normalised` β the real case.
The docstring explains why the distinction matters: *"'the sensor did not measure this' is not the same
fact as 'the sensor measured a flat band', and the trace must not merge them."* The mechanism is the
degenerate-window test, which needs **no mask** β deliberately, because `CROMAEncoder.encode` has a
pinned signature of exactly `(self, optical, sar)` and finding C-1 says CROMA never receives a mask.
Inferring unavailability from the data keeps that guard intact.
A non-finite pixel (NaN/Inf) is *"a defective measurement, not a measurement of zero"*; it is pinned to
`0.0` rather than allowed to propagate a NaN into the frozen encoder, where *"one NaN would poison the
whole forward pass."* The `finite_pixels` count in the report keeps the gap detectable.
`resolve_use_8_bit(config)` resolves the flag in a documented order β environment variable first, then
`croma.use_8_bit` when the key exists and is a bool, then `DEFAULT_USE_8_BIT` β and returns
`(value, source)` where `source` is `"env"`, `"config"` or `"default"`. The reason it is not simply a
config key is recorded: adding `croma.use_8_bit` to `configs/base.yaml` would move `Config.hash` away
from the frozen `78f1e3700da15aa1` and detach the Phase 9 benchmark. The hash-exempt channel is the same
pattern already used for `SATQUERY_DEVICE`.
`summarise(optical_report, sar_report)` bundles both modalities under one `use_8_bit` and **asserts** they
agree, raising `ValueError` if they do not: *"upstream applies one transform to both."*
### 7.5 The display stretches β a different transform
Two other percentile stretches exist in the codebase. Neither is the CROMA transform, and the DEV-2
ruling names reusing one of them as **forbidden**.
**`preprocessing/imagery.py::load_image_array`** produces a displayable `(H, W, 3)` uint8 array:
```python
arr = array.astype(np.float32)
finite = arr[np.isfinite(arr)]
if finite.size:
lo, hi = np.percentile(finite, (2, 98))
if hi > lo:
arr = (arr - lo) / (hi - lo)
arr = np.clip(arr, 0.0, 1.0)
return (arr * 255).astype(np.uint8)
```
Its docstring is explicit about its scope: it *"does not resample, crop, or reproject. Those change the
pixel grid, and the grounding specialist converts normalized boxes to pixel coordinates using the
ORIGINAL raster's dimensions β a silent resize here would put every box in the wrong place."* The
determinism claim is stated: *"the same file always yields the same array, so a grounding box and a VQA
answer describe identical pixels."*
`docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§1.1 records the three disqualifying differences between
this stretch and CROMA's: it is **per-scene** (one `(lo, hi)` over all bands) rather than **per-channel**;
it returns **uint8 3-band** rather than float 12-channel; and its **purpose** is presentation, not
radiometry.
**`specialists/change/postprocess.py::to_grayscale_float`** applies a per-scene 2/98 percentile stretch
to a grayscale collapse, so that *"a uint16 raster and a uint8 raster of the same scene correlate
identically."* Its comment records a measured defect fixed in place: `np.percentile` returns float64
scalars even for float32 input, so the array is promoted; without the explicit
`.astype(np.float32, copy=False)` the function returns float64 and `cv2.phaseCorrelate` fails its own
type assertion against a `CV_32F` Hanning window β *"Measured, not theorised."*
### 7.6 The unimplemented stage, and its consequence
The percentile/dB conditioning stage is **not implemented**. The radiometry module's docstring states the
consequence plainly:
> **Consequence: the two `optical.*` and three `sar.*` config keys are still read by no code.** That is
> recorded, not fixed.
`docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§6 item 3 adds: *"A config key that no code reads is worse
than a missing key: it reads as a satisfied requirement."* The ruling's follow-on table assigns the
implementation to the engineer, with the alternative of explicitly retiring the keys.
The status of the whole chain is therefore:
| Stage | Declared | Implemented | Read by code |
|---|---|---|---|
| Optical percentile 2/98 conditioning | β
`optical.*` | β | β |
| SAR dB clip β30/+5 conditioning | β
`sar.*` | β | β |
| Per-channel `mean Β± 2Β·std` β [0,1] | β
(ruling Β§3.1) | β
`radiometry.py` | β
(via `training/fusion/extract.py`; see below) |
**Where the implemented stretch actually runs.** The serving-path pipeline
(`specialists/optical_sar/inference.py::run_pipeline`) does **not** call `normalise_for_croma`. It
canonicalises, resizes to 120Γ120 and calls `encoder.encode(...)`. The stretch is applied by the
**extraction** path (`training/fusion/extract.py`), which states: *"This pipeline applies the DEV-2
encoder-input stretch itself (`radiometry.normalise_for_croma`), because the arm's `use_8_bit` has to be
applied at exactly one place."* `extract.py` also refuses an encoder that would stretch a second time:
`_assert_raw_encoder` *"refuses it by name rather than trusting the caller."* So the implemented
radiometric transform is a **training-time** transform that feeds the cached features, and the serving
path's input scaling is whatever `CROMAEncoder.encode` does internally (see Β§8.9 and the CROMA module).
`UNKNOWN β not established from the available evidence` whether the serving path's encoder applies the
same `mean Β± 2Β·std` stretch internally; `specialists/optical_sar/croma.py::CROMAEncoder.encode` was
verified to have the pinned signature `(self, optical, sar)` with no mask parameter, and the extraction
module's guard implies the encoder *can* apply a stretch when `normalize_input=True` (its default), but
the exact serving configuration of that flag is not established from the files read for this chapter.
---
## 8. The CROMA sensor adapter contract
`specialists/optical_sar/sensor_adapter.py` is *"the one place where 'some sensor's bands' becomes
'CROMA's canonical channels', and it exists so that translation is explicit, inspectable and testable."*
Its module docstring states **the two hard rules** from the architecture plan:
> "No invented missing bands."
> "Do not fabricate spectral bands."
The failure mode the module prevents is described with unusual clarity:
> A 4-band Cartosat-2S scene maps cleanly onto canonical channels 1-4. Filling channels 5-12 with
> *something* β a copy of B3 as a fake "red-edge", a zero-order hold, an interpolation β produces a tensor
> whose shape is correct and whose statistics look plausible. CROMA would run. Nothing would raise. And
> every downstream number would be computed from channels that no sensor ever measured, presented as if
> they were measurements.
So: **missing channels are zero-filled, and the availability mask says which ones were real.**
### 8.1 `SensorDescriptor`
The schema object (`core/schemas.py:138`) is `extra="forbid"`, described as the *"Frozen sensor-adapter
contract (plan section 18)"*:
| Field | Type | Default | Meaning |
|---|---|---|---|
| `sensor` | `str` | required | Sensor identifier, e.g. `"cartosat_2s"` |
| `available_bands` | `list[str]` | `[]` | The bands the sensor actually delivers |
| `band_map` | `dict[str, str]` | `{}` | Source band name β canonical band name |
| `normalization` | `str` | `"percentile"` | Identifier recorded for the trace |
| `availability_mask` | `list[bool]` | `[]` | Which canonical channels are real |
| `resolution` | `float \| None` | `None` | Ground sample distance in metres |
| `canonical_channels` | `int` | `0` | Target channel count (12 optical, 2 SAR) |
| `missing_channels_zero_filled` | `bool` | `True` | The zero-fill convention |
**Why `band_map` is canonical-channel-keyed.** The module docstring addresses a direction question: the
plan's example writes `band_map` sensor-keyed (`{"B1": "c1", ...}`), and the schema types it as
`dict[str, str]`. This module uses the schema's direction because the question asked at load time is
forward-looking: *"canonical channel 3 β did I get a real band for it, and if so which one?"*
`canonical_to_sensor` is exposed as the inverse for reporting.
### 8.2 The canonical channel orders
```python
OPTICAL_CANONICAL: tuple[str, ...] = (
"B01", # coastal aerosol
"B02", # blue
"B03", # green
"B04", # red
"B05", # red edge 1
"B06", # red edge 2
"B07", # red edge 3
"B08", # NIR
"B8A", # narrow NIR
"B09", # water vapour
"B11", # SWIR 1
"B12", # SWIR 2
)
SAR_CANONICAL: tuple[str, ...] = ("VV", "VH")
```
The comment above `OPTICAL_CANONICAL` states the contract: *"The ORDER is part of the contract: channel
index i must always mean the same physical band, or a pretrained encoder's weights would be applied to
the wrong input."* This is Sentinel-2 L2A with the cirrus band removed, per CROMA's README: *"Sentinel-2
must be 12 channels (remove the cirrus band if necessary)."* Note that `B10` (cirrus) is **not** in the
canonical order β that is the removal.
`SAR_CANONICAL` is described as *"Sentinel-1 dual-pol order CROMA was pretrained with. Positional fallback
ONLY."*
`_BAND_ALIASES` maps common spellings onto the canonical form: `B1`β`B01`, `BLUE`β`B02`, `NIR`β`B08`,
`SWIR1`β`B11`, `VVPOL`β`VV`, and so on. The comment is a caution: *"Kept small and explicit: a guess here
becomes a silently mislabelled band."* `_canonicalise` strips whitespace and underscores and uppercases
before lookup, and **never invents** a band β an unmapped name passes through unchanged and is later
rejected.
### 8.3 Building an optical adapter
`build_optical_adapter(sensor, *, available_bands, canonical_channels=12, normalization="percentile",
resolution=None)` (`specialists/optical_sar/sensor_adapter.py:205`):
- When `available_bands` is `None`, the sensor is looked up in `_KNOWN_OPTICAL_SENSORS`; an unknown
sensor with no explicit band list raises `UnsupportedBandsError` β *"refusing to guess a band
arrangement."*
- Each band is canonicalised. Any canonical name **not** in `OPTICAL_CANONICAL` raises
`UnsupportedBandsError`: *"do not map onto the canonical Sentinel-2 order ... supply an explicit
band_map rather than guessing a position."*
- Duplicate bands raise.
- More bands than `canonical_channels` raises.
- `band_map` is built as `{source: canonical for source, canonical in zip(bands, canonical)}` β keyed by
the **source** name, because *"that is what an array row can actually be looked up by."*
- `availability` is `[canonical_name in band_map.values() for canonical_name in OPTICAL_CANONICAL[:canonical_channels]]`.
The returned descriptor has `missing_channels_zero_filled=True`.
`_KNOWN_OPTICAL_SENSORS` is *"deliberately tiny and explicit. Adding one is a code change, not a guess at
runtime"*:
| Sensor | Bands |
|---|---|
| `sentinel_2` | All 12 canonical |
| `cartosat_2s` | `["B1", "B2", "B3", "B4"]` |
| `cartosat_2` | `["B1", "B2", "B3", "B4"]` |
| `landsat_8` | `["B1", "B2", "B3", "B4"]` |
The comment on `cartosat_2s` is a standing instruction: *"Cartosat-2S is a 4-band VNIR imager. It is NOT a
Sentinel-2 clone and must not be treated as one: the 8 channels it lacks stay masked-off."*
### 8.4 Describing a SAR sensor β inspect, do not assume
`describe_sar(sensor, *, available_bands, canonical_channels=2, normalization="db", resolution=None)`
(`specialists/optical_sar/sensor_adapter.py:318`) implements the plan's instruction:
> "RISAT imagery may vary in acquisition/polarization characteristics... the SAR adapter must inspect
> actual available channels rather than assuming one fixed polarization pair."
The module docstring explains why a hardcoded `["VV", "VH"]` would be wrong: *"ISRO describes RISAT-1 as
a C-band SAR mission with multiple polarization configurations (single HH, single VV, dual HH+HV, dual
VV+VH, ...). A hardcoded `["VV", "VH"]` would silently mislabel an HH/HV scene, and the resulting tensor
would be a *confidently wrong* input rather than a recognisably broken one."*
When `available_bands` names the polarisations, they are honoured. When it does not, the **channel count**
is used to select the most likely convention, and **the choice is recorded on the descriptor's
`available_bands`** so the assumption appears in the trace rather than being invisible.
The **slot order is the sensor's own**:
```python
band_map: dict[str, str] = {source: source for source in bands}
availability = [i < len(bands) for i in range(canonical_channels)]
```
The comment explains: *"Forcing HH/HV into a VV/VH frame would mean either relabelling the polarisations
(a lie the mask would then contradict) or masking both off and discarding real data."* So the tensor holds
the radar measurements that exist, and `available_bands[i]` says which polarisation slot `i` is.
`_known_sar(sensor)` is described as *"Best-effort polarisation layout for a named SAR sensor, or `[]`"*:
| Sensor key | Returns |
|---|---|
| `sentinel_1`, `sentinel1` | `["VV", "VH"]` |
| `risat*` | `["VV", "VH"]` β **recorded as an ASSUMPTION, not a fact** |
| `alos_palsar`, `palsar`, `alos2` | `["HH", "HV"]` |
| anything else | `[]` β *"An empty return is honest: it says 'I do not know this sensor's polarisations'"* |
The `risat` case carries an explicit comment: *"RISAT-1's most common operational mode is dual-pol.
Recorded as an ASSUMPTION, not a fact β see the module docstring. Callers with real band metadata should
pass `available_bands` and override this."* This is an `ASSUMPTION` in the project's status vocabulary,
and it is the reason the hidden-set behaviour on RISAT cannot be claimed as measured.
### 8.5 Applying the canonical layout β `_apply_canonical`
This is *"the function the two hard rules live in"* (`specialists/optical_sar/sensor_adapter.py:428`). It
places real bands in canonical positions and zero-fills the rest:
```python
canonical = np.zeros((n_canonical, height, width), dtype=np.float32)
mask = np.zeros((n_canonical,), dtype=bool)
for source_index, source_name in enumerate(descriptor.available_bands):
if source_index >= n_source:
continue
canonical_name = descriptor.band_map.get(source_name)
if canonical_name is None or canonical_name not in order:
continue
channel_index = order.index(canonical_name)
if channel_index >= n_canonical:
continue
canonical[channel_index] = array[source_index].astype(np.float32)
mask[channel_index] = True
```
The docstring's key claim: *"There is no branch anywhere below that writes a value into an unavailable
channel."* The absence of such a branch **is** the mask. If no band could be placed while the array
carries bands, it raises `UnsupportedBandsError` rather than returning an all-zero tensor silently.
`_canonical_order(descriptor)` decides the slot order: for a descriptor whose `available_bands` are all
SAR polarisations, the order is the descriptor's own list (padded to canonical width); otherwise it is
`OPTICAL_CANONICAL`.
`adapt_optical` and `adapt_sar` are thin wrappers that call `_apply_canonical` with the right modality
label. `adapt_sar`'s docstring names the plan's internal representation exactly:
`canonical_sar[2, H, W]` + `sar_channel_mask[2]`.
### 8.6 `SensorAdapterOutput`
The dataclass returned by `adapt_*` carries `canonical` `(C, H, W)` float32, `mask` `(C,)` bool, and the
`descriptor`. It exposes `n_channels`, `missing_indices` (`[i for i, ok in enumerate(self.mask) if not
ok]`), `available_bands`, and a `to_dict()` that includes `n_available` and `missing_channels`.
### 8.7 `positional_fallback_descriptor`
Used when a raster has the right **channel count** for a modality but **no band labels**. The fallback
order is chosen from `_FALLBACK_SAR_ORDER` (`("VV", "VH")`, `("HH", "HV")`) for 1β2 channels, or the
canonical optical prefix for β₯3 channels. The fact that it was a fallback is **recorded in the `sensor`
field** so it *"cannot be mistaken for measured metadata."* The `complement=True` flag selects the other
convention, which lets a caller test both HH/HV and VV/VH orderings without asserting which is correct.
The optical-SAR specialist uses this in `_descriptor_for`
(`specialists/optical_sar/specialist.py:398`) and appends a warning that names the assumption: *"The
polarisation IDENTITY of each channel is NOT known from a band count alone β supply a sensor descriptor
to state it."*
### 8.8 Finding C-1 β the mask goes to the fusion head, not to CROMA
This is the single most important fact in this section. `core/schemas.py`'s module docstring lists it
among the findings baked into the schemas: *"C-1 channel-availability mask is a first-class fusion input,
never a CROMA input."*
`specialists/optical_sar/fusion_head.py` devotes a section to the rationale:
> CROMA is a masked autoencoder. Handing it an availability mask invites it to reconstruct the missing
> channels β which is precisely the fabrication the sensor adapter exists to prevent: the model would
> output plausible values for bands no sensor measured, and those values would then be treated as data.
> The mask is therefore consumed HERE, where it can only do one thing: tell the classifier which inputs
> to distrust.
Mechanically: `CROMAEncoder.encode` has a pinned signature of exactly `(self, optical, sar)` β the
radiometry module records that `tests/unit/test_optical_sar_croma.py` *"asserts the absence of any mask
parameter, because freeze finding C-1 says CROMA never receives one."* The mask instead enters at
`assemble_fusion_input` (Β§8.9). The optical-SAR specialist's `JOINT_FEATURE_REGION` evidence item records
the fact explicitly in its payload: `"mask_consumed_by": "fusion_head"` and `"croma_received_mask":
False` (`specialists/optical_sar/specialist.py:816`).
`EvidenceType` even has a dedicated member for the mask as trust evidence:
`AVAILABILITY_MASK = "availability_mask" # C-1: modality trust evidence` (`core/schemas.py:76`). The
optical-SAR specialist emits one `AVAILABILITY_MASK` item per modality, carrying the sensor, canonical
channel count, available bands, band map, the mask itself, the missing indices,
`missing_channels_zero_filled`, `normalization` and `resolution`.
### 8.9 The frozen fusion concatenation
`specialists/optical_sar/fusion_head.py` defines the concatenation that the mask feeds into:
```
optical_GAP (B, 768)
SAR_GAP (B, 768)
joint_GAP (B, 768)
optical_mask (B, 12) <- availability, from the sensor adapter
sar_mask (B, 2) <- availability, from the sensor adapter
---------
concat (B, 2318)
```
`expected_fusion_dim(encoder_dim=768, optical_channels=12, sar_channels=2, modalities_used=3)` returns
`3*768 + 12 + 2 = 2318`. The module recomputes the width and **refuses to build on a mismatch**, so *"a
config edit cannot silently reshape the first Linear layer into something that trains but means
nothing."* `core/config.py` recomputes it independently at load time from `croma.encoder_dim`,
`croma.modalities_used`, `croma.optical_channels` and `croma.sar_channels`, and rejects a config whose
`fusion.input_dim` disagrees (finding C-1 in the validator's own message).
`assemble_fusion_input` asserts the order rather than assuming it β *"a permutation here produces a tensor
of exactly the right shape that trains to a worse number β the hardest kind of bug to notice"* β and
validates that the three GAP vectors agree on batch size and that both masks agree with it.
`channel_dropout(features, mask, *, keep_probabilities, rng)` is the mechanism that *"teaches the head to
trust the availability mask"*. Its load-bearing line is:
```python
keep = draws & (msk > 0.0)
out_mask = keep.astype(np.float32)
out = out * out_mask
```
Dropped channels are **zeroed, not renormalised** β *"renormalising would fabricate a scale that the real
missing-channel case does not have"* β and the mask is updated in lockstep. `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md`
Β§5 independently verified this against freeze Β§2.5 and ruled: *"the engineer's implementation SATISFIES
Β§2.5's intent. No change entry required."* Its table records that the `& (msk > 0.0)` guard is the
load-bearing one: *"no training transform can resurrect a band the sensor did not measure."*
The schedule (plan section 20) is optical at `100/80/60/40%` and SAR at `100/50%`, declared in
`training/fusion/train.py` as `OPTICAL_DROPOUT_RATES = (1.0, 0.8, 0.6, 0.4)` and
`SAR_DROPOUT_RATES = (1.0, 0.5)`.
`mask_availability_stats(optical_mask, sar_mask)` returns `optical_available_fraction`,
`sar_available_fraction`, `optical_channels_present` and `sar_channels_present`. These feed the
confidence system: the two modality-confidence terms in the optical-SAR specialist are exactly the
availability fractions, which the specialist's docstring justifies: *"On the hidden set the optical side
is Cartosat-2S with 4 bands against a canonical 12, so 8 of 12 channels are always absent. A classifier
resting on 4 measured channels is not the same claim as one resting on 12."*
### 8.10 The optical-SAR confidence components
`OpticalSarSpecialist._confidence_components` (`specialists/optical_sar/specialist.py:481`) implements
plan section 26's four components with these weights:
```python
CONFIDENCE_WEIGHTS: dict[str, float] = {
"fusion_margin": 0.40,
"optical_confidence": 0.20,
"sar_confidence": 0.20,
"cross_modal_agreement": 0.20,
}
```
`cross_modal_agreement` is `cosine(optical_GAP, SAR_GAP)`, rescaled from `[-1, 1]` to `[0, 1]` via
`AGREEMENT_FLOOR = -1.0`. `cross_modal_agreement` in `inference.py` returns `None` on any exception or a
shape mismatch, and the specialist treats `None` as a zero component. The composition is a weighted sum,
**floored to 0.0** when there is no prediction or no trained head β *"Both are signal GAPS, not weak
signals, so they resolve to 0.0 rather than merely discounting."* The components dict records
`has_croma`, `trained_head`, `no_prediction` and the raw availability counts, so *"the trace states every
reason the score is zero, not just the first one encountered."*
---
## 9. Pair-compatibility rules
Two workflows take a **pair** of assets, and both enforce compatibility before reading pixels.
### 9.1 Change β exactly two, temporally distinct, co-registered
`ChangeSpecialist.validate_request` (`specialists/change/specialist.py:190`) enforces **exactly two**
assets. The comment names both common mistakes: *"ONE asset is the common mistake: the caller treated
this as a single-image task. THREE is the other: they attached a time series. Both are rejected with the
same typed error."* The user-facing message is *"Change detection needs exactly two images: an earlier one
and a later one."*
`_assess_pair` then runs three checks in order:
1. **Temporal distinctness.** Identical `path` raises `TemporalPairError`. Identical `sha256` (when both
carry one) also raises: *"a repeated acquisition is not a temporal pair."*
2. **CRS compatibility** via `compare_crs` (Β§2.5). Incompatible or reprojection-requiring pairs get a
warning, not a refusal β pixel-domain change detection can still proceed, but geographic claims are
withheld.
3. **Registration quality** via `measure_registration` (Β§9.3).
If registration is **not usable**, the specialist **suppresses spatial claims** while still computing and
returning the change maps:
> The maps are still returned (the caller may want to look at them), but the REGION CLAIMS are withheld:
> a region is an assertion about *where* something changed on the ground, and we cannot make it.
`ChangeSpecialist._geospatial` drops the transform in that case, returning a `GeoMetadata` with
`is_georeferenced=False` β *"a transform invites the caller to place a region on a map, and we have just
said the regions are not placeable."* `_write_artifacts` writes the map **without** georeferencing when
suppressed, and the comment explains why copying T1's transform would be wrong: *"Copying T1's transform
onto it would produce a file that unlocks exactly the placed-on-a-map reading the suppression exists to
forbid."* The confidence is driven to the floor via a **minimum** against the registration factor, so
alignment is a gate, not one vote among three.
The change confidence weights are:
```python
CONFIDENCE_WEIGHTS: dict[str, float] = {
"registration_quality": 0.50,
"mean_change_probability": 0.30,
"component_stability": 0.20,
}
```
with `STABILITY_HEADROOM = 0.5`, `DEFAULT_MAX_SHIFT_PX = 8`, `MIN_REGION_PIXELS = 32`.
### 9.2 Optical-SAR β exactly two, one of each modality
`OpticalSarSpecialist.validate_request` (`specialists/optical_sar/specialist.py:192`) enforces **exactly
two** assets **and** that they are one optical and one SAR. The docstring explains the case worth
explaining:
> Two optical images is not a malformed request in the way that three is β it is a plausible mistake,
> from an operator who uploaded a bi-temporal pair to a single-modality workflow, or who labelled a
> four-band Cartosat scene as SAR.
>
> If such a request were accepted, the adapter would place the second optical image into the 2-channel SAR
> slot, zero-fill, and produce a fused tensor of the correct shape. CROMA would run. The answer would be
> about one modality while claiming to fuse two.
So the modality check is done on the **assets, before any pixels are read**. `_assess_pair` raises
`InvalidRequestError` with a machine-readable `reason` of `"two_optical"`, `"two_sar"` or
`"indeterminate"`, and each carries a user-facing message.
### 9.3 The equal-dimensions requirement, and the 726Β²-vs-736Β² error
The change detector requires T1 and T2 to have **identical shapes**. The guard is in the model itself,
`specialists/change/stanet.py:415`:
```python
if t1.shape != t2.shape:
raise SpecialistError(
f"T1 and T2 must have the same shape; got {tuple(t1.shape)} "
f"and {tuple(t2.shape)}",
specialist="change",
)
```
This is a **hard** requirement, and it is where the bundled demo pair fails. `docs/FINAL_DELIVERY_REPORT.md`
Β§6 records it as `DEGRADED`:
| Finding | Status | Detail |
|---|---|---|
| Change on the bundled EO pair | DEGRADED | `delta-growth-t0-1975.jpg` (726Β²) and `t1-2025.jpg` (736Β²) differ in shape; the change specialist errors (`T1 and T2 must have the same shape`). A same-shape pair, or a resize step, is needed for a clean change demo. |
There is **no resize step on the change path**. `load_image_array` explicitly does not resample
(Β§7.5) β its docstring states *"a silent resize here would put every box in the wrong place"* β and the
change specialist's `_predict` only pads for the encoder's stride-8 requirement, cropping back to the
original size:
```python
h, w = t1.shape[-2:]
if h % 8 or w % 8:
ph, pw = (-h) % 8, (-w) % 8
t1 = torch.nn.functional.pad(t1, (0, pw, 0, ph), mode="reflect")
t2 = torch.nn.functional.pad(t2, (0, pw, 0, ph), mode="reflect")
```
That padding does **not** reconcile two different input sizes: T1 and T2 are each padded by their own
`(-dim) % 8`, so a 726Β² and a 736Β² input become 728Β² and 736Β² respectively β still different β and the
model raises. The reflect-pad is a *stride* accommodation, not a *pair* accommodation, and the two must
not be conflated.
**Consequence for operators.** The change workflow requires a same-shape pair, and the failure is
`IMPLEMENTED` (it raises a typed error with a readable message) rather than silent. A resize or
co-registration step that reconciles differing dimensions is **not implemented** (Β§13). The
`FINAL_DELIVERY_REPORT` records the demo as `DEGRADED` and prescribes *"a same-shape pair, or a resize
step"* β either would fix the demo; neither is currently in the code path.
### 9.4 Registration measurement
`measure_registration` (`specialists/change/postprocess.py:134`) uses `cv2.phaseCorrelate` on grayscale
collapses of the two images, returning a `RegistrationQuality` with `shift_x`, `shift_y`, `response`,
`max_shift_px`, `is_usable` and `reason`. Defaults: `max_shift_px=8`, `min_response=0.15`.
The module docstring is candid about the method's limit:
> Deliberately a TRANSLATION model. Real mis-registration includes rotation and warp, which phase
> correlation does not recover; a small residual after a translation correction is therefore not proof of
> good alignment. What the measurement can do honestly is flag the *bad* cases, and that is how it is
> used: as a gate, not a certificate.
`RegistrationQuality.confidence_factor()` applies two independent penalties and takes the worse:
a shift penalty (`1.0` at zero shift, `0.0` at `max_shift_px`) and a response factor (`0.0` at zero
response, `1.0` at response β₯ 0.15). A pair that is offset *and* uncorrelated takes the worse of the two β
*"the conservative choice."*
A failure to *measure* is treated as a measurement of failure β both mean the pair cannot be trusted
spatially β but the distinction is recorded in the warning: *"A failure to MEASURE is not the same as a
measurement of failure."*
Two registration-gate hypotheses were tested and both are recorded as **REJECTED** or
**INSUFFICIENT**: `docs/CHANGE_REGISTRATION_GATE_COHERENCE_TEST.md` records a rejected hypothesis (AUC S2
vs S3 β€ 0.59), and `docs/CHANGE_REGISTRATION_GATE_TEXTURE_TEST.md` records a low-texture signal that was
**CONFIRMED but INSUFFICIENT**, with option 3 **ELIMINATED**. So the gate as implemented (phase
correlation) is the surviving method, and no additional coherence- or texture-based gate was adopted.
---
## 10. Large-image tiling policy
### 10.1 What is declared
`configs/base.yaml` declares the full tiling policy:
```yaml
image:
max_pixels: 25000000
tile_size: 512
tile_overlap: 128
max_tiles: 64
# plan section 9.1 tile policy: whole-image thumbnail first, then top-K tiles.
# max_tiles is the hard ceiling on tiles *examined*; top_k_tiles is how many
# are actually sent through a specialist.
top_k_tiles: 4
```
| Key | Value | Declared meaning |
|---|---|---|
| `max_pixels` | 25,000,000 | Pixel budget; exceeding it is recoverable (the caller may downscale) |
| `tile_size` | 512 | Tile edge in pixels |
| `tile_overlap` | 128 | Overlap between adjacent tiles |
| `max_tiles` | 64 | Hard ceiling on tiles **examined** |
| `top_k_tiles` | 4 | How many tiles are actually sent through a specialist |
The comment states the intended policy: *"whole-image thumbnail first, then top-K tiles."* The change
workflow declares a different tiling resolution, `change.tile_size: 256` with `change.tile_overlap: 0`
β 256 is *"the change model's OWN training resolution"* (LEVIR-CD patches are 256Γ256).
### 10.2 What is enforced
`core/config.py` validates the tiling policy at load time, but only one relationship:
```python
# --- tiling policy -------------------------------------------------
if self.get("image.top_k_tiles", 0) > self.get("image.max_tiles", 0):
errors.append("image.top_k_tiles cannot exceed image.max_tiles")
```
So the config loader guarantees that the number of tiles sent through a specialist does not exceed the
number examined. It does not implement tiling.
The pixel budget **is** enforced, but at the raster-inspection layer rather than as a tiling step:
`inspect_raster` raises `OversizedImageError` when `width * height > max_pixels`, and the error is
`recoverable` β the docstring says *"the caller may downscale."* The controller passes the configured
budget into `_inspect_asset` during VALIDATE.
### 10.3 What is not implemented β the honest status
**The tiling policy is `DECLARED (not read)`.** `image.tile_size`, `image.tile_overlap`, `image.max_tiles`
and `image.top_k_tiles` appear only in `configs/base.yaml` and in the `core/config.py` validation above.
No code on the serving path slices a raster into tiles, selects the top-K, or stitches results.
`architecture/03-request-lifecycle.md` reaches the same conclusion: the tiling policy is referenced only
in config validation and not implemented on the serving path.
The practical consequence: a raster within the `max_pixels` budget is handed to a specialist **whole**.
The VLM path pins `processor_longest_edge: 512` so the processor does not upscale and split a tile, but
that pin assumes a 512-px input β it is a *control against* the processor's default 2048, not a tiling
step. `core/config.py` enforces the relationship (`processor_longest_edge` must not exceed
`image.tile_size`), with the measured rationale recorded: the default 2048 upscales a 512 tile 4Γ and then
splits it into ~17 sub-images, *"the real figure is ~17x"* against the plan's estimated 4Γ.
`UNKNOWN β not established from the available evidence` how a raster between the tile size and the pixel
budget is handled by the VLM path in practice; the config implies the tile policy would govern it, but
the policy is not implemented, so the behaviour is whatever the specialist does with the whole image.
---
## 11. Grounding resolution β the frozen 224 decision
Grounding runs at **224 px**, and this is one of the project's cleanest pre-registered results.
`configs/base.yaml`:
```yaml
grounding:
...
# RESOLVED 2026-09-16. 448 did NOT earn its cost over all 16,159 VRSBench
# eval records: mean best IoU -0.0147, Recall@0.5 -0.0022, and every recall
# threshold lower (-0.0699 @0.10, -0.0243 @0.25), at 1.59x the latency.
# Paired over identical samples: mean diff -0.0147, 95% CI [-0.0160,-0.0134],
# t = -22.63. 448 was better on 8.5% of records, worse on 20.9%.
# Pre-registered rule and the paired test AGREE on 224.
# Evidence: docs/PHASE7_RESOLUTION_DECISION.md
image_size: 224
resolution_frozen: true
nms_iou: 0.50
max_candidates: 20
confidence_threshold: 0.40
benchmark_box_scale: 100.0
coordinate_system: normalized_0_1
encoder_projected_dim: 512
```
### 11.1 The pre-registered rule
`docs/PHASE7_RESOLUTION_DECISION.md` records the rule, **fixed before the result was seen**:
```
448 WINS if Recall@0.5 improves by >= 0.05 absolute
OR mean best IoU improves by >= 0.05 absolute
224 WINS otherwise
INCONCLUSIVE if fewer than 30 samples were scored
```
The artifact records `rule_changed_since_preregistration: false`.
### 11.2 The result
```
224 WINS
recall@0.5 gain 448/224 : -0.0022
bestIoU gain 448/224 : -0.0147
latency ratio : 1.59x
```
Neither component came close to the +0.05 margin; both were **negative**. The measured detail:
| Metric | 224 | 448 | Delta |
|---|---|---|---|
| Token grid | 7 Γ 7 = 49 | 14 Γ 14 = 196 | 4.0Γ tokens |
| Mean best IoU | **0.0972** | 0.0825 | **β0.0147** |
| Recall@0.10 | **0.3298** | 0.2599 | **β0.0699** |
| Recall@0.25 | **0.1187** | 0.0944 | **β0.0243** |
| Recall@0.50 | **0.0234** | 0.0212 | **β0.0022** |
| Latency mean | **20.0 ms** | 31.8 ms | 1.59Γ |
| Latency p90 | **20.9 ms** | 32.9 ms | 1.57Γ |
| Peak VRAM | **592.1 MB** | 599.8 MB | +7.7 MB |
All 16,159 / 16,159 VRSBench eval records were scored at both resolutions on a Tesla T4.
### 11.3 The paired test
Because both resolutions scored the **same 16,159 samples**, the paired test is the stronger statistic:
```
paired samples : 16159
mean 224 : 0.0972
mean 448 : 0.0825
mean paired diff : -0.0147 (95% CI -0.0160 .. -0.0134)
t statistic : -22.63
CI excludes zero : True
448 better on : 1371/16159 ( 8.5%)
448 worse on : 3372/16159 (20.9%)
identical : 11416/16159 (70.6%)
```
The paired test and the pre-registered rule **agree**, so there is *"no rule-versus-evidence disagreement
to escalate."* The win/loss split is informative on its own: 448 wins on only 8.5% and loses on 20.9% β
*"the finer grid is not merely neutral, it is actively harmful on a fifth of the corpus."*
The recall ladder narrows as the threshold rises (β0.0699 at 0.10, β0.0243 at 0.25, β0.0022 at 0.50),
which the document reads as *"the signature of a method that cannot reach high IoU either way."*
### 11.4 Why 448 did not help
The document's honest reading: the zero-shot method selects a patch by text similarity and returns that
patch's box. At 224 a box is 1/7 of the image; at 448 it is 1/14. A finer grid is only better if the
target is small **and** the similarity peak lands on the correct fine cell. Two things work against that:
the peak is not sharper at 448 (the similarity field on frozen features is smooth, so the argmax moves
around), and Recall@0.10 β the loosest threshold β degrades *most*, which means *"the fine grid is adding
positional noise rather than positional precision."*
The document is careful to attribute this to the zero-shot baseline, not to RemoteCLIP: *"A **learned**
head trained to regress boxes from these features may respond differently to resolution."*
### 11.5 What the decision does not establish
- **Whether the zero-shot baseline is good.** It is not: mean best IoU 0.0972 and Recall@0.5 0.0234 are
*"weak"*. It is an ablation floor for the Phase 8 head.
- **Whether a trained head has the same resolution sensitivity.** Re-opening the question is legitimate
**if** the head's validation curve suggests it, and would be *"a new pre-registered experiment, not a
silent retune."*
- **Anything about hidden ISRO/SAC imagery.** *"VRSBench is overhead optical. The hidden set is
Cartosat-2S + RISAT, a different distribution entirely."*
### 11.6 Downstream consequences of the frozen resolution
The frozen resolution propagates into concrete serving parameters:
- **The decode grid.** `build_grounding_specialist` sets `grid = resolution // 32`, so at 224 the grid is
**7Γ7 = 49 cells**. The head emits `(1, 49, 5)` β `[tx, ty, tw, th, objectness]` per cell.
- **The projected dimension.** `grounding.encoder_projected_dim: 512` is the measured `visual.proj`
output of RemoteCLIP ViT-B/32 (the transformer width is 768; `visual.proj` maps to 512). The
per-cell feature is `concat([patch, text, patch*text, global_pool]) = 4 Γ 512 = 2048`, which
`grounding_head.feature_dim: 2048` must equal β enforced at load time by `core/config.py`, which
rejects any other value, because *"a mismatch here is a SILENT shape error"* that torch raises only at
the similarity step, *"after the patch features are already cached."*
- **NMS.** `grounding.nms_iou: 0.50` is passed to `nms(boxes, scores, iou_threshold)`
(`specialists/grounding/head.py:294`).
- **The serving candidate budget.** `GroundingSpecialist` defaults `max_candidates=6`, **not** the
config's `20`. The class docstring records why: *"The Phase 8 decision measured that a 20-box budget
inflates mean best IoU relative to the 5.99-box baseline; 6 is the decode-matched setting."* The builder
reads `grounding.serving_max_candidates` with a default of 6.
- **The degenerate-box guard.** `DEGENERATE_AREA_FRACTION = 0.9` drops any zero-shot candidate covering
more than 90% of the frame, *"the shape a failed localisation takes once it has been smoothed into a
rectangle."* This guard applies to the **zero-shot fallback path only** β *"the trained-head decode is
a regressed box with its own learned prior; filtering it would change the frozen Phase 8 benchmark
numbers."*
- **The decode function.** `decode_cell_relative` (`specialists/grounding/head.py:204`) turns
`(B, N, 4)` cell-relative parameters into `(B, N, 4)` normalised xyxy, clamped to `[0, 1]` and
order-corrected by `enforce_order` (which swaps inverted corners rather than letting IoU go silently to
zero). The decode is *"differentiable: the output feeds the box and GIoU losses directly, so the loss is
computed on exactly the coordinates the metric measures."*
---
## 12. Worked examples
### 12.1 A 4-band Cartosat-2S optical scene
1. **Inspect.** `inspect_raster` reads `band_count=4`, no CRS or with CRS. `infer_modality(4)` returns
`OPTICAL` (4 is in `_OPTICAL_BAND_COUNTS`).
2. **Descriptor.** With no `SensorDescriptor`, `_descriptor_for` maps the 4 bands positionally onto
`OPTICAL_CANONICAL[:4]` = `B01, B02, B03, B04` and appends a warning naming the assumption. With a
descriptor built by `build_optical_adapter("cartosat_2s")`, the bands are `["B1","B2","B3","B4"]`
canonicalised to `B01..B04`, and `availability = [True, True, True, True, False Γ 8]`.
3. **Canonicalise.** `adapt_optical` returns a `(12, H, W)` tensor with channels 0β3 carrying the real
bands and channels 4β11 exactly zero, plus a 12-element mask with four `True`.
4. **Resize.** `resize_to_canonical` bilinearly resizes to `120Γ120`. The zero channels stay zero.
5. **Encode.** `encoder.encode(optical_tensor[None], sar_tensor[None])` β **no mask** (C-1).
6. **Assemble.** `assemble_fusion_input` concatenates the three 768-d GAPs with the 12-element optical
mask and the 2-element SAR mask β `(1, 2318)`.
7. **Head.** With a trained head, the classifier emits 19 logits β softmax β a label and a margin. With
no head, `probabilities` stays `None`, `degraded=True`, and no label is invented.
8. **Evidence.** An `AVAILABILITY_MASK` item per modality records the sensor, the 4 real bands, the band
map, the mask (`[true Γ 4, false Γ 8]`), and `missing_channels: [4,5,6,7,8,9,10,11]`. A
`JOINT_FEATURE_REGION` item records `fusion_input_dim: 2318`, `mask_consumed_by: "fusion_head"`,
`croma_received_mask: false`.
9. **Confidence.** `optical_confidence = 4/12 = 0.3333`; the SAR side contributes its own fraction; the
fusion margin (if a head exists) is the largest weight. If no head, the raw confidence is **0.0** with
`no_prediction: 1.0` recorded.
### 12.2 A RISAT SAR scene with no band labels
1. **Descriptor.** `positional_fallback_descriptor(stem, 2)` selects `("VV", "VH")` from
`_FALLBACK_SAR_ORDER` and records the fallback in the `sensor` field. The specialist appends: *"The
polarisation IDENTITY of each channel is NOT known from a band count alone."*
2. **Canonicalise.** `adapt_sar` returns `(2, H, W)` with both slots filled, mask `[True, True]`.
3. **Trace.** `available_bands = ["VV", "VH"]` is the **assumption**, not a measurement. A caller with
real polarisation labels passes `available_bands=["HH","HV"]`, in which case the descriptor's own order
becomes the slot order and the mask reflects it β no relabelling, no discarded data.
### 12.3 A change pair with differing dimensions
1. **Validate.** Exactly two assets, distinct paths/SHA-256, CRS compared.
2. **Register.** `measure_registration` collapses both to grayscale float32 and runs `phaseCorrelate`.
3. **Shape check.** If the two images differ in size, `STANetStyleChangeDetector.forward` raises
`SpecialistError("T1 and T2 must have the same shape; got ...")` β the 726Β²/736Β² case (Β§9.3). No resize
step exists to reconcile them.
4. **If shapes match but registration is poor:** the maps are still computed, the regions are withheld,
the transform is dropped from the `geospatial` block, and the confidence is floored at 0.0 with
`suppressed_by_registration: 1.0` recorded.
---
## 13. What is NOT implemented β exhaustive
Each item below was checked against the code. Where a capability is absent, it is named rather than
implied; where the absence cannot be confirmed from the files read, it is marked `UNKNOWN`.
| Capability | Status | Evidence |
|---|---|---|
| **Reprojection pipeline** | **NOT IMPLEMENTED** | `geospatial/crs.py::compare_crs` *reports* `requires_reprojection`; no code performs a reprojection. `preprocessing/imagery.py` explicitly does not reproject. |
| **Cloud masking** | **NOT IMPLEMENTED** | No cloud, cirrus, QA or `SCL` handling anywhere in `preprocessing/` or the specialists. `UNKNOWN β not established from the available evidence` whether any dataset adapter applies a cloud mask. |
| **Atmospheric correction** | **NOT IMPLEMENTED** | No atmospheric, surface-reflectance or `L2A`-processing code. Rasters are consumed as delivered. |
| **Mosaicking** | **NOT IMPLEMENTED** | No mosaicking, seamline or composite code exists. |
| **Tiling / top-K tile selection** | **DECLARED (not read)** | `image.tile_size`, `image.tile_overlap`, `image.max_tiles`, `image.top_k_tiles` are read only by `core/config.py`'s `top_k_tiles <= max_tiles` check. No serving-path tiling. |
| **Optical percentile 2/98 conditioning** | **DECLARED (not read)** | `optical.normalization`, `optical.lower_percentile`, `optical.upper_percentile` are read by no code (`radiometry.py` docstring; PHASE14 Β§1). |
| **SAR dB clip β30/+5 conditioning** | **DECLARED (not read)** | `sar.representation`, `sar.clip_min_db`, `sar.clip_max_db` are read by no code. |
| **Resampling / cropping on the display path** | **DELIBERATELY ABSENT** | `preprocessing/imagery.py` docstring: *"It does not resample, crop, or reproject."* |
| **Resize to reconcile differing pair dimensions** | **NOT IMPLEMENTED** | The change path requires equal shapes and has no reconciling resize (Β§9.3). |
| **Nodata filling** | **NOT IMPLEMENTED** | `quality.py` refuses any non-finite input outright; its comment records: *"Filling nodata from the raster profile belongs in the tiling work."* |
| **Area computation for geographic CRS** | **NOT IMPLEMENTED** | `pixel_area_m2` returns `None` for a non-metric CRS rather than approximating. |
| **GeoJSON / WKT / KML output** | **NOT IMPLEMENTED** | Coordinates are emitted as lists; no geometry serialisation exists. |
| **Multi-band raster β RGB band selection policy** | **PARTIAL** | `load_image_array` takes the first three bands as RGB (or repeats a single band); there is no configurable band combination. |
| **Sensor adapter for sensors beyond the known list** | **REFUSED BY DESIGN** | An unknown optical sensor with no explicit band list raises `UnsupportedBandsError`; `_known_sar` returns `[]` for an unknown SAR sensor. |
**A note on the "not implemented" list.** None of these omissions is hidden. The two most consequential β
the tiling policy and the radiometric conditioning β are declared in the frozen config and read by no
code, and both the config comments and the DEV-2 ruling name that state explicitly rather than letting the
key read as a satisfied requirement.
---
## 14. What is NOT RUN / OPEN / BLOCKED for this topic
### NOT RUN
- **The optical-SAR forward pass at serving time with a trained head.** `CROMA_base.pt` is not present
locally and `use_croma.py` must be vendored; the specialist runs sensor-only and produces no label.
- **A system-level geospatial accuracy benchmark.** No end-to-end benchmark exists; no system-level
accuracy is claimed anywhere.
- **Reprojection of a genuinely mismatched-CRS pair.** The code path that would consume a reprojection is
absent, so the case has never been exercised.
- **Change on a bundled pair with equal dimensions.** The only bundled demo pair is 726Β² vs 736Β² and
errors; a same-shape pair has not been run as a demo.
### OPEN
- **The `optical.*` / `sar.*` conditioning stage.** `OPEN` β either it gets an implementation or the keys
are retired (`docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§6 item 3).
- **The arm-constructibility question.** The Β§4-registered arm B (*"percentile/dB only, no encoder-input
stretch"*) is **not constructible** while stage 1 is unimplemented; the arm now named B corresponds to
the registered arm C. The ruling records this as *"a human decision and deliberately not resolved."*
- **The optical-SAR ruling.** Accuracy **0.931** with macro-F1 **0.434161**; ruling **OPEN**. Never quote
the accuracy without the macro-F1.
- **The grounding resolution question after a trained head.** Re-opening 224 is legitimate only as a new
pre-registered experiment.
- **The change-VQA metric ruling.** `OPEN` and owner-gated.
- **Whether any call site builds a `Box` outside the two conversion functions** (Β§5.4) β `UNKNOWN`.
### BLOCKED
- **A clean change demo.** Blocked on a same-shape pair or a resize step.
- **The CROMA normalisation experiment (option c).** Blocked on arm B being non-constructible and on the
Arm-B feature cache, which does not exist (`artifacts/optical_sar/fusion_features/` holds only the
Arm-A cache).
- **The gated experiment's sample size.** Cannot be specified until a variance-estimation pass is run;
the ruling deliberately declines to invent a number.
### REJECTED
- **448 px grounding resolution.** `REJECTED` by the pre-registered rule and the paired test.
- **Option (b) β match `base.yaml` alone, accept the shift.** `REJECTED` as a standalone path
(`docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§2.1).
- **Reusing `preprocessing/imagery.py`'s stretch for CROMA.** `REJECTED` as `forbidden` (Β§3.3 of the
ruling).
- **Coherence- and texture-based registration gates.** One `REJECTED`, one `CONFIRMED but INSUFFICIENT`.
---
## 15. Where the evidence lives
**Source modules (the code this chapter describes):**
| Path | What it defines |
|---|---|
| `preprocessing/raster.py` | The validation chain, `inspect_raster`, `read_bands`, `write_raster`, `infer_modality`, `file_sha256`, `pixel_area_m2` |
| `preprocessing/imagery.py` | `load_image_array`, `load_pil_image`, the per-scene 2/98 display stretch |
| `preprocessing/quality.py` | The deterministic quality gate, `MIN_AUTOCORRELATION=0.10`, `NOISE_ENTROPY_BITS=7.8` |
| `geospatial/crs.py` | `parse_crs`, `describe_crs`, `compare_crs`, `is_metric`, `CRSCompatibility` |
| `geospatial/transform.py` | `PixelWindow`, `GeoBounds`, `to_normalized`/`to_pixel`/`to_geo`, `intersect`, `iou`, `benchmark_boxes_to_normalized`, `affine_from_metadata` |
| `core/schemas.py` | `CoordinateSystem`, `GeoMetadata`, `SensorDescriptor`, `AssetMetadata`, `Box`, `Region`, `ChangeRegion`, `Evidence._spatial_needs_crs` |
| `core/config.py` | Config loading, validation, `Config.hash`, `device_preference` |
| `specialists/optical_sar/sensor_adapter.py` | `OPTICAL_CANONICAL`, `SAR_CANONICAL`, `build_optical_adapter`, `describe_sar`, `_apply_canonical`, `positional_fallback_descriptor` |
| `specialists/optical_sar/radiometry.py` | `TRANSFORM_NAME`, `normalise_for_croma`, `resolve_use_8_bit`, the zero-channel rule |
| `specialists/optical_sar/fusion_head.py` | `expected_fusion_dim`, `assemble_fusion_input`, `channel_dropout`, `build_fusion_head` |
| `specialists/optical_sar/inference.py` | `run_pipeline`, `resize_to_canonical`, `cross_modal_agreement`, `build_head_from_config` |
| `specialists/optical_sar/specialist.py` | `_assess_pair`, `_descriptor_for`, `_confidence_components`, `_geospatial` |
| `specialists/change/postprocess.py` | `RegistrationQuality`, `measure_registration`, `to_grayscale_float`, `regions_to_schema` |
| `specialists/change/specialist.py` | `_assess_pair`, `_predict`, `_geospatial`, the suppression path |
| `specialists/change/stanet.py:415` | The equal-shape guard |
| `specialists/grounding/specialist.py` | `_to_geo_boxes`, the `NORMALIZED_0_1` box construction, `DEGENERATE_AREA_FRACTION` |
| `specialists/grounding/head.py` | `decode_cell_relative`, `enforce_order`, `nms` |
| `specialists/grounding/inference.py` | `decode_candidates_from_features`, `ground_phrase` |
| `evaluation/metrics/grounding.py` | `benchmark_to_normalized(box, scale=100.0)` |
**Configuration:**
- `configs/base.yaml` β `image.*` (tiling), `optical.*`, `sar.*`, `croma.*`, `fusion.*`, `grounding.*`,
`grounding_head.*`, `change.*`.
**Documents:**
- `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` β the DEV-2 ruling, the two ordered stages, the Gate F arm
amendment.
- `docs/CROMA_NORMALISATION_UPSTREAM_EVIDENCE.md` β what is and is not established about CROMA's
pretraining distribution.
- `docs/PHASE7_RESOLUTION_DECISION.md` β the 224/448 decision in full.
- `docs/CHANGE_REGISTRATION_GATE_COHERENCE_TEST.md`, `docs/CHANGE_REGISTRATION_GATE_TEXTURE_TEST.md` β the
rejected and insufficient registration gates.
- `docs/API_CONTRACT.md` Β§2.5 (asset allowlist), Β§3.2 (coordinate systems).
- `docs/FINAL_DELIVERY_REPORT.md` Β§6 β the 726Β²/736Β² change demo recorded as DEGRADED.
- `docs/PHASE10_CDVQA_DATA_STATUS.md` β the CDVQA/SECOND pair layout and temporal semantics.
- `docs/PHASE11_CROMA_STATUS.md` β the reproduced CROMA forward pass.
**Sibling release chapters:**
- [`DATA_PIPELINE.md`](DATA_PIPELINE.md) β the end-to-end data path, caches and artefacts.
- [`architecture/03-request-lifecycle.md`](architecture/03-request-lifecycle.md) β the nine controller
states, including VALIDATE and PREPROCESS.
- [`architecture/05-specialists.md`](architecture/05-specialists.md) β the six specialists.
- [`architecture/07-configuration-freeze.md`](architecture/07-configuration-freeze.md) β the frozen hash
and every config key.
- [`DATASETS.md`](DATASETS.md) β the corpora and their measured shapes.
|