Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
release: add docs/GEOSPATIAL.md
Browse files- docs/GEOSPATIAL.md +1747 -0
docs/GEOSPATIAL.md
ADDED
|
@@ -0,0 +1,1747 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Geospatial β deep reference
|
| 2 |
+
|
| 3 |
+
**Status tags used on every substantive claim:** `IMPLEMENTED` Β· `VERIFIED` Β· `MEASURED` Β· `ATTEMPTED`
|
| 4 |
+
Β· `NOT RUN` Β· `BLOCKED` Β· `DEFERRED` Β· `REJECTED` Β· `OPEN` Β· `RESOLVED` Β· `CLOSED`.
|
| 5 |
+
|
| 6 |
+
Every statement in this document was produced by reading a file in `C:/Users/anish/satquery-ai/`. Where
|
| 7 |
+
a fact is not established from the available evidence, this document writes
|
| 8 |
+
`UNKNOWN β not established from the available evidence`. Where a transform is *declared* in configuration
|
| 9 |
+
but read by no code, that is stated explicitly rather than presented as a working feature. Where a
|
| 10 |
+
capability is absent, it is listed in [Β§13](#13-what-is-not-implemented-exhaustive) rather than implied.
|
| 11 |
+
|
| 12 |
+
This chapter is the reference for the **spatial contract**: how a raster becomes an analysed image, what
|
| 13 |
+
georeferencing survives, what coordinate system a box is in, how a sensor's bands become CROMA's canonical
|
| 14 |
+
channels, and which of those steps are actually implemented on the serving path. It is written to be
|
| 15 |
+
sufficient to reconstruct the subsystem from the document alone.
|
| 16 |
+
|
| 17 |
+
**Sibling documents:** [`DATA_PIPELINE.md`](DATA_PIPELINE.md) (the end-to-end data path, caches and
|
| 18 |
+
artefacts), [`architecture/03-request-lifecycle.md`](architecture/03-request-lifecycle.md) (where the
|
| 19 |
+
PREPROCESS and VALIDATE states sit), [`architecture/05-specialists.md`](architecture/05-specialists.md)
|
| 20 |
+
(the six specialists in full), [`architecture/07-configuration-freeze.md`](architecture/07-configuration-freeze.md)
|
| 21 |
+
(the frozen config identity and every key), [`DATASETS.md`](DATASETS.md) (the corpora and their measured
|
| 22 |
+
shapes).
|
| 23 |
+
|
| 24 |
+
---
|
| 25 |
+
|
| 26 |
+
## Table of contents
|
| 27 |
+
|
| 28 |
+
0. [How to read this document](#0-how-to-read-this-document)
|
| 29 |
+
1. [The raster input contract](#1-the-raster-input-contract)
|
| 30 |
+
2. [CRS handling](#2-crs-handling-geospatialcrs-py)
|
| 31 |
+
3. [Affine transforms and coordinate frames](#3-affine-transforms-and-coordinate-frames)
|
| 32 |
+
4. [The coordinate-system contract](#4-the-coordinate-system-contract)
|
| 33 |
+
5. [Benchmark box scale β 0β100 versus 0β1](#5-benchmark-box-scale--0100-versus-01)
|
| 34 |
+
6. [Band-count modality inference](#6-band-count-modality-inference)
|
| 35 |
+
7. [Radiometric normalisation and representation](#7-radiometric-normalisation-and-representation)
|
| 36 |
+
8. [The CROMA sensor adapter contract](#8-the-croma-sensor-adapter-contract)
|
| 37 |
+
9. [Pair-compatibility rules](#9-pair-compatibility-rules)
|
| 38 |
+
10. [Large-image tiling policy](#10-large-image-tiling-policy)
|
| 39 |
+
11. [Grounding resolution β the frozen 224 decision](#11-grounding-resolution--the-frozen-224-decision)
|
| 40 |
+
12. [Worked examples](#12-worked-examples)
|
| 41 |
+
13. [What is NOT implemented β exhaustive](#13-what-is-not-implemented-exhaustive)
|
| 42 |
+
14. [What is NOT RUN / OPEN / BLOCKED for this topic](#14-what-is-not-run--open--blocked-for-this-topic)
|
| 43 |
+
15. [Where the evidence lives](#15-where-the-evidence-lives)
|
| 44 |
+
|
| 45 |
+
---
|
| 46 |
+
|
| 47 |
+
## 0. How to read this document
|
| 48 |
+
|
| 49 |
+
Three distinctions are load-bearing throughout, and conflating them is the single most common way to
|
| 50 |
+
misread this system:
|
| 51 |
+
|
| 52 |
+
1. **Declared versus executed.** `configs/base.yaml` is the authoritative registry, and `core/config.py`
|
| 53 |
+
validates and hashes it at load time. But *declaring* a policy and *executing* it are different
|
| 54 |
+
claims. The tiling policy, the optical percentile stretch and the SAR dB clip are all declared in
|
| 55 |
+
`configs/base.yaml` and are all read by no code on the serving path. This document says so at each
|
| 56 |
+
point rather than presenting the key as a feature.
|
| 57 |
+
2. **Adapter versus encoder.** The CROMA sensor adapter (`specialists/optical_sar/sensor_adapter.py`)
|
| 58 |
+
performs band mapping, zero-fill and availability masking. CROMA the encoder does **not** receive the
|
| 59 |
+
availability mask (finding C-1). The mask is consumed by the fusion head. These are separate contracts
|
| 60 |
+
and Β§8.9 treats the distinction in full.
|
| 61 |
+
3. **Presentation versus radiometry.** There are two different percentile stretches in this codebase:
|
| 62 |
+
a per-*scene* 2/98 display stretch in `preprocessing/imagery.py` (and a grayscale cousin in
|
| 63 |
+
`specialists/change/postprocess.py`), and a per-*channel* `mean Β± 2Β·std` encoder-input stretch in
|
| 64 |
+
`specialists/optical_sar/radiometry.py`. They serve different purposes and are **not**
|
| 65 |
+
interchangeable. Β§7.5 is explicit about why.
|
| 66 |
+
|
| 67 |
+
The status vocabulary is used on every substantive claim. `MEASURED` means a number was produced by
|
| 68 |
+
executing code; `VERIFIED` means the claim was checked against a second independent source; `IMPLEMENTED`
|
| 69 |
+
means the code exists on the serving path; `DECLARED (not read)` means the value lives in
|
| 70 |
+
`configs/base.yaml` and no code reads it.
|
| 71 |
+
|
| 72 |
+
---
|
| 73 |
+
|
| 74 |
+
## 1. The raster input contract
|
| 75 |
+
|
| 76 |
+
### 1.1 What arrives, and what is accepted
|
| 77 |
+
|
| 78 |
+
An asset reaches the analysis path through `POST /v1/assets`, which mints an opaque `asset_id`. The
|
| 79 |
+
gateway enforces a **closed content-type allowlist of exactly five types**
|
| 80 |
+
(`docs/API_CONTRACT.md` Β§2.5):
|
| 81 |
+
|
| 82 |
+
| Content type | Notes |
|
| 83 |
+
|---|---|
|
| 84 |
+
| `image/tiff` | **The type the geospatial specialists need.** A client that uploads only PNG/JPEG can serve `vqa`, `caption` and `grounding`, but not `change` or `optical_sar` |
|
| 85 |
+
| `image/geotiff` | Accepted; media-type parameters are ignored, so `image/tiff; charset=binary` is accepted |
|
| 86 |
+
| `image/png` | Accepted |
|
| 87 |
+
| `image/jpeg` | Accepted |
|
| 88 |
+
| `application/octet-stream` | Accepted, but the raster reader still has to be able to open the bytes |
|
| 89 |
+
|
| 90 |
+
Two behaviours are specified rather than defaulted: a request with **no** declared content type is
|
| 91 |
+
**refused** rather than defaulted (defaulting is how a PDF reaches a raster reader), and a disallowed
|
| 92 |
+
type gets `415`. The stored filename is derived from the **content type**, never from the client's
|
| 93 |
+
filename, so a client-supplied `../../` cannot influence where bytes land (`docs/API_CONTRACT.md` Β§2.5).
|
| 94 |
+
The size cap is enforced at two layers β the gateway and the Space β and both now refuse an unparsable
|
| 95 |
+
or non-positive value rather than silently falling back to a default
|
| 96 |
+
(`docs/API_CONTRACT.md` Β§2.5, "Size cap").
|
| 97 |
+
|
| 98 |
+
`artifact_ref` on every `Evidence` item is **always `null` in v1** (`core/schemas.py:225`): v1 exposes no
|
| 99 |
+
artifact-serving endpoint, so a reference a client cannot retrieve is not fabricated. The schema field's
|
| 100 |
+
own description states this β *"Reference to an externally retrievable artifact. NEVER a filesystem path
|
| 101 |
+
(F-16, owner ruling 2026-09-23)"* β and the change and optical-SAR specialists both emit `null` on every
|
| 102 |
+
path (Β§4.3).
|
| 103 |
+
|
| 104 |
+
### 1.2 The validation chain
|
| 105 |
+
|
| 106 |
+
`preprocessing/raster.py` is the first thing every uploaded asset meets. Its module docstring names the
|
| 107 |
+
chain it implements, verbatim:
|
| 108 |
+
|
| 109 |
+
```
|
| 110 |
+
file -> dimensions -> bands -> dtype -> CRS -> transform -> bounds -> nodata
|
| 111 |
+
-> modality -> temporal metadata
|
| 112 |
+
```
|
| 113 |
+
|
| 114 |
+
The module's three design rules are stated at the top of the file and are worth quoting because they
|
| 115 |
+
govern every failure mode below:
|
| 116 |
+
|
| 117 |
+
> * Never raise a bare exception. Every failure is a typed SatQueryError.
|
| 118 |
+
> * Never silently drop geospatial metadata. If the source had a CRS, the returned AssetMetadata says so.
|
| 119 |
+
> * A missing CRS degrades to non-geospatial mode; it does not abort.
|
| 120 |
+
|
| 121 |
+
### 1.3 `inspect_raster` β header-only validation
|
| 122 |
+
|
| 123 |
+
`inspect_raster` (`preprocessing/raster.py:96`) validates and describes a raster **without loading its
|
| 124 |
+
pixel data**. Its exact signature:
|
| 125 |
+
|
| 126 |
+
```python
|
| 127 |
+
def inspect_raster(
|
| 128 |
+
path: str | Path,
|
| 129 |
+
*,
|
| 130 |
+
max_pixels: int | None = None,
|
| 131 |
+
explicit_modality: str | None = None,
|
| 132 |
+
compute_hash: bool = False,
|
| 133 |
+
sensor: SensorDescriptor | None = None,
|
| 134 |
+
) -> AssetMetadata:
|
| 135 |
+
```
|
| 136 |
+
|
| 137 |
+
The behaviour, field by field:
|
| 138 |
+
|
| 139 |
+
- The path is checked with `Path(path).exists()` and `is_file()`; a missing path raises `RasterReadError`.
|
| 140 |
+
- `rasterio` is imported lazily; an `ImportError` raises `RasterReadError` rather than a bare error.
|
| 141 |
+
- Inside `rasterio.open(p)`, the reader extracts `width`, `height`, `count`, the sorted set of `dtypes`,
|
| 142 |
+
the CRS as a string (`ds.crs.to_string() if ds.crs else None`), the affine transform truncated to its
|
| 143 |
+
first six coefficients (`list(ds.transform)[:6] if ds.transform else None`), the bounds
|
| 144 |
+
(`list(ds.bounds)`), `nodata`, `resolution` (`list(ds.res)`), the `driver` string, and a tiled flag
|
| 145 |
+
read through `_safe_is_tiled`.
|
| 146 |
+
- `_safe_is_tiled` (`preprocessing/raster.py:42`) exists because `DatasetReader.is_tiled` is scheduled for
|
| 147 |
+
removal; it suppresses the pending deprecation and **degrades to `False`** rather than aborting, since
|
| 148 |
+
the flag is informational metadata only.
|
| 149 |
+
- Degenerate dimensions (`width <= 0 or height <= 0`) raise `RasterReadError`.
|
| 150 |
+
- If `max_pixels` is given and `width * height > max_pixels`, it raises `OversizedImageError` β and the
|
| 151 |
+
docstring records that this is **recoverable**: *"the caller may downscale"*.
|
| 152 |
+
- A band count `<= 0` raises `UnsupportedBandsError`.
|
| 153 |
+
- Anything else rasterio raises is wrapped: `RasterReadError(f"could not open raster '{p}': {exc}")`.
|
| 154 |
+
|
| 155 |
+
The return is an `AssetMetadata` (see Β§1.4) with `modality` from `infer_modality` (Β§6) and, when
|
| 156 |
+
`compute_hash=True`, the streaming SHA-256 of the file (`file_sha256`, `preprocessing/raster.py:82`,
|
| 157 |
+
chunked at 1 MiB).
|
| 158 |
+
|
| 159 |
+
`inspect_raster` is called from `core/controller.py::_inspect_asset` during the VALIDATE state, with the
|
| 160 |
+
configured pixel budget passed through. That means the **header-only** nature of the inspection is a real
|
| 161 |
+
property of the serving path: a request never decodes a raster just to learn its shape.
|
| 162 |
+
|
| 163 |
+
### 1.4 `AssetMetadata` and `GeoMetadata`
|
| 164 |
+
|
| 165 |
+
`GeoMetadata` (`core/schemas.py:120`) is the geospatial detail carrier. It is `extra="allow"`, which is
|
| 166 |
+
deliberate β the optical-SAR specialist adds cross-sensor keys to it (Β§4.3):
|
| 167 |
+
|
| 168 |
+
| Field | Type | Meaning |
|
| 169 |
+
|---|---|---|
|
| 170 |
+
| `crs` | `str \| None` | CRS as a string (e.g. `EPSG:4326`), or `None` |
|
| 171 |
+
| `transform` | `list[float] \| None` | Affine coefficients, row-major, six values |
|
| 172 |
+
| `bounds` | `list[float] \| None` | `[minx, miny, maxx, maxy]` |
|
| 173 |
+
| `width`, `height` | `int \| None` | Raster dimensions in pixels |
|
| 174 |
+
| `band_count` | `int \| None` | Number of bands |
|
| 175 |
+
| `dtype` | `str \| None` | Comma-joined distinct dtypes when the bands disagree |
|
| 176 |
+
| `nodata` | `float \| None` | The nodata sentinel, or `None` |
|
| 177 |
+
| `resolution` | `list[float] \| None` | `(xres, yres)` |
|
| 178 |
+
| `has_crs` | `bool` | `crs_value is not None` |
|
| 179 |
+
| `is_georeferenced` | `bool` | `bool(crs_value and transform)` |
|
| 180 |
+
|
| 181 |
+
`AssetMetadata` (`core/schemas.py:153`) is `extra="forbid"` and carries:
|
| 182 |
+
|
| 183 |
+
| Field | Type | Meaning |
|
| 184 |
+
|---|---|---|
|
| 185 |
+
| `asset_id` | `str` | Minted `asset_<12 hex>`, opaque |
|
| 186 |
+
| `path` | `str` | Resolved local path |
|
| 187 |
+
| `modality` | `Modality` | From `infer_modality`, or declared |
|
| 188 |
+
| `sha256` | `str \| None` | Present only when `compute_hash=True` |
|
| 189 |
+
| `geo` | `GeoMetadata` | The block above |
|
| 190 |
+
| `sensor` | `SensorDescriptor \| None` | Present only when a sensor descriptor was supplied |
|
| 191 |
+
| `acquisition_date` | `str \| None` | Temporal metadata |
|
| 192 |
+
| `scene_id`, `dataset_id`, `split` | `str \| None` | Provenance for leakage control |
|
| 193 |
+
|
| 194 |
+
Note the distinction between `has_crs` and `is_georeferenced`: a raster can carry a CRS but no affine
|
| 195 |
+
transform (or vice versa), and both are recorded. Only when **both** are present is
|
| 196 |
+
`is_georeferenced=True`. `geospatial/transform.py::to_geo` requires the transform, so a raster with
|
| 197 |
+
`has_crs=True` but `transform=None` cannot be converted to geographic coordinates β it will raise
|
| 198 |
+
`CoordinateError("cannot convert to geo: raster has no affine transform")`.
|
| 199 |
+
|
| 200 |
+
### 1.5 Band reading, writing, and pixel area
|
| 201 |
+
|
| 202 |
+
`read_bands` (`preprocessing/raster.py:188`) returns `(bands, H, W)` plus a profile that **retains**
|
| 203 |
+
`crs`, `transform` and `bounds` as plain values, so downstream code never loses georeferencing. The
|
| 204 |
+
profile is a copy of `ds.profile` with those three keys overwritten from the live dataset.
|
| 205 |
+
|
| 206 |
+
`write_raster` (`preprocessing/raster.py:221`) writes an array using a profile captured by `read_bands`,
|
| 207 |
+
and re-applies `crs` and `transform` explicitly in the `rasterio.open(..., "w", ...)` call. It is used by
|
| 208 |
+
evidence generation so change maps stay georeferenced β **except** when the pair is suppressed for poor
|
| 209 |
+
registration, in which case the transform is deliberately dropped (Β§9.1).
|
| 210 |
+
|
| 211 |
+
`pixel_area_m2` (`preprocessing/raster.py:259`) returns the area of one pixel in square metres **when the
|
| 212 |
+
CRS is metric**, and `None` otherwise:
|
| 213 |
+
|
| 214 |
+
```python
|
| 215 |
+
def pixel_area_m2(meta: AssetMetadata) -> float | None:
|
| 216 |
+
from geospatial.crs import is_metric
|
| 217 |
+
if not is_metric(meta.geo.crs):
|
| 218 |
+
return None
|
| 219 |
+
if not meta.geo.resolution or len(meta.geo.resolution) < 2:
|
| 220 |
+
return None
|
| 221 |
+
xres, yres = meta.geo.resolution[0], meta.geo.resolution[1]
|
| 222 |
+
return abs(xres * yres)
|
| 223 |
+
```
|
| 224 |
+
|
| 225 |
+
The docstring is explicit: *"Returns None for geographic CRS or unknown resolution β callers must not
|
| 226 |
+
guess."* A degree-based CRS has no meaningful pixel area in metres, so the function refuses rather than
|
| 227 |
+
returning a plausible-looking number.
|
| 228 |
+
|
| 229 |
+
---
|
| 230 |
+
|
| 231 |
+
## 2. CRS handling (`geospatial/crs.py`)
|
| 232 |
+
|
| 233 |
+
CRS comparison is kept **separate** from pixel arithmetic so that *"can these two rasters be compared at
|
| 234 |
+
all?"* is answerable without touching pixels (`geospatial/crs.py` docstring). The module depends on
|
| 235 |
+
`pyproj`.
|
| 236 |
+
|
| 237 |
+
### 2.1 `parse_crs`
|
| 238 |
+
|
| 239 |
+
```python
|
| 240 |
+
def parse_crs(value: str | None) -> CRS | None:
|
| 241 |
+
if value is None or value == "":
|
| 242 |
+
return None
|
| 243 |
+
try:
|
| 244 |
+
return CRS.from_user_input(value)
|
| 245 |
+
except PyProjCRSError as exc:
|
| 246 |
+
raise MissingCRSError(f"unparseable CRS '{value}': {exc}") from exc
|
| 247 |
+
```
|
| 248 |
+
|
| 249 |
+
Absent is `None`; malformed raises. That distinction matters: a missing CRS is a legitimate state
|
| 250 |
+
(degrades to non-geospatial), while a *malformed* one is a defect that must surface.
|
| 251 |
+
|
| 252 |
+
### 2.2 `describe_crs`
|
| 253 |
+
|
| 254 |
+
`describe_crs` returns a dict for the execution trace: `present`, `srs`, `name`, `is_geographic`,
|
| 255 |
+
`is_projected`, and `axis_units` (the first axis's `unit_name`). For an absent CRS it returns
|
| 256 |
+
`{"present": False}`.
|
| 257 |
+
|
| 258 |
+
### 2.3 `compare_crs` and `CRSCompatibility`
|
| 259 |
+
|
| 260 |
+
`CRSCompatibility` (`geospatial/crs.py:19`) is a frozen dataclass with `compatible`, `identical`,
|
| 261 |
+
`same_units`, `left`, `right`, `reason`, and a derived property:
|
| 262 |
+
|
| 263 |
+
```python
|
| 264 |
+
@property
|
| 265 |
+
def requires_reprojection(self) -> bool:
|
| 266 |
+
return self.compatible and not self.identical
|
| 267 |
+
```
|
| 268 |
+
|
| 269 |
+
`compare_crs` (`geospatial/crs.py:59`) decides:
|
| 270 |
+
|
| 271 |
+
- **A missing CRS on either side is NOT compatible.** The reason string is
|
| 272 |
+
`"one or both rasters lack a CRS; spatial comparison is unsafe"`. The docstring records that this is
|
| 273 |
+
*"not fatal for non-spatial tasks"*.
|
| 274 |
+
- **Identical CRS** (`crs_l.equals(crs_r)`) returns `identical=True`, reason `"identical CRS"`.
|
| 275 |
+
- **Same kind** (both projected or both geographic) returns `compatible=True, identical=False`, reason
|
| 276 |
+
`"reprojection required (<left name> -> <right name>)"`.
|
| 277 |
+
- **Mixed projected/geographic** returns `compatible=True, identical=False`, with a reason that adds
|
| 278 |
+
*"reprojection required and should be verified"*. The comment records the rationale: *"Mixing a
|
| 279 |
+
projected CRS with a geographic one is legal to reproject but signals a pipeline mistake."*
|
| 280 |
+
|
| 281 |
+
`same_units` compares the first axis's `unit_name` on each side (defaulting to `"unknown"` when there is
|
| 282 |
+
no axis info).
|
| 283 |
+
|
| 284 |
+
**Crucially: `compare_crs` never reprojects.** It reports whether reprojection is *required*; the actual
|
| 285 |
+
transform is not implemented anywhere (Β§13).
|
| 286 |
+
|
| 287 |
+
### 2.4 `is_metric`
|
| 288 |
+
|
| 289 |
+
```python
|
| 290 |
+
def is_metric(crs_value: str | None) -> bool:
|
| 291 |
+
crs = parse_crs(crs_value)
|
| 292 |
+
if crs is None:
|
| 293 |
+
return False
|
| 294 |
+
if not crs.is_projected:
|
| 295 |
+
return False
|
| 296 |
+
if not crs.axis_info:
|
| 297 |
+
return False
|
| 298 |
+
return crs.axis_info[0].unit_name in {"metre", "meter", "m"}
|
| 299 |
+
```
|
| 300 |
+
|
| 301 |
+
A geographic CRS is never metric here, and the unit check is a closed set of three spellings.
|
| 302 |
+
|
| 303 |
+
### 2.5 Where CRS comparison is used
|
| 304 |
+
|
| 305 |
+
The only consumer on the serving path is the change specialist's `_assess_pair`
|
| 306 |
+
(`specialists/change/specialist.py:251`):
|
| 307 |
+
|
| 308 |
+
```python
|
| 309 |
+
compat = compare_crs(t1.geo.crs, t2.geo.crs)
|
| 310 |
+
if not compat.compatible:
|
| 311 |
+
warnings.append(
|
| 312 |
+
f"CRS comparison unavailable ({compat.reason}); change is "
|
| 313 |
+
f"reported in normalized pixel coordinates only"
|
| 314 |
+
)
|
| 315 |
+
elif compat.requires_reprojection:
|
| 316 |
+
warnings.append(
|
| 317 |
+
f"T1 and T2 use different CRS ({compat.reason}); change is "
|
| 318 |
+
f"valid only if the rasters already share a pixel grid"
|
| 319 |
+
)
|
| 320 |
+
```
|
| 321 |
+
|
| 322 |
+
The comment above it states the policy precisely: *"No CRS on either side is NOT fatal for pixel-domain
|
| 323 |
+
change detection β the two rasters still tile to the same grid. It IS fatal for any claim about ground
|
| 324 |
+
coordinates, so the warning is recorded and the geospatial block is withheld downstream."* The
|
| 325 |
+
withholding happens in `ChangeSpecialist._geospatial`, which drops the transform when registration is
|
| 326 |
+
unusable (Β§9.1).
|
| 327 |
+
|
| 328 |
+
---
|
| 329 |
+
|
| 330 |
+
## 3. Affine transforms and coordinate frames
|
| 331 |
+
|
| 332 |
+
`geospatial/transform.py` is where a normalised 0β1 box, a pixel box and a geographic box are converted
|
| 333 |
+
between. Its module docstring states the non-negotiable rules (citing `docs/ARCHITECTURE_FREEZE.md`
|
| 334 |
+
Β§2.6):
|
| 335 |
+
|
| 336 |
+
> * CRS and affine transform are preserved, never silently stripped.
|
| 337 |
+
> * Conversions are explicit about their source and target frames.
|
| 338 |
+
> * A conversion that cannot be performed raises, rather than returning a plausible lie.
|
| 339 |
+
|
| 340 |
+
That last line is the module's governing principle, and it is why `to_normalized`, `to_pixel` and
|
| 341 |
+
`to_geo` all raise `CoordinateError` when they lack the context a conversion needs.
|
| 342 |
+
|
| 343 |
+
### 3.1 `PixelWindow` and `GeoBounds`
|
| 344 |
+
|
| 345 |
+
`PixelWindow` (`geospatial/transform.py:29`) is a frozen dataclass of `[col_min, row_min, col_max,
|
| 346 |
+
row_max]`. Its `__post_init__` **raises** on an inverted window:
|
| 347 |
+
|
| 348 |
+
```python
|
| 349 |
+
if self.col_max < self.col_min or self.row_max < self.row_min:
|
| 350 |
+
raise CoordinateError(f"pixel window is inverted: {self.as_list()}")
|
| 351 |
+
```
|
| 352 |
+
|
| 353 |
+
It exposes `width`, `height` and `area` properties. `GeoBounds` (`geospatial/transform.py:60`) is the
|
| 354 |
+
geographic cousin: `(minx, miny, maxx, maxy)`.
|
| 355 |
+
|
| 356 |
+
### 3.2 Normalised β pixel
|
| 357 |
+
|
| 358 |
+
`normalized_to_pixel(box, width, height)` multiplies each coordinate by the corresponding dimension. Its
|
| 359 |
+
docstring records a deliberate choice:
|
| 360 |
+
|
| 361 |
+
> Values are NOT clipped: a box slightly outside the frame is preserved so the caller can decide whether
|
| 362 |
+
> that is a bug or a legitimate edge case.
|
| 363 |
+
|
| 364 |
+
It raises `CoordinateError` if `len(box) != 4` or if either dimension is `<= 0`.
|
| 365 |
+
`pixel_to_normalized` is the inverse, dividing by the dimensions.
|
| 366 |
+
|
| 367 |
+
### 3.3 Pixel β geographic
|
| 368 |
+
|
| 369 |
+
`pixel_to_geo(window, transform)` applies the affine transform to two corners. The docstring explains the
|
| 370 |
+
axis convention:
|
| 371 |
+
|
| 372 |
+
> rasterio/affine transforms map (col, row) -> (x, y). The y axis usually points down in pixel space, so
|
| 373 |
+
> the geographic miny comes from the BOTTOM row.
|
| 374 |
+
|
| 375 |
+
It uses the non-deprecated matmul form `transform @ (col, row)` and then takes `min`/`max` across the two
|
| 376 |
+
corners so the returned `GeoBounds` is ordered. It raises `CoordinateError` if the transform is `None`.
|
| 377 |
+
|
| 378 |
+
`geo_to_pixel(bounds, transform)` inverts the transform with `~transform`, raising `CoordinateError` if
|
| 379 |
+
the inverse cannot be computed (`"affine transform is not invertible: ..."`).
|
| 380 |
+
|
| 381 |
+
### 3.4 Schema-aware conversions
|
| 382 |
+
|
| 383 |
+
`to_normalized`, `to_pixel` and `to_geo` operate on `Box` objects and are schema-aware β they read
|
| 384 |
+
`box.coordinate_system` and return a `Box` with the target system set. The context each requires is
|
| 385 |
+
documented in the docstring:
|
| 386 |
+
|
| 387 |
+
- `pixel -> normalized` requires `width` and `height`.
|
| 388 |
+
- `geo -> normalized` requires `width`, `height` **and** the affine `transform`.
|
| 389 |
+
|
| 390 |
+
Each returns early when the box is already in the target frame, and otherwise uses
|
| 391 |
+
`box.model_copy(update={...})` to produce a new box with the new coordinates and the new
|
| 392 |
+
`coordinate_system`.
|
| 393 |
+
|
| 394 |
+
### 3.5 Intersection and IoU
|
| 395 |
+
|
| 396 |
+
`intersect(a, b)` returns the axis-aligned intersection of two `PixelWindow`s, or `None` when they are
|
| 397 |
+
disjoint. `iou(a, b)` computes intersection-over-union in `[0, 1]`, returning `0.0` for a disjoint pair
|
| 398 |
+
or a zero-area union. These are used by the grounding NMS path and by change-region overlap.
|
| 399 |
+
|
| 400 |
+
### 3.6 `affine_from_metadata`
|
| 401 |
+
|
| 402 |
+
```python
|
| 403 |
+
def affine_from_metadata(geo: GeoMetadata) -> Affine | None:
|
| 404 |
+
if not geo.transform or len(geo.transform) != 6:
|
| 405 |
+
return None
|
| 406 |
+
a, b, c, d, e, f = geo.transform
|
| 407 |
+
return Affine(a, b, c, d, e, f)
|
| 408 |
+
```
|
| 409 |
+
|
| 410 |
+
A `GeoMetadata` whose transform is missing or not exactly six values yields `None`; callers treat that as
|
| 411 |
+
"cannot georeference" rather than as an error. This is the bridge the grounding specialist uses in
|
| 412 |
+
`_to_geo_boxes` (Β§4.3).
|
| 413 |
+
|
| 414 |
+
---
|
| 415 |
+
|
| 416 |
+
## 4. The coordinate-system contract
|
| 417 |
+
|
| 418 |
+
### 4.1 The enum
|
| 419 |
+
|
| 420 |
+
`CoordinateSystem` (`core/schemas.py:57`) has exactly three members, and the class docstring is a warning
|
| 421 |
+
in itself β *"Never omit this. A bare box is meaningless without it. (C-5)"*:
|
| 422 |
+
|
| 423 |
+
```python
|
| 424 |
+
class CoordinateSystem(str, Enum):
|
| 425 |
+
NORMALIZED_0_1 = "normalized_0_1"
|
| 426 |
+
PIXEL = "pixel"
|
| 427 |
+
GEO = "geo"
|
| 428 |
+
```
|
| 429 |
+
|
| 430 |
+
`docs/API_CONTRACT.md` Β§3.2 fixes the meaning of each value:
|
| 431 |
+
|
| 432 |
+
| Value | Meaning | Rendering |
|
| 433 |
+
|---|---|---|
|
| 434 |
+
| `normalized_0_1` | `[0, 1]`, origin **top-left** | Multiply by image width/height |
|
| 435 |
+
| `pixel` | Absolute pixel coordinates | Use directly |
|
| 436 |
+
| `geo` | CRS coordinates (usually EPSG:4326) | Requires a map, not a 2-D canvas |
|
| 437 |
+
|
| 438 |
+
The contract adds a hard warning: *"The frontend must read the `coordinate_system` field on each
|
| 439 |
+
`Box`/`Region` and must not assume one convention. A box drawn with the wrong assumption lands in
|
| 440 |
+
plausible-looking wrong places."* It also records that earlier drafts used shorthand (`normalized`,
|
| 441 |
+
`geographic`) and that those strings are **wrong** β they would fail schema validation, because the
|
| 442 |
+
models are `extra="forbid"` and the enum is closed.
|
| 443 |
+
|
| 444 |
+
### 4.2 Why it is mandatory β the schema validator
|
| 445 |
+
|
| 446 |
+
The `Box` model (`core/schemas.py:171`) defaults `coordinate_system` to
|
| 447 |
+
`CoordinateSystem.NORMALIZED_0_1`, and its `_ordered` validator rejects a box whose coordinates do not
|
| 448 |
+
satisfy `x2 >= x1` and `y2 >= y1`. `Region` (`core/schemas.py:189`) and `ChangeRegion`
|
| 449 |
+
(`core/schemas.py:202`) both carry the same field with the same default.
|
| 450 |
+
|
| 451 |
+
The enforcement that matters is on `Evidence`. `Evidence._spatial_needs_crs`
|
| 452 |
+
(`core/schemas.py:238`) **rejects a spatial evidence item that carries coordinates but no
|
| 453 |
+
coordinate system**:
|
| 454 |
+
|
| 455 |
+
```python
|
| 456 |
+
@model_validator(mode="after")
|
| 457 |
+
def _spatial_needs_crs(self) -> "Evidence":
|
| 458 |
+
spatial = {
|
| 459 |
+
EvidenceType.BOUNDING_BOX,
|
| 460 |
+
EvidenceType.MASK,
|
| 461 |
+
EvidenceType.CHANGE_MAP,
|
| 462 |
+
EvidenceType.TILE,
|
| 463 |
+
EvidenceType.IMAGE_CROP,
|
| 464 |
+
EvidenceType.JOINT_FEATURE_REGION,
|
| 465 |
+
}
|
| 466 |
+
if self.type in spatial and self.coordinates and self.coordinate_system is None:
|
| 467 |
+
raise ValueError(
|
| 468 |
+
f"evidence type '{self.type.value}' carries coordinates "
|
| 469 |
+
"but no coordinate_system"
|
| 470 |
+
)
|
| 471 |
+
return self
|
| 472 |
+
```
|
| 473 |
+
|
| 474 |
+
So a bare box cannot be published as evidence. The validator fires only when the item both is of a
|
| 475 |
+
spatial type **and** carries coordinates β a `STATISTIC` or `GEOLOCATION` item with no coordinates is
|
| 476 |
+
unaffected, and a spatial item with no coordinates (e.g. an availability-mask record) is also unaffected.
|
| 477 |
+
`GEOLOCATION` is deliberately **not** in the spatial set: a geolocation is a point reference, not a
|
| 478 |
+
rectangle, and the grounding specialist sets its `coordinate_system` to `GEO` explicitly anyway.
|
| 479 |
+
|
| 480 |
+
Note the interaction with `extra="forbid"`: `Evidence` forbids extra keys, so an author cannot smuggle a
|
| 481 |
+
convention through an undeclared field. `GeoMetadata`, by contrast, is `extra="allow"`, which is how the
|
| 482 |
+
optical-SAR specialist attaches cross-sensor keys (Β§4.3).
|
| 483 |
+
|
| 484 |
+
### 4.3 What each specialist emits
|
| 485 |
+
|
| 486 |
+
| Specialist | Boxes / regions | Evidence coordinate systems |
|
| 487 |
+
|---|---|---|
|
| 488 |
+
| `grounding` | `Box` and `Region` in `NORMALIZED_0_1` | `BOUNDING_BOX` in `NORMALIZED_0_1`; a `GEOLOCATION` item in `GEO` when the CRS allows; `STATISTIC` items with no coordinates |
|
| 489 |
+
| `change` | `ChangeRegion` and `Region` in `NORMALIZED_0_1` | `CHANGE_MAP` and `BOUNDING_BOX` in `NORMALIZED_0_1`; `STATISTIC` items with no coordinates |
|
| 490 |
+
| `optical_sar` | No `regions` (empty list) | `AVAILABILITY_MASK` items (no coordinates); `OPTICAL_VIEW` / `SAR_VIEW` with `coordinates=[0,0,1,1]` in `NORMALIZED_0_1`; a `JOINT_FEATURE_REGION` in `NORMALIZED_0_1`; `STATISTIC` items |
|
| 491 |
+
|
| 492 |
+
**Grounding.** `GroundingSpecialist.execute` (`specialists/grounding/specialist.py:253`) builds every box
|
| 493 |
+
with an explicit system:
|
| 494 |
+
|
| 495 |
+
```python
|
| 496 |
+
boxes = [
|
| 497 |
+
Box(
|
| 498 |
+
x1=d.box[0], y1=d.box[1], x2=d.box[2], y2=d.box[3],
|
| 499 |
+
score=d.score, label=request.query,
|
| 500 |
+
coordinate_system=CoordinateSystem.NORMALIZED_0_1,
|
| 501 |
+
)
|
| 502 |
+
for d in decoded
|
| 503 |
+
]
|
| 504 |
+
```
|
| 505 |
+
|
| 506 |
+
Geographic boxes are **derived, never replacements**: `_to_geo_boxes` (`specialists/grounding/specialist.py:431`)
|
| 507 |
+
converts the normalised boxes to geographic bounds *only* when the raster has a CRS, a transform and
|
| 508 |
+
dimensions, and emits a `GEOLOCATION` evidence item with `coordinate_system=CoordinateSystem.GEO`. When
|
| 509 |
+
the raster has no CRS/transform, the method appends a warning β
|
| 510 |
+
*"raster has no CRS/transform; boxes are reported in normalized coordinates only and cannot be placed on a
|
| 511 |
+
map"* β and returns an empty list. The result's boxes stay normalised; the geo boxes are additional
|
| 512 |
+
evidence for a consumer with a CRS. Its own comment states the rule: *"Pixel and geo coordinates are
|
| 513 |
+
DERIVED, never replacements."*
|
| 514 |
+
|
| 515 |
+
**Change.** `regions_to_schema` (`specialists/change/postprocess.py:443`) builds `ChangeRegion` objects in
|
| 516 |
+
the requested system. It supports `NORMALIZED_0_1` and `PIXEL` and **raises** for anything else:
|
| 517 |
+
|
| 518 |
+
```python
|
| 519 |
+
else:
|
| 520 |
+
raise ValueError(
|
| 521 |
+
f"regions_to_schema cannot produce {coordinate_system.value} "
|
| 522 |
+
f"coordinates; convert after this call"
|
| 523 |
+
)
|
| 524 |
+
```
|
| 525 |
+
|
| 526 |
+
The change specialist calls it with the default (`NORMALIZED_0_1`). Geographic conversion is a deliberate
|
| 527 |
+
post-step β the docstring records: *"Callers that want geographic coordinates convert afterwards via
|
| 528 |
+
`geospatial.transform`, which keeps this function free of rasterio and CRS concerns."*
|
| 529 |
+
|
| 530 |
+
**Optical-SAR.** The specialist emits no `regions`. Its view evidence items carry
|
| 531 |
+
`coordinates=[0.0, 0.0, 1.0, 1.0]` in `NORMALIZED_0_1` β the whole frame β because the view is a
|
| 532 |
+
presentation of the canonical tensor, not a localisation. Its `geospatial` block is the **optical**
|
| 533 |
+
asset's georeferencing (chosen because it is the asset a reader can navigate by), with cross-sensor keys
|
| 534 |
+
added via `GeoMetadata`'s `extra="allow"`: `optical_crs`, `sar_crs`, `sar_bounds`, `sar_resolution`,
|
| 535 |
+
`optical_availability`, `sar_availability`, `optical_sensor`, `sar_sensor`, and
|
| 536 |
+
`resolution_ratio_sar_to_optical` when both resolutions are known
|
| 537 |
+
(`specialists/optical_sar/specialist.py:597`). The comment records the reason the two footprints are
|
| 538 |
+
carried separately rather than fused: *"RISAT and Cartosat-2S are different missions with different
|
| 539 |
+
orbits."*
|
| 540 |
+
|
| 541 |
+
---
|
| 542 |
+
|
| 543 |
+
## 5. Benchmark box scale β 0β100 versus 0β1
|
| 544 |
+
|
| 545 |
+
This is a small contract with an outsized failure mode, and the codebase treats it as such.
|
| 546 |
+
|
| 547 |
+
### 5.1 The two conventions
|
| 548 |
+
|
| 549 |
+
- **VRSBench** (the grounding benchmark) states verbatim that *"all box coordinates are normalized to
|
| 550 |
+
0-100"*.
|
| 551 |
+
- **This project** stores every box normalised to **0β1**.
|
| 552 |
+
|
| 553 |
+
Treating a 0β100 box as a pixel box, or as a 0β1 box, is a silent, catastrophic bug: the numbers are all
|
| 554 |
+
in range, the schema accepts them, and the box lands in the wrong place.
|
| 555 |
+
|
| 556 |
+
### 5.2 The conversion functions
|
| 557 |
+
|
| 558 |
+
`geospatial/transform.py:286` provides the conversion, and its docstring names the failure mode it exists
|
| 559 |
+
to prevent:
|
| 560 |
+
|
| 561 |
+
```python
|
| 562 |
+
def benchmark_boxes_to_normalized(
|
| 563 |
+
boxes: Iterable[Sequence[float]], scale: float = 100.0
|
| 564 |
+
) -> list[list[float]]:
|
| 565 |
+
"""Convert benchmark boxes (e.g. VRSBench's 0-100) to our internal 0-1.
|
| 566 |
+
|
| 567 |
+
Finding from Phase 0: VRSBench states verbatim that "all box coordinates are
|
| 568 |
+
normalized to 0-100". Treating those numbers as pixels is a silent, catastrophic
|
| 569 |
+
bug β this function exists so that mistake can be made once, explicitly.
|
| 570 |
+
"""
|
| 571 |
+
if scale <= 0:
|
| 572 |
+
raise CoordinateError(f"box scale must be positive, got {scale}")
|
| 573 |
+
...
|
| 574 |
+
out.append([float(v) / scale for v in b])
|
| 575 |
+
```
|
| 576 |
+
|
| 577 |
+
`normalized_boxes_to_benchmark` is the exact inverse (multiply by `scale`). Both raise on a
|
| 578 |
+
non-positive scale and on a box that is not four values.
|
| 579 |
+
|
| 580 |
+
A second, independent conversion lives in the evaluation metrics:
|
| 581 |
+
`evaluation/metrics/grounding.py::benchmark_to_normalized(box, scale=100.0)` divides by the scale and
|
| 582 |
+
raises `ValueError(f"scale must be positive, got {scale}")` on a non-positive value. `docs/API_CONTRACT.md`
|
| 583 |
+
Β§3.2 names it explicitly as the site of the VRSBench conversion.
|
| 584 |
+
|
| 585 |
+
### 5.3 The declared scale
|
| 586 |
+
|
| 587 |
+
`configs/base.yaml` declares the conversion factor in the grounding block, with a comment that restates
|
| 588 |
+
the convention:
|
| 589 |
+
|
| 590 |
+
```yaml
|
| 591 |
+
grounding:
|
| 592 |
+
...
|
| 593 |
+
# VRSBench stores boxes normalised to 0-100; we store 0-1.
|
| 594 |
+
benchmark_box_scale: 100.0
|
| 595 |
+
coordinate_system: normalized_0_1
|
| 596 |
+
```
|
| 597 |
+
|
| 598 |
+
`benchmark_box_scale: 100.0` is therefore the **declared** value, and `100.0` is also the **default
|
| 599 |
+
argument** in both conversion functions. The two agree.
|
| 600 |
+
|
| 601 |
+
### 5.4 The double-application risk
|
| 602 |
+
|
| 603 |
+
Because the scale appears in three places β the config key, the `geospatial.transform` default and the
|
| 604 |
+
`evaluation.metrics.grounding` default β there is a real hazard that a caller converts 0β100 β 0β1 and
|
| 605 |
+
then a downstream stage divides by 100 again, producing a box 100Γ too small in each coordinate. The
|
| 606 |
+
codebase's mitigations are structural rather than documentary:
|
| 607 |
+
|
| 608 |
+
1. **The conversion is named for its direction.** `benchmark_boxes_to_normalized` and
|
| 609 |
+
`normalized_boxes_to_benchmark` are inverses with self-describing names, so a reader cannot easily
|
| 610 |
+
apply the wrong one without noticing.
|
| 611 |
+
2. **The default is the correct scale, not `1.0`.** A caller that forgets to pass `scale` gets `100.0`,
|
| 612 |
+
which is right for VRSBench and would produce an obviously wrong box (values > 1) for already-0β1
|
| 613 |
+
input β a loud failure, not a silent one.
|
| 614 |
+
3. **The schema rejects out-of-range values indirectly.** A `Box` built from un-divided VRSBench
|
| 615 |
+
coordinates would have `score` in range but coordinates in `[0, 100]`; the `_ordered` validator would
|
| 616 |
+
pass it (the ordering is still valid), so the schema alone does **not** catch this. The catch is that
|
| 617 |
+
the metrics module's own `benchmark_to_normalized` is the documented entry point, and the resolution
|
| 618 |
+
experiment and the grounding evaluation both go through it.
|
| 619 |
+
|
| 620 |
+
**Residual risk β stated honestly.** There is no runtime assertion that a `Box` in
|
| 621 |
+
`NORMALIZED_0_1` actually has coordinates within `[0, 1]`. `Box._ordered` checks only ordering. A
|
| 622 |
+
mis-scaled box would therefore pass schema validation. `UNKNOWN β not established from the available
|
| 623 |
+
evidence` whether any production call site currently constructs a `Box` without going through one of the
|
| 624 |
+
two conversion functions; the serving path (`GroundingSpecialist.execute`) builds boxes from the decoder's
|
| 625 |
+
output, which is already in `[0, 1]`, so the serving path is not exposed. The risk is confined to
|
| 626 |
+
evaluation and experiment code that reads VRSBench directly.
|
| 627 |
+
|
| 628 |
+
---
|
| 629 |
+
|
| 630 |
+
## 6. Band-count modality inference
|
| 631 |
+
|
| 632 |
+
### 6.1 The heuristics
|
| 633 |
+
|
| 634 |
+
`preprocessing/raster.py` defines two closed sets:
|
| 635 |
+
|
| 636 |
+
```python
|
| 637 |
+
_OPTICAL_BAND_COUNTS = {3, 4, 8, 11, 12, 13}
|
| 638 |
+
_SAR_BAND_COUNTS = {1, 2}
|
| 639 |
+
```
|
| 640 |
+
|
| 641 |
+
with the comment: *"Band-count heuristics for modality detection when explicit metadata is absent. These
|
| 642 |
+
are heuristics, not ground truth β the sensor adapter is authoritative when a sensor descriptor is
|
| 643 |
+
supplied."*
|
| 644 |
+
|
| 645 |
+
### 6.2 `infer_modality`
|
| 646 |
+
|
| 647 |
+
```python
|
| 648 |
+
def infer_modality(band_count: int, explicit: str | None = None) -> Modality:
|
| 649 |
+
if explicit:
|
| 650 |
+
try:
|
| 651 |
+
return Modality(explicit.lower())
|
| 652 |
+
except ValueError:
|
| 653 |
+
pass
|
| 654 |
+
|
| 655 |
+
if band_count in _SAR_BAND_COUNTS:
|
| 656 |
+
return Modality.SAR
|
| 657 |
+
if band_count in _OPTICAL_BAND_COUNTS:
|
| 658 |
+
return Modality.OPTICAL
|
| 659 |
+
# 12-band optical and 2-band SAR are both plausible; default to optical only
|
| 660 |
+
# when the count clearly favours it.
|
| 661 |
+
if band_count >= 4:
|
| 662 |
+
return Modality.OPTICAL
|
| 663 |
+
return Modality.UNKNOWN
|
| 664 |
+
```
|
| 665 |
+
|
| 666 |
+
The decision order is worth reading carefully:
|
| 667 |
+
|
| 668 |
+
1. An **explicit** label wins if it parses to a `Modality` member. An explicit label that does *not* parse
|
| 669 |
+
(e.g. `"radar"`) is silently ignored and inference proceeds β the `except ValueError: pass` is
|
| 670 |
+
deliberate.
|
| 671 |
+
2. `{1, 2}` β `SAR`.
|
| 672 |
+
3. `{3, 4, 8, 11, 12, 13}` β `OPTICAL`.
|
| 673 |
+
4. `>= 4` (and not already matched) β `OPTICAL`.
|
| 674 |
+
5. Otherwise β `UNKNOWN`.
|
| 675 |
+
|
| 676 |
+
So the effective mapping is: 1β2 β SAR; 3β13 β optical (by set or by the `>= 4` fallback); 0 or a
|
| 677 |
+
non-positive count β `UNKNOWN` (though `inspect_raster` rejects `band_count <= 0` before this point).
|
| 678 |
+
|
| 679 |
+
### 6.3 The declared-wins rule
|
| 680 |
+
|
| 681 |
+
In the optical-SAR specialist, `_assess_pair` (`specialists/optical_sar/specialist.py:223`) uses the
|
| 682 |
+
asset's **declared** modality first, falling back to band-count inference only when it is `UNKNOWN`:
|
| 683 |
+
|
| 684 |
+
```python
|
| 685 |
+
for asset in assets:
|
| 686 |
+
modality = asset.modality
|
| 687 |
+
if modality is Modality.UNKNOWN:
|
| 688 |
+
inferred = self._infer_modality(asset, warnings)
|
| 689 |
+
resolved.append(inferred)
|
| 690 |
+
else:
|
| 691 |
+
resolved.append(modality)
|
| 692 |
+
```
|
| 693 |
+
|
| 694 |
+
`_infer_modality` appends a warning that names the heuristic for what it is: *"A band count is a
|
| 695 |
+
heuristic, not a sensor declaration β label the asset explicitly if this is wrong."* The docstring above
|
| 696 |
+
`_assess_pair` states the rationale: *"The declared value wins when present, because the caller may know
|
| 697 |
+
something the band count does not β a 2-band Cartosat stack, or a 12-band decomposition product."*
|
| 698 |
+
|
| 699 |
+
### 6.4 Why the heuristic is not authoritative
|
| 700 |
+
|
| 701 |
+
The comment on the sets and the warning text both point at the same thing: a band count does not identify
|
| 702 |
+
a *sensor*. A 2-band file could be a dual-pol SAR pair or a two-band optical stack; a 4-band file could be
|
| 703 |
+
RGB+NIR or a four-polarisation SAR product. The adapter (Β§8) is the authoritative path, because it maps
|
| 704 |
+
named bands onto CROMA's canonical channels and records what it could not place.
|
| 705 |
+
|
| 706 |
+
---
|
| 707 |
+
|
| 708 |
+
## 7. Radiometric normalisation and representation
|
| 709 |
+
|
| 710 |
+
This section covers a genuine specification gap and its partial resolution. It is presented in full
|
| 711 |
+
because the gap is load-bearing and the honest state is more useful than a tidy summary.
|
| 712 |
+
|
| 713 |
+
### 7.1 What configuration declares
|
| 714 |
+
|
| 715 |
+
`configs/base.yaml` declares two radiometric transforms:
|
| 716 |
+
|
| 717 |
+
```yaml
|
| 718 |
+
optical:
|
| 719 |
+
normalization: percentile
|
| 720 |
+
lower_percentile: 2
|
| 721 |
+
upper_percentile: 98
|
| 722 |
+
# CROMA expects exactly 12 optical channels.
|
| 723 |
+
canonical_channels: 12
|
| 724 |
+
|
| 725 |
+
sar:
|
| 726 |
+
representation: db
|
| 727 |
+
clip_min_db: -30
|
| 728 |
+
clip_max_db: 5
|
| 729 |
+
# CROMA expects exactly 2 SAR channels (VV, VH).
|
| 730 |
+
canonical_channels: 2
|
| 731 |
+
```
|
| 732 |
+
|
| 733 |
+
So the declared optical normalisation is a **percentile stretch at the 2nd and 98th percentiles**, and the
|
| 734 |
+
declared SAR representation is **decibels clipped to [β30, +5]**.
|
| 735 |
+
|
| 736 |
+
### 7.2 The DEV-2 ruling β two ordered stages
|
| 737 |
+
|
| 738 |
+
`docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` records the ruling that governs this. The gap it confirmed:
|
| 739 |
+
CROMA's own example preprocessing is a **per-channel** dynamic-range stretch
|
| 740 |
+
(`mean Β± 2Β·std β [0, 1]`, optionally through a uint8 round-trip), while `configs/base.yaml` specifies
|
| 741 |
+
something structurally different. The ruling's verification table records that a grep of
|
| 742 |
+
`specialists/optical_sar/` for `NORM_*`, `clip_min_db`, `lower_percentile`, `use_8_bit` returned **six
|
| 743 |
+
hits, all constant definitions/labels** at `sensor_adapter.py:125,126,210,323,547,548`, and **zero
|
| 744 |
+
application sites**.
|
| 745 |
+
|
| 746 |
+
The ruling resolves this by declaring the pipeline to have **two ordered stages**, and it states
|
| 747 |
+
explicitly that they are *not alternatives*:
|
| 748 |
+
|
| 749 |
+
| Stage | Transform | Owner | Status |
|
| 750 |
+
|---|---|---|---|
|
| 751 |
+
| **Radiometric conditioning** | per-scene percentile / dB-clip, per `configs/base.yaml` | preprocessing / sensor adapter | **specified in the ruling, must be implemented** |
|
| 752 |
+
| **Encoder-input normalisation** | per-channel `mean Β± 2Β·std` β `[0,1]`, then optional `Γ255` uint8 β `/255` | a new function consumed by `croma.py` | **mandatory, immediate** |
|
| 753 |
+
|
| 754 |
+
The rationale for the CROMA stage being mandatory: *"it is the only stage that bounds the input range.
|
| 755 |
+
Without it the encoder receives unbounded input."* The `use_8_bit` branch is the safer default because
|
| 756 |
+
the uint8 round-trip **guarantees** `[0, 1]`, whereas the float path's clip applies to a stretch that is
|
| 757 |
+
unbounded by construction (`mean Β± 2Β·std` is not a min/max clamp). The decision on `use_8_bit` is
|
| 758 |
+
**ENABLED**.
|
| 759 |
+
|
| 760 |
+
### 7.3 The implemented stage β `specialists/optical_sar/radiometry.py`
|
| 761 |
+
|
| 762 |
+
The **encoder-input** stage is implemented, in full, at `specialists/optical_sar/radiometry.py`. Its
|
| 763 |
+
module docstring states the boundary between what is and is not implemented:
|
| 764 |
+
|
| 765 |
+
> Implemented: the **encoder-input** stage β per-channel `mean +/- 2*std` -> `[0, 1]`, the transform the
|
| 766 |
+
> CROMA authors' own README instructs users to apply.
|
| 767 |
+
>
|
| 768 |
+
> NOT implemented: the percentile / dB **conditioning** stage that `base.yaml` names.
|
| 769 |
+
|
| 770 |
+
The transform is quoted verbatim from the CROMA README in the docstring:
|
| 771 |
+
|
| 772 |
+
```
|
| 773 |
+
min_value = x[:, c].mean() - 2 * x[:, c].std()
|
| 774 |
+
max_value = x[:, c].mean() + 2 * x[:, c].std()
|
| 775 |
+
img = (x[:, c] - min_value) / (max_value - min_value) * 255.0
|
| 776 |
+
img = clip(img, 0, 255).to(uint8)
|
| 777 |
+
# before the forward pass:
|
| 778 |
+
x = x.float() / 255
|
| 779 |
+
```
|
| 780 |
+
|
| 781 |
+
The module makes two **deliberate deviations** from the README, both documented:
|
| 782 |
+
|
| 783 |
+
1. **Per sample, not per batch.** The README computes `x[:, c].mean()` over the whole batch, which makes
|
| 784 |
+
one image's encoding a function of its neighbours. This module computes the window **per sample**,
|
| 785 |
+
because *"a serving system whose whole point is reproducibility"* must not produce a different answer
|
| 786 |
+
because another request was batched alongside it. At `N = 1` the two are identical, and the README's
|
| 787 |
+
own example runs at `N = 1`.
|
| 788 |
+
2. **Unbiased standard deviation (`ddof=1`).** Upstream is torch, whose `.std()` defaults to the unbiased
|
| 789 |
+
estimator; NumPy's defaults to `ddof=0`. The module uses `ddof=1` to match the reference
|
| 790 |
+
implementation, and states the numerical immateriality (~1.00003 factor at 120Γ120) so nobody has to
|
| 791 |
+
guess which convention a number came from.
|
| 792 |
+
|
| 793 |
+
The key constants:
|
| 794 |
+
|
| 795 |
+
| Constant | Value | Meaning |
|
| 796 |
+
|---|---|---|
|
| 797 |
+
| `TRANSFORM_NAME` | `"per_channel_mean_pm_2std"` | Recorded on every trace |
|
| 798 |
+
| `WINDOW_SIGMAS` | `2.0` | The window half-width |
|
| 799 |
+
| `DEFAULT_USE_8_BIT` | `True` | The ruling's decision |
|
| 800 |
+
| `ENV_USE_8_BIT` | `"SATQUERY_CROMA_USE_8_BIT"` | Hash-exempt override channel |
|
| 801 |
+
| `MIN_FINITE_PIXELS` | `2` | Below this, a std is not meaningful |
|
| 802 |
+
| `STATUS_NORMALISED` / `STATUS_UNAVAILABLE` / `STATUS_DEGENERATE` | strings | Per-channel status |
|
| 803 |
+
|
| 804 |
+
`normalise_for_croma(array, *, use_8_bit, modality)` is **pure and deterministic** β the docstring
|
| 805 |
+
states: *"no sampling, no learned statistics, no RNG, and no dependence on batch composition. Same input
|
| 806 |
+
-> same output."* It accepts `(B, C, H, W)` or `(C, H, W)`, returns a `float32` array of the same shape
|
| 807 |
+
plus a `RadiometryReport`, and leaves skipped channels **exactly zero**.
|
| 808 |
+
|
| 809 |
+
### 7.4 The zero-channel rule
|
| 810 |
+
|
| 811 |
+
This is the load-bearing part, and it is the C-1 discipline applied to radiometry. An unavailable channel
|
| 812 |
+
is all zeros (the sensor adapter's convention). Such a channel has `std = 0`, so `max_value - min_value
|
| 813 |
+
== 0` and the stretch divides by zero. The module therefore **skips and leaves at exactly zero** any
|
| 814 |
+
channel with no dynamic range, and records a **different status** for each reason:
|
| 815 |
+
|
| 816 |
+
- `unavailable` β the channel is all zeros (`if not np.any(channel)`), i.e. the sensor did not measure it.
|
| 817 |
+
- `degenerate` β the channel is present but constant (`_channel_window` returns `None` because
|
| 818 |
+
`upper - lower <= 0`), i.e. the sensor measured a flat band.
|
| 819 |
+
- `normalised` β the real case.
|
| 820 |
+
|
| 821 |
+
The docstring explains why the distinction matters: *"'the sensor did not measure this' is not the same
|
| 822 |
+
fact as 'the sensor measured a flat band', and the trace must not merge them."* The mechanism is the
|
| 823 |
+
degenerate-window test, which needs **no mask** β deliberately, because `CROMAEncoder.encode` has a
|
| 824 |
+
pinned signature of exactly `(self, optical, sar)` and finding C-1 says CROMA never receives a mask.
|
| 825 |
+
Inferring unavailability from the data keeps that guard intact.
|
| 826 |
+
|
| 827 |
+
A non-finite pixel (NaN/Inf) is *"a defective measurement, not a measurement of zero"*; it is pinned to
|
| 828 |
+
`0.0` rather than allowed to propagate a NaN into the frozen encoder, where *"one NaN would poison the
|
| 829 |
+
whole forward pass."* The `finite_pixels` count in the report keeps the gap detectable.
|
| 830 |
+
|
| 831 |
+
`resolve_use_8_bit(config)` resolves the flag in a documented order β environment variable first, then
|
| 832 |
+
`croma.use_8_bit` when the key exists and is a bool, then `DEFAULT_USE_8_BIT` β and returns
|
| 833 |
+
`(value, source)` where `source` is `"env"`, `"config"` or `"default"`. The reason it is not simply a
|
| 834 |
+
config key is recorded: adding `croma.use_8_bit` to `configs/base.yaml` would move `Config.hash` away
|
| 835 |
+
from the frozen `78f1e3700da15aa1` and detach the Phase 9 benchmark. The hash-exempt channel is the same
|
| 836 |
+
pattern already used for `SATQUERY_DEVICE`.
|
| 837 |
+
|
| 838 |
+
`summarise(optical_report, sar_report)` bundles both modalities under one `use_8_bit` and **asserts** they
|
| 839 |
+
agree, raising `ValueError` if they do not: *"upstream applies one transform to both."*
|
| 840 |
+
|
| 841 |
+
### 7.5 The display stretches β a different transform
|
| 842 |
+
|
| 843 |
+
Two other percentile stretches exist in the codebase. Neither is the CROMA transform, and the DEV-2
|
| 844 |
+
ruling names reusing one of them as **forbidden**.
|
| 845 |
+
|
| 846 |
+
**`preprocessing/imagery.py::load_image_array`** produces a displayable `(H, W, 3)` uint8 array:
|
| 847 |
+
|
| 848 |
+
```python
|
| 849 |
+
arr = array.astype(np.float32)
|
| 850 |
+
finite = arr[np.isfinite(arr)]
|
| 851 |
+
if finite.size:
|
| 852 |
+
lo, hi = np.percentile(finite, (2, 98))
|
| 853 |
+
if hi > lo:
|
| 854 |
+
arr = (arr - lo) / (hi - lo)
|
| 855 |
+
arr = np.clip(arr, 0.0, 1.0)
|
| 856 |
+
return (arr * 255).astype(np.uint8)
|
| 857 |
+
```
|
| 858 |
+
|
| 859 |
+
Its docstring is explicit about its scope: it *"does not resample, crop, or reproject. Those change the
|
| 860 |
+
pixel grid, and the grounding specialist converts normalized boxes to pixel coordinates using the
|
| 861 |
+
ORIGINAL raster's dimensions β a silent resize here would put every box in the wrong place."* The
|
| 862 |
+
determinism claim is stated: *"the same file always yields the same array, so a grounding box and a VQA
|
| 863 |
+
answer describe identical pixels."*
|
| 864 |
+
|
| 865 |
+
`docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§1.1 records the three disqualifying differences between
|
| 866 |
+
this stretch and CROMA's: it is **per-scene** (one `(lo, hi)` over all bands) rather than **per-channel**;
|
| 867 |
+
it returns **uint8 3-band** rather than float 12-channel; and its **purpose** is presentation, not
|
| 868 |
+
radiometry.
|
| 869 |
+
|
| 870 |
+
**`specialists/change/postprocess.py::to_grayscale_float`** applies a per-scene 2/98 percentile stretch
|
| 871 |
+
to a grayscale collapse, so that *"a uint16 raster and a uint8 raster of the same scene correlate
|
| 872 |
+
identically."* Its comment records a measured defect fixed in place: `np.percentile` returns float64
|
| 873 |
+
scalars even for float32 input, so the array is promoted; without the explicit
|
| 874 |
+
`.astype(np.float32, copy=False)` the function returns float64 and `cv2.phaseCorrelate` fails its own
|
| 875 |
+
type assertion against a `CV_32F` Hanning window β *"Measured, not theorised."*
|
| 876 |
+
|
| 877 |
+
### 7.6 The unimplemented stage, and its consequence
|
| 878 |
+
|
| 879 |
+
The percentile/dB conditioning stage is **not implemented**. The radiometry module's docstring states the
|
| 880 |
+
consequence plainly:
|
| 881 |
+
|
| 882 |
+
> **Consequence: the two `optical.*` and three `sar.*` config keys are still read by no code.** That is
|
| 883 |
+
> recorded, not fixed.
|
| 884 |
+
|
| 885 |
+
`docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§6 item 3 adds: *"A config key that no code reads is worse
|
| 886 |
+
than a missing key: it reads as a satisfied requirement."* The ruling's follow-on table assigns the
|
| 887 |
+
implementation to the engineer, with the alternative of explicitly retiring the keys.
|
| 888 |
+
|
| 889 |
+
The status of the whole chain is therefore:
|
| 890 |
+
|
| 891 |
+
| Stage | Declared | Implemented | Read by code |
|
| 892 |
+
|---|---|---|---|
|
| 893 |
+
| Optical percentile 2/98 conditioning | β
`optical.*` | β | β |
|
| 894 |
+
| SAR dB clip β30/+5 conditioning | β
`sar.*` | β | β |
|
| 895 |
+
| Per-channel `mean Β± 2Β·std` β [0,1] | β
(ruling Β§3.1) | β
`radiometry.py` | β
(via `training/fusion/extract.py`; see below) |
|
| 896 |
+
|
| 897 |
+
**Where the implemented stretch actually runs.** The serving-path pipeline
|
| 898 |
+
(`specialists/optical_sar/inference.py::run_pipeline`) does **not** call `normalise_for_croma`. It
|
| 899 |
+
canonicalises, resizes to 120Γ120 and calls `encoder.encode(...)`. The stretch is applied by the
|
| 900 |
+
**extraction** path (`training/fusion/extract.py`), which states: *"This pipeline applies the DEV-2
|
| 901 |
+
encoder-input stretch itself (`radiometry.normalise_for_croma`), because the arm's `use_8_bit` has to be
|
| 902 |
+
applied at exactly one place."* `extract.py` also refuses an encoder that would stretch a second time:
|
| 903 |
+
`_assert_raw_encoder` *"refuses it by name rather than trusting the caller."* So the implemented
|
| 904 |
+
radiometric transform is a **training-time** transform that feeds the cached features, and the serving
|
| 905 |
+
path's input scaling is whatever `CROMAEncoder.encode` does internally (see Β§8.9 and the CROMA module).
|
| 906 |
+
|
| 907 |
+
`UNKNOWN β not established from the available evidence` whether the serving path's encoder applies the
|
| 908 |
+
same `mean Β± 2Β·std` stretch internally; `specialists/optical_sar/croma.py::CROMAEncoder.encode` was
|
| 909 |
+
verified to have the pinned signature `(self, optical, sar)` with no mask parameter, and the extraction
|
| 910 |
+
module's guard implies the encoder *can* apply a stretch when `normalize_input=True` (its default), but
|
| 911 |
+
the exact serving configuration of that flag is not established from the files read for this chapter.
|
| 912 |
+
|
| 913 |
+
---
|
| 914 |
+
|
| 915 |
+
## 8. The CROMA sensor adapter contract
|
| 916 |
+
|
| 917 |
+
`specialists/optical_sar/sensor_adapter.py` is *"the one place where 'some sensor's bands' becomes
|
| 918 |
+
'CROMA's canonical channels', and it exists so that translation is explicit, inspectable and testable."*
|
| 919 |
+
Its module docstring states **the two hard rules** from the architecture plan:
|
| 920 |
+
|
| 921 |
+
> "No invented missing bands."
|
| 922 |
+
> "Do not fabricate spectral bands."
|
| 923 |
+
|
| 924 |
+
The failure mode the module prevents is described with unusual clarity:
|
| 925 |
+
|
| 926 |
+
> A 4-band Cartosat-2S scene maps cleanly onto canonical channels 1-4. Filling channels 5-12 with
|
| 927 |
+
> *something* β a copy of B3 as a fake "red-edge", a zero-order hold, an interpolation β produces a tensor
|
| 928 |
+
> whose shape is correct and whose statistics look plausible. CROMA would run. Nothing would raise. And
|
| 929 |
+
> every downstream number would be computed from channels that no sensor ever measured, presented as if
|
| 930 |
+
> they were measurements.
|
| 931 |
+
|
| 932 |
+
So: **missing channels are zero-filled, and the availability mask says which ones were real.**
|
| 933 |
+
|
| 934 |
+
### 8.1 `SensorDescriptor`
|
| 935 |
+
|
| 936 |
+
The schema object (`core/schemas.py:138`) is `extra="forbid"`, described as the *"Frozen sensor-adapter
|
| 937 |
+
contract (plan section 18)"*:
|
| 938 |
+
|
| 939 |
+
| Field | Type | Default | Meaning |
|
| 940 |
+
|---|---|---|---|
|
| 941 |
+
| `sensor` | `str` | required | Sensor identifier, e.g. `"cartosat_2s"` |
|
| 942 |
+
| `available_bands` | `list[str]` | `[]` | The bands the sensor actually delivers |
|
| 943 |
+
| `band_map` | `dict[str, str]` | `{}` | Source band name β canonical band name |
|
| 944 |
+
| `normalization` | `str` | `"percentile"` | Identifier recorded for the trace |
|
| 945 |
+
| `availability_mask` | `list[bool]` | `[]` | Which canonical channels are real |
|
| 946 |
+
| `resolution` | `float \| None` | `None` | Ground sample distance in metres |
|
| 947 |
+
| `canonical_channels` | `int` | `0` | Target channel count (12 optical, 2 SAR) |
|
| 948 |
+
| `missing_channels_zero_filled` | `bool` | `True` | The zero-fill convention |
|
| 949 |
+
|
| 950 |
+
**Why `band_map` is canonical-channel-keyed.** The module docstring addresses a direction question: the
|
| 951 |
+
plan's example writes `band_map` sensor-keyed (`{"B1": "c1", ...}`), and the schema types it as
|
| 952 |
+
`dict[str, str]`. This module uses the schema's direction because the question asked at load time is
|
| 953 |
+
forward-looking: *"canonical channel 3 β did I get a real band for it, and if so which one?"*
|
| 954 |
+
`canonical_to_sensor` is exposed as the inverse for reporting.
|
| 955 |
+
|
| 956 |
+
### 8.2 The canonical channel orders
|
| 957 |
+
|
| 958 |
+
```python
|
| 959 |
+
OPTICAL_CANONICAL: tuple[str, ...] = (
|
| 960 |
+
"B01", # coastal aerosol
|
| 961 |
+
"B02", # blue
|
| 962 |
+
"B03", # green
|
| 963 |
+
"B04", # red
|
| 964 |
+
"B05", # red edge 1
|
| 965 |
+
"B06", # red edge 2
|
| 966 |
+
"B07", # red edge 3
|
| 967 |
+
"B08", # NIR
|
| 968 |
+
"B8A", # narrow NIR
|
| 969 |
+
"B09", # water vapour
|
| 970 |
+
"B11", # SWIR 1
|
| 971 |
+
"B12", # SWIR 2
|
| 972 |
+
)
|
| 973 |
+
|
| 974 |
+
SAR_CANONICAL: tuple[str, ...] = ("VV", "VH")
|
| 975 |
+
```
|
| 976 |
+
|
| 977 |
+
The comment above `OPTICAL_CANONICAL` states the contract: *"The ORDER is part of the contract: channel
|
| 978 |
+
index i must always mean the same physical band, or a pretrained encoder's weights would be applied to
|
| 979 |
+
the wrong input."* This is Sentinel-2 L2A with the cirrus band removed, per CROMA's README: *"Sentinel-2
|
| 980 |
+
must be 12 channels (remove the cirrus band if necessary)."* Note that `B10` (cirrus) is **not** in the
|
| 981 |
+
canonical order β that is the removal.
|
| 982 |
+
|
| 983 |
+
`SAR_CANONICAL` is described as *"Sentinel-1 dual-pol order CROMA was pretrained with. Positional fallback
|
| 984 |
+
ONLY."*
|
| 985 |
+
|
| 986 |
+
`_BAND_ALIASES` maps common spellings onto the canonical form: `B1`β`B01`, `BLUE`β`B02`, `NIR`β`B08`,
|
| 987 |
+
`SWIR1`β`B11`, `VVPOL`β`VV`, and so on. The comment is a caution: *"Kept small and explicit: a guess here
|
| 988 |
+
becomes a silently mislabelled band."* `_canonicalise` strips whitespace and underscores and uppercases
|
| 989 |
+
before lookup, and **never invents** a band β an unmapped name passes through unchanged and is later
|
| 990 |
+
rejected.
|
| 991 |
+
|
| 992 |
+
### 8.3 Building an optical adapter
|
| 993 |
+
|
| 994 |
+
`build_optical_adapter(sensor, *, available_bands, canonical_channels=12, normalization="percentile",
|
| 995 |
+
resolution=None)` (`specialists/optical_sar/sensor_adapter.py:205`):
|
| 996 |
+
|
| 997 |
+
- When `available_bands` is `None`, the sensor is looked up in `_KNOWN_OPTICAL_SENSORS`; an unknown
|
| 998 |
+
sensor with no explicit band list raises `UnsupportedBandsError` β *"refusing to guess a band
|
| 999 |
+
arrangement."*
|
| 1000 |
+
- Each band is canonicalised. Any canonical name **not** in `OPTICAL_CANONICAL` raises
|
| 1001 |
+
`UnsupportedBandsError`: *"do not map onto the canonical Sentinel-2 order ... supply an explicit
|
| 1002 |
+
band_map rather than guessing a position."*
|
| 1003 |
+
- Duplicate bands raise.
|
| 1004 |
+
- More bands than `canonical_channels` raises.
|
| 1005 |
+
- `band_map` is built as `{source: canonical for source, canonical in zip(bands, canonical)}` β keyed by
|
| 1006 |
+
the **source** name, because *"that is what an array row can actually be looked up by."*
|
| 1007 |
+
- `availability` is `[canonical_name in band_map.values() for canonical_name in OPTICAL_CANONICAL[:canonical_channels]]`.
|
| 1008 |
+
|
| 1009 |
+
The returned descriptor has `missing_channels_zero_filled=True`.
|
| 1010 |
+
|
| 1011 |
+
`_KNOWN_OPTICAL_SENSORS` is *"deliberately tiny and explicit. Adding one is a code change, not a guess at
|
| 1012 |
+
runtime"*:
|
| 1013 |
+
|
| 1014 |
+
| Sensor | Bands |
|
| 1015 |
+
|---|---|
|
| 1016 |
+
| `sentinel_2` | All 12 canonical |
|
| 1017 |
+
| `cartosat_2s` | `["B1", "B2", "B3", "B4"]` |
|
| 1018 |
+
| `cartosat_2` | `["B1", "B2", "B3", "B4"]` |
|
| 1019 |
+
| `landsat_8` | `["B1", "B2", "B3", "B4"]` |
|
| 1020 |
+
|
| 1021 |
+
The comment on `cartosat_2s` is a standing instruction: *"Cartosat-2S is a 4-band VNIR imager. It is NOT a
|
| 1022 |
+
Sentinel-2 clone and must not be treated as one: the 8 channels it lacks stay masked-off."*
|
| 1023 |
+
|
| 1024 |
+
### 8.4 Describing a SAR sensor β inspect, do not assume
|
| 1025 |
+
|
| 1026 |
+
`describe_sar(sensor, *, available_bands, canonical_channels=2, normalization="db", resolution=None)`
|
| 1027 |
+
(`specialists/optical_sar/sensor_adapter.py:318`) implements the plan's instruction:
|
| 1028 |
+
|
| 1029 |
+
> "RISAT imagery may vary in acquisition/polarization characteristics... the SAR adapter must inspect
|
| 1030 |
+
> actual available channels rather than assuming one fixed polarization pair."
|
| 1031 |
+
|
| 1032 |
+
The module docstring explains why a hardcoded `["VV", "VH"]` would be wrong: *"ISRO describes RISAT-1 as
|
| 1033 |
+
a C-band SAR mission with multiple polarization configurations (single HH, single VV, dual HH+HV, dual
|
| 1034 |
+
VV+VH, ...). A hardcoded `["VV", "VH"]` would silently mislabel an HH/HV scene, and the resulting tensor
|
| 1035 |
+
would be a *confidently wrong* input rather than a recognisably broken one."*
|
| 1036 |
+
|
| 1037 |
+
When `available_bands` names the polarisations, they are honoured. When it does not, the **channel count**
|
| 1038 |
+
is used to select the most likely convention, and **the choice is recorded on the descriptor's
|
| 1039 |
+
`available_bands`** so the assumption appears in the trace rather than being invisible.
|
| 1040 |
+
|
| 1041 |
+
The **slot order is the sensor's own**:
|
| 1042 |
+
|
| 1043 |
+
```python
|
| 1044 |
+
band_map: dict[str, str] = {source: source for source in bands}
|
| 1045 |
+
availability = [i < len(bands) for i in range(canonical_channels)]
|
| 1046 |
+
```
|
| 1047 |
+
|
| 1048 |
+
The comment explains: *"Forcing HH/HV into a VV/VH frame would mean either relabelling the polarisations
|
| 1049 |
+
(a lie the mask would then contradict) or masking both off and discarding real data."* So the tensor holds
|
| 1050 |
+
the radar measurements that exist, and `available_bands[i]` says which polarisation slot `i` is.
|
| 1051 |
+
|
| 1052 |
+
`_known_sar(sensor)` is described as *"Best-effort polarisation layout for a named SAR sensor, or `[]`"*:
|
| 1053 |
+
|
| 1054 |
+
| Sensor key | Returns |
|
| 1055 |
+
|---|---|
|
| 1056 |
+
| `sentinel_1`, `sentinel1` | `["VV", "VH"]` |
|
| 1057 |
+
| `risat*` | `["VV", "VH"]` β **recorded as an ASSUMPTION, not a fact** |
|
| 1058 |
+
| `alos_palsar`, `palsar`, `alos2` | `["HH", "HV"]` |
|
| 1059 |
+
| anything else | `[]` β *"An empty return is honest: it says 'I do not know this sensor's polarisations'"* |
|
| 1060 |
+
|
| 1061 |
+
The `risat` case carries an explicit comment: *"RISAT-1's most common operational mode is dual-pol.
|
| 1062 |
+
Recorded as an ASSUMPTION, not a fact β see the module docstring. Callers with real band metadata should
|
| 1063 |
+
pass `available_bands` and override this."* This is an `ASSUMPTION` in the project's status vocabulary,
|
| 1064 |
+
and it is the reason the hidden-set behaviour on RISAT cannot be claimed as measured.
|
| 1065 |
+
|
| 1066 |
+
### 8.5 Applying the canonical layout β `_apply_canonical`
|
| 1067 |
+
|
| 1068 |
+
This is *"the function the two hard rules live in"* (`specialists/optical_sar/sensor_adapter.py:428`). It
|
| 1069 |
+
places real bands in canonical positions and zero-fills the rest:
|
| 1070 |
+
|
| 1071 |
+
```python
|
| 1072 |
+
canonical = np.zeros((n_canonical, height, width), dtype=np.float32)
|
| 1073 |
+
mask = np.zeros((n_canonical,), dtype=bool)
|
| 1074 |
+
|
| 1075 |
+
for source_index, source_name in enumerate(descriptor.available_bands):
|
| 1076 |
+
if source_index >= n_source:
|
| 1077 |
+
continue
|
| 1078 |
+
canonical_name = descriptor.band_map.get(source_name)
|
| 1079 |
+
if canonical_name is None or canonical_name not in order:
|
| 1080 |
+
continue
|
| 1081 |
+
channel_index = order.index(canonical_name)
|
| 1082 |
+
if channel_index >= n_canonical:
|
| 1083 |
+
continue
|
| 1084 |
+
canonical[channel_index] = array[source_index].astype(np.float32)
|
| 1085 |
+
mask[channel_index] = True
|
| 1086 |
+
```
|
| 1087 |
+
|
| 1088 |
+
The docstring's key claim: *"There is no branch anywhere below that writes a value into an unavailable
|
| 1089 |
+
channel."* The absence of such a branch **is** the mask. If no band could be placed while the array
|
| 1090 |
+
carries bands, it raises `UnsupportedBandsError` rather than returning an all-zero tensor silently.
|
| 1091 |
+
|
| 1092 |
+
`_canonical_order(descriptor)` decides the slot order: for a descriptor whose `available_bands` are all
|
| 1093 |
+
SAR polarisations, the order is the descriptor's own list (padded to canonical width); otherwise it is
|
| 1094 |
+
`OPTICAL_CANONICAL`.
|
| 1095 |
+
|
| 1096 |
+
`adapt_optical` and `adapt_sar` are thin wrappers that call `_apply_canonical` with the right modality
|
| 1097 |
+
label. `adapt_sar`'s docstring names the plan's internal representation exactly:
|
| 1098 |
+
`canonical_sar[2, H, W]` + `sar_channel_mask[2]`.
|
| 1099 |
+
|
| 1100 |
+
### 8.6 `SensorAdapterOutput`
|
| 1101 |
+
|
| 1102 |
+
The dataclass returned by `adapt_*` carries `canonical` `(C, H, W)` float32, `mask` `(C,)` bool, and the
|
| 1103 |
+
`descriptor`. It exposes `n_channels`, `missing_indices` (`[i for i, ok in enumerate(self.mask) if not
|
| 1104 |
+
ok]`), `available_bands`, and a `to_dict()` that includes `n_available` and `missing_channels`.
|
| 1105 |
+
|
| 1106 |
+
### 8.7 `positional_fallback_descriptor`
|
| 1107 |
+
|
| 1108 |
+
Used when a raster has the right **channel count** for a modality but **no band labels**. The fallback
|
| 1109 |
+
order is chosen from `_FALLBACK_SAR_ORDER` (`("VV", "VH")`, `("HH", "HV")`) for 1β2 channels, or the
|
| 1110 |
+
canonical optical prefix for β₯3 channels. The fact that it was a fallback is **recorded in the `sensor`
|
| 1111 |
+
field** so it *"cannot be mistaken for measured metadata."* The `complement=True` flag selects the other
|
| 1112 |
+
convention, which lets a caller test both HH/HV and VV/VH orderings without asserting which is correct.
|
| 1113 |
+
|
| 1114 |
+
The optical-SAR specialist uses this in `_descriptor_for`
|
| 1115 |
+
(`specialists/optical_sar/specialist.py:398`) and appends a warning that names the assumption: *"The
|
| 1116 |
+
polarisation IDENTITY of each channel is NOT known from a band count alone β supply a sensor descriptor
|
| 1117 |
+
to state it."*
|
| 1118 |
+
|
| 1119 |
+
### 8.8 Finding C-1 β the mask goes to the fusion head, not to CROMA
|
| 1120 |
+
|
| 1121 |
+
This is the single most important fact in this section. `core/schemas.py`'s module docstring lists it
|
| 1122 |
+
among the findings baked into the schemas: *"C-1 channel-availability mask is a first-class fusion input,
|
| 1123 |
+
never a CROMA input."*
|
| 1124 |
+
|
| 1125 |
+
`specialists/optical_sar/fusion_head.py` devotes a section to the rationale:
|
| 1126 |
+
|
| 1127 |
+
> CROMA is a masked autoencoder. Handing it an availability mask invites it to reconstruct the missing
|
| 1128 |
+
> channels β which is precisely the fabrication the sensor adapter exists to prevent: the model would
|
| 1129 |
+
> output plausible values for bands no sensor measured, and those values would then be treated as data.
|
| 1130 |
+
> The mask is therefore consumed HERE, where it can only do one thing: tell the classifier which inputs
|
| 1131 |
+
> to distrust.
|
| 1132 |
+
|
| 1133 |
+
Mechanically: `CROMAEncoder.encode` has a pinned signature of exactly `(self, optical, sar)` β the
|
| 1134 |
+
radiometry module records that `tests/unit/test_optical_sar_croma.py` *"asserts the absence of any mask
|
| 1135 |
+
parameter, because freeze finding C-1 says CROMA never receives one."* The mask instead enters at
|
| 1136 |
+
`assemble_fusion_input` (Β§8.9). The optical-SAR specialist's `JOINT_FEATURE_REGION` evidence item records
|
| 1137 |
+
the fact explicitly in its payload: `"mask_consumed_by": "fusion_head"` and `"croma_received_mask":
|
| 1138 |
+
False` (`specialists/optical_sar/specialist.py:816`).
|
| 1139 |
+
|
| 1140 |
+
`EvidenceType` even has a dedicated member for the mask as trust evidence:
|
| 1141 |
+
`AVAILABILITY_MASK = "availability_mask" # C-1: modality trust evidence` (`core/schemas.py:76`). The
|
| 1142 |
+
optical-SAR specialist emits one `AVAILABILITY_MASK` item per modality, carrying the sensor, canonical
|
| 1143 |
+
channel count, available bands, band map, the mask itself, the missing indices,
|
| 1144 |
+
`missing_channels_zero_filled`, `normalization` and `resolution`.
|
| 1145 |
+
|
| 1146 |
+
### 8.9 The frozen fusion concatenation
|
| 1147 |
+
|
| 1148 |
+
`specialists/optical_sar/fusion_head.py` defines the concatenation that the mask feeds into:
|
| 1149 |
+
|
| 1150 |
+
```
|
| 1151 |
+
optical_GAP (B, 768)
|
| 1152 |
+
SAR_GAP (B, 768)
|
| 1153 |
+
joint_GAP (B, 768)
|
| 1154 |
+
optical_mask (B, 12) <- availability, from the sensor adapter
|
| 1155 |
+
sar_mask (B, 2) <- availability, from the sensor adapter
|
| 1156 |
+
---------
|
| 1157 |
+
concat (B, 2318)
|
| 1158 |
+
```
|
| 1159 |
+
|
| 1160 |
+
`expected_fusion_dim(encoder_dim=768, optical_channels=12, sar_channels=2, modalities_used=3)` returns
|
| 1161 |
+
`3*768 + 12 + 2 = 2318`. The module recomputes the width and **refuses to build on a mismatch**, so *"a
|
| 1162 |
+
config edit cannot silently reshape the first Linear layer into something that trains but means
|
| 1163 |
+
nothing."* `core/config.py` recomputes it independently at load time from `croma.encoder_dim`,
|
| 1164 |
+
`croma.modalities_used`, `croma.optical_channels` and `croma.sar_channels`, and rejects a config whose
|
| 1165 |
+
`fusion.input_dim` disagrees (finding C-1 in the validator's own message).
|
| 1166 |
+
|
| 1167 |
+
`assemble_fusion_input` asserts the order rather than assuming it β *"a permutation here produces a tensor
|
| 1168 |
+
of exactly the right shape that trains to a worse number β the hardest kind of bug to notice"* β and
|
| 1169 |
+
validates that the three GAP vectors agree on batch size and that both masks agree with it.
|
| 1170 |
+
|
| 1171 |
+
`channel_dropout(features, mask, *, keep_probabilities, rng)` is the mechanism that *"teaches the head to
|
| 1172 |
+
trust the availability mask"*. Its load-bearing line is:
|
| 1173 |
+
|
| 1174 |
+
```python
|
| 1175 |
+
keep = draws & (msk > 0.0)
|
| 1176 |
+
out_mask = keep.astype(np.float32)
|
| 1177 |
+
out = out * out_mask
|
| 1178 |
+
```
|
| 1179 |
+
|
| 1180 |
+
Dropped channels are **zeroed, not renormalised** β *"renormalising would fabricate a scale that the real
|
| 1181 |
+
missing-channel case does not have"* β and the mask is updated in lockstep. `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md`
|
| 1182 |
+
Β§5 independently verified this against freeze Β§2.5 and ruled: *"the engineer's implementation SATISFIES
|
| 1183 |
+
Β§2.5's intent. No change entry required."* Its table records that the `& (msk > 0.0)` guard is the
|
| 1184 |
+
load-bearing one: *"no training transform can resurrect a band the sensor did not measure."*
|
| 1185 |
+
|
| 1186 |
+
The schedule (plan section 20) is optical at `100/80/60/40%` and SAR at `100/50%`, declared in
|
| 1187 |
+
`training/fusion/train.py` as `OPTICAL_DROPOUT_RATES = (1.0, 0.8, 0.6, 0.4)` and
|
| 1188 |
+
`SAR_DROPOUT_RATES = (1.0, 0.5)`.
|
| 1189 |
+
|
| 1190 |
+
`mask_availability_stats(optical_mask, sar_mask)` returns `optical_available_fraction`,
|
| 1191 |
+
`sar_available_fraction`, `optical_channels_present` and `sar_channels_present`. These feed the
|
| 1192 |
+
confidence system: the two modality-confidence terms in the optical-SAR specialist are exactly the
|
| 1193 |
+
availability fractions, which the specialist's docstring justifies: *"On the hidden set the optical side
|
| 1194 |
+
is Cartosat-2S with 4 bands against a canonical 12, so 8 of 12 channels are always absent. A classifier
|
| 1195 |
+
resting on 4 measured channels is not the same claim as one resting on 12."*
|
| 1196 |
+
|
| 1197 |
+
### 8.10 The optical-SAR confidence components
|
| 1198 |
+
|
| 1199 |
+
`OpticalSarSpecialist._confidence_components` (`specialists/optical_sar/specialist.py:481`) implements
|
| 1200 |
+
plan section 26's four components with these weights:
|
| 1201 |
+
|
| 1202 |
+
```python
|
| 1203 |
+
CONFIDENCE_WEIGHTS: dict[str, float] = {
|
| 1204 |
+
"fusion_margin": 0.40,
|
| 1205 |
+
"optical_confidence": 0.20,
|
| 1206 |
+
"sar_confidence": 0.20,
|
| 1207 |
+
"cross_modal_agreement": 0.20,
|
| 1208 |
+
}
|
| 1209 |
+
```
|
| 1210 |
+
|
| 1211 |
+
`cross_modal_agreement` is `cosine(optical_GAP, SAR_GAP)`, rescaled from `[-1, 1]` to `[0, 1]` via
|
| 1212 |
+
`AGREEMENT_FLOOR = -1.0`. `cross_modal_agreement` in `inference.py` returns `None` on any exception or a
|
| 1213 |
+
shape mismatch, and the specialist treats `None` as a zero component. The composition is a weighted sum,
|
| 1214 |
+
**floored to 0.0** when there is no prediction or no trained head β *"Both are signal GAPS, not weak
|
| 1215 |
+
signals, so they resolve to 0.0 rather than merely discounting."* The components dict records
|
| 1216 |
+
`has_croma`, `trained_head`, `no_prediction` and the raw availability counts, so *"the trace states every
|
| 1217 |
+
reason the score is zero, not just the first one encountered."*
|
| 1218 |
+
|
| 1219 |
+
---
|
| 1220 |
+
|
| 1221 |
+
## 9. Pair-compatibility rules
|
| 1222 |
+
|
| 1223 |
+
Two workflows take a **pair** of assets, and both enforce compatibility before reading pixels.
|
| 1224 |
+
|
| 1225 |
+
### 9.1 Change β exactly two, temporally distinct, co-registered
|
| 1226 |
+
|
| 1227 |
+
`ChangeSpecialist.validate_request` (`specialists/change/specialist.py:190`) enforces **exactly two**
|
| 1228 |
+
assets. The comment names both common mistakes: *"ONE asset is the common mistake: the caller treated
|
| 1229 |
+
this as a single-image task. THREE is the other: they attached a time series. Both are rejected with the
|
| 1230 |
+
same typed error."* The user-facing message is *"Change detection needs exactly two images: an earlier one
|
| 1231 |
+
and a later one."*
|
| 1232 |
+
|
| 1233 |
+
`_assess_pair` then runs three checks in order:
|
| 1234 |
+
|
| 1235 |
+
1. **Temporal distinctness.** Identical `path` raises `TemporalPairError`. Identical `sha256` (when both
|
| 1236 |
+
carry one) also raises: *"a repeated acquisition is not a temporal pair."*
|
| 1237 |
+
2. **CRS compatibility** via `compare_crs` (Β§2.5). Incompatible or reprojection-requiring pairs get a
|
| 1238 |
+
warning, not a refusal β pixel-domain change detection can still proceed, but geographic claims are
|
| 1239 |
+
withheld.
|
| 1240 |
+
3. **Registration quality** via `measure_registration` (Β§9.3).
|
| 1241 |
+
|
| 1242 |
+
If registration is **not usable**, the specialist **suppresses spatial claims** while still computing and
|
| 1243 |
+
returning the change maps:
|
| 1244 |
+
|
| 1245 |
+
> The maps are still returned (the caller may want to look at them), but the REGION CLAIMS are withheld:
|
| 1246 |
+
> a region is an assertion about *where* something changed on the ground, and we cannot make it.
|
| 1247 |
+
|
| 1248 |
+
`ChangeSpecialist._geospatial` drops the transform in that case, returning a `GeoMetadata` with
|
| 1249 |
+
`is_georeferenced=False` β *"a transform invites the caller to place a region on a map, and we have just
|
| 1250 |
+
said the regions are not placeable."* `_write_artifacts` writes the map **without** georeferencing when
|
| 1251 |
+
suppressed, and the comment explains why copying T1's transform would be wrong: *"Copying T1's transform
|
| 1252 |
+
onto it would produce a file that unlocks exactly the placed-on-a-map reading the suppression exists to
|
| 1253 |
+
forbid."* The confidence is driven to the floor via a **minimum** against the registration factor, so
|
| 1254 |
+
alignment is a gate, not one vote among three.
|
| 1255 |
+
|
| 1256 |
+
The change confidence weights are:
|
| 1257 |
+
|
| 1258 |
+
```python
|
| 1259 |
+
CONFIDENCE_WEIGHTS: dict[str, float] = {
|
| 1260 |
+
"registration_quality": 0.50,
|
| 1261 |
+
"mean_change_probability": 0.30,
|
| 1262 |
+
"component_stability": 0.20,
|
| 1263 |
+
}
|
| 1264 |
+
```
|
| 1265 |
+
|
| 1266 |
+
with `STABILITY_HEADROOM = 0.5`, `DEFAULT_MAX_SHIFT_PX = 8`, `MIN_REGION_PIXELS = 32`.
|
| 1267 |
+
|
| 1268 |
+
### 9.2 Optical-SAR β exactly two, one of each modality
|
| 1269 |
+
|
| 1270 |
+
`OpticalSarSpecialist.validate_request` (`specialists/optical_sar/specialist.py:192`) enforces **exactly
|
| 1271 |
+
two** assets **and** that they are one optical and one SAR. The docstring explains the case worth
|
| 1272 |
+
explaining:
|
| 1273 |
+
|
| 1274 |
+
> Two optical images is not a malformed request in the way that three is β it is a plausible mistake,
|
| 1275 |
+
> from an operator who uploaded a bi-temporal pair to a single-modality workflow, or who labelled a
|
| 1276 |
+
> four-band Cartosat scene as SAR.
|
| 1277 |
+
>
|
| 1278 |
+
> If such a request were accepted, the adapter would place the second optical image into the 2-channel SAR
|
| 1279 |
+
> slot, zero-fill, and produce a fused tensor of the correct shape. CROMA would run. The answer would be
|
| 1280 |
+
> about one modality while claiming to fuse two.
|
| 1281 |
+
|
| 1282 |
+
So the modality check is done on the **assets, before any pixels are read**. `_assess_pair` raises
|
| 1283 |
+
`InvalidRequestError` with a machine-readable `reason` of `"two_optical"`, `"two_sar"` or
|
| 1284 |
+
`"indeterminate"`, and each carries a user-facing message.
|
| 1285 |
+
|
| 1286 |
+
### 9.3 The equal-dimensions requirement, and the 726Β²-vs-736Β² error
|
| 1287 |
+
|
| 1288 |
+
The change detector requires T1 and T2 to have **identical shapes**. The guard is in the model itself,
|
| 1289 |
+
`specialists/change/stanet.py:415`:
|
| 1290 |
+
|
| 1291 |
+
```python
|
| 1292 |
+
if t1.shape != t2.shape:
|
| 1293 |
+
raise SpecialistError(
|
| 1294 |
+
f"T1 and T2 must have the same shape; got {tuple(t1.shape)} "
|
| 1295 |
+
f"and {tuple(t2.shape)}",
|
| 1296 |
+
specialist="change",
|
| 1297 |
+
)
|
| 1298 |
+
```
|
| 1299 |
+
|
| 1300 |
+
This is a **hard** requirement, and it is where the bundled demo pair fails. `docs/FINAL_DELIVERY_REPORT.md`
|
| 1301 |
+
Β§6 records it as `DEGRADED`:
|
| 1302 |
+
|
| 1303 |
+
| Finding | Status | Detail |
|
| 1304 |
+
|---|---|---|
|
| 1305 |
+
| Change on the bundled EO pair | DEGRADED | `delta-growth-t0-1975.jpg` (726Β²) and `t1-2025.jpg` (736Β²) differ in shape; the change specialist errors (`T1 and T2 must have the same shape`). A same-shape pair, or a resize step, is needed for a clean change demo. |
|
| 1306 |
+
|
| 1307 |
+
There is **no resize step on the change path**. `load_image_array` explicitly does not resample
|
| 1308 |
+
(Β§7.5) β its docstring states *"a silent resize here would put every box in the wrong place"* β and the
|
| 1309 |
+
change specialist's `_predict` only pads for the encoder's stride-8 requirement, cropping back to the
|
| 1310 |
+
original size:
|
| 1311 |
+
|
| 1312 |
+
```python
|
| 1313 |
+
h, w = t1.shape[-2:]
|
| 1314 |
+
if h % 8 or w % 8:
|
| 1315 |
+
ph, pw = (-h) % 8, (-w) % 8
|
| 1316 |
+
t1 = torch.nn.functional.pad(t1, (0, pw, 0, ph), mode="reflect")
|
| 1317 |
+
t2 = torch.nn.functional.pad(t2, (0, pw, 0, ph), mode="reflect")
|
| 1318 |
+
```
|
| 1319 |
+
|
| 1320 |
+
That padding does **not** reconcile two different input sizes: T1 and T2 are each padded by their own
|
| 1321 |
+
`(-dim) % 8`, so a 726Β² and a 736Β² input become 728Β² and 736Β² respectively β still different β and the
|
| 1322 |
+
model raises. The reflect-pad is a *stride* accommodation, not a *pair* accommodation, and the two must
|
| 1323 |
+
not be conflated.
|
| 1324 |
+
|
| 1325 |
+
**Consequence for operators.** The change workflow requires a same-shape pair, and the failure is
|
| 1326 |
+
`IMPLEMENTED` (it raises a typed error with a readable message) rather than silent. A resize or
|
| 1327 |
+
co-registration step that reconciles differing dimensions is **not implemented** (Β§13). The
|
| 1328 |
+
`FINAL_DELIVERY_REPORT` records the demo as `DEGRADED` and prescribes *"a same-shape pair, or a resize
|
| 1329 |
+
step"* β either would fix the demo; neither is currently in the code path.
|
| 1330 |
+
|
| 1331 |
+
### 9.4 Registration measurement
|
| 1332 |
+
|
| 1333 |
+
`measure_registration` (`specialists/change/postprocess.py:134`) uses `cv2.phaseCorrelate` on grayscale
|
| 1334 |
+
collapses of the two images, returning a `RegistrationQuality` with `shift_x`, `shift_y`, `response`,
|
| 1335 |
+
`max_shift_px`, `is_usable` and `reason`. Defaults: `max_shift_px=8`, `min_response=0.15`.
|
| 1336 |
+
|
| 1337 |
+
The module docstring is candid about the method's limit:
|
| 1338 |
+
|
| 1339 |
+
> Deliberately a TRANSLATION model. Real mis-registration includes rotation and warp, which phase
|
| 1340 |
+
> correlation does not recover; a small residual after a translation correction is therefore not proof of
|
| 1341 |
+
> good alignment. What the measurement can do honestly is flag the *bad* cases, and that is how it is
|
| 1342 |
+
> used: as a gate, not a certificate.
|
| 1343 |
+
|
| 1344 |
+
`RegistrationQuality.confidence_factor()` applies two independent penalties and takes the worse:
|
| 1345 |
+
a shift penalty (`1.0` at zero shift, `0.0` at `max_shift_px`) and a response factor (`0.0` at zero
|
| 1346 |
+
response, `1.0` at response β₯ 0.15). A pair that is offset *and* uncorrelated takes the worse of the two β
|
| 1347 |
+
*"the conservative choice."*
|
| 1348 |
+
|
| 1349 |
+
A failure to *measure* is treated as a measurement of failure β both mean the pair cannot be trusted
|
| 1350 |
+
spatially β but the distinction is recorded in the warning: *"A failure to MEASURE is not the same as a
|
| 1351 |
+
measurement of failure."*
|
| 1352 |
+
|
| 1353 |
+
Two registration-gate hypotheses were tested and both are recorded as **REJECTED** or
|
| 1354 |
+
**INSUFFICIENT**: `docs/CHANGE_REGISTRATION_GATE_COHERENCE_TEST.md` records a rejected hypothesis (AUC S2
|
| 1355 |
+
vs S3 β€ 0.59), and `docs/CHANGE_REGISTRATION_GATE_TEXTURE_TEST.md` records a low-texture signal that was
|
| 1356 |
+
**CONFIRMED but INSUFFICIENT**, with option 3 **ELIMINATED**. So the gate as implemented (phase
|
| 1357 |
+
correlation) is the surviving method, and no additional coherence- or texture-based gate was adopted.
|
| 1358 |
+
|
| 1359 |
+
---
|
| 1360 |
+
|
| 1361 |
+
## 10. Large-image tiling policy
|
| 1362 |
+
|
| 1363 |
+
### 10.1 What is declared
|
| 1364 |
+
|
| 1365 |
+
`configs/base.yaml` declares the full tiling policy:
|
| 1366 |
+
|
| 1367 |
+
```yaml
|
| 1368 |
+
image:
|
| 1369 |
+
max_pixels: 25000000
|
| 1370 |
+
tile_size: 512
|
| 1371 |
+
tile_overlap: 128
|
| 1372 |
+
max_tiles: 64
|
| 1373 |
+
# plan section 9.1 tile policy: whole-image thumbnail first, then top-K tiles.
|
| 1374 |
+
# max_tiles is the hard ceiling on tiles *examined*; top_k_tiles is how many
|
| 1375 |
+
# are actually sent through a specialist.
|
| 1376 |
+
top_k_tiles: 4
|
| 1377 |
+
```
|
| 1378 |
+
|
| 1379 |
+
| Key | Value | Declared meaning |
|
| 1380 |
+
|---|---|---|
|
| 1381 |
+
| `max_pixels` | 25,000,000 | Pixel budget; exceeding it is recoverable (the caller may downscale) |
|
| 1382 |
+
| `tile_size` | 512 | Tile edge in pixels |
|
| 1383 |
+
| `tile_overlap` | 128 | Overlap between adjacent tiles |
|
| 1384 |
+
| `max_tiles` | 64 | Hard ceiling on tiles **examined** |
|
| 1385 |
+
| `top_k_tiles` | 4 | How many tiles are actually sent through a specialist |
|
| 1386 |
+
|
| 1387 |
+
The comment states the intended policy: *"whole-image thumbnail first, then top-K tiles."* The change
|
| 1388 |
+
workflow declares a different tiling resolution, `change.tile_size: 256` with `change.tile_overlap: 0`
|
| 1389 |
+
β 256 is *"the change model's OWN training resolution"* (LEVIR-CD patches are 256Γ256).
|
| 1390 |
+
|
| 1391 |
+
### 10.2 What is enforced
|
| 1392 |
+
|
| 1393 |
+
`core/config.py` validates the tiling policy at load time, but only one relationship:
|
| 1394 |
+
|
| 1395 |
+
```python
|
| 1396 |
+
# --- tiling policy -------------------------------------------------
|
| 1397 |
+
if self.get("image.top_k_tiles", 0) > self.get("image.max_tiles", 0):
|
| 1398 |
+
errors.append("image.top_k_tiles cannot exceed image.max_tiles")
|
| 1399 |
+
```
|
| 1400 |
+
|
| 1401 |
+
So the config loader guarantees that the number of tiles sent through a specialist does not exceed the
|
| 1402 |
+
number examined. It does not implement tiling.
|
| 1403 |
+
|
| 1404 |
+
The pixel budget **is** enforced, but at the raster-inspection layer rather than as a tiling step:
|
| 1405 |
+
`inspect_raster` raises `OversizedImageError` when `width * height > max_pixels`, and the error is
|
| 1406 |
+
`recoverable` β the docstring says *"the caller may downscale."* The controller passes the configured
|
| 1407 |
+
budget into `_inspect_asset` during VALIDATE.
|
| 1408 |
+
|
| 1409 |
+
### 10.3 What is not implemented β the honest status
|
| 1410 |
+
|
| 1411 |
+
**The tiling policy is `DECLARED (not read)`.** `image.tile_size`, `image.tile_overlap`, `image.max_tiles`
|
| 1412 |
+
and `image.top_k_tiles` appear only in `configs/base.yaml` and in the `core/config.py` validation above.
|
| 1413 |
+
No code on the serving path slices a raster into tiles, selects the top-K, or stitches results.
|
| 1414 |
+
`architecture/03-request-lifecycle.md` reaches the same conclusion: the tiling policy is referenced only
|
| 1415 |
+
in config validation and not implemented on the serving path.
|
| 1416 |
+
|
| 1417 |
+
The practical consequence: a raster within the `max_pixels` budget is handed to a specialist **whole**.
|
| 1418 |
+
The VLM path pins `processor_longest_edge: 512` so the processor does not upscale and split a tile, but
|
| 1419 |
+
that pin assumes a 512-px input β it is a *control against* the processor's default 2048, not a tiling
|
| 1420 |
+
step. `core/config.py` enforces the relationship (`processor_longest_edge` must not exceed
|
| 1421 |
+
`image.tile_size`), with the measured rationale recorded: the default 2048 upscales a 512 tile 4Γ and then
|
| 1422 |
+
splits it into ~17 sub-images, *"the real figure is ~17x"* against the plan's estimated 4Γ.
|
| 1423 |
+
|
| 1424 |
+
`UNKNOWN β not established from the available evidence` how a raster between the tile size and the pixel
|
| 1425 |
+
budget is handled by the VLM path in practice; the config implies the tile policy would govern it, but
|
| 1426 |
+
the policy is not implemented, so the behaviour is whatever the specialist does with the whole image.
|
| 1427 |
+
|
| 1428 |
+
---
|
| 1429 |
+
|
| 1430 |
+
## 11. Grounding resolution β the frozen 224 decision
|
| 1431 |
+
|
| 1432 |
+
Grounding runs at **224 px**, and this is one of the project's cleanest pre-registered results.
|
| 1433 |
+
`configs/base.yaml`:
|
| 1434 |
+
|
| 1435 |
+
```yaml
|
| 1436 |
+
grounding:
|
| 1437 |
+
...
|
| 1438 |
+
# RESOLVED 2026-09-16. 448 did NOT earn its cost over all 16,159 VRSBench
|
| 1439 |
+
# eval records: mean best IoU -0.0147, Recall@0.5 -0.0022, and every recall
|
| 1440 |
+
# threshold lower (-0.0699 @0.10, -0.0243 @0.25), at 1.59x the latency.
|
| 1441 |
+
# Paired over identical samples: mean diff -0.0147, 95% CI [-0.0160,-0.0134],
|
| 1442 |
+
# t = -22.63. 448 was better on 8.5% of records, worse on 20.9%.
|
| 1443 |
+
# Pre-registered rule and the paired test AGREE on 224.
|
| 1444 |
+
# Evidence: docs/PHASE7_RESOLUTION_DECISION.md
|
| 1445 |
+
image_size: 224
|
| 1446 |
+
resolution_frozen: true
|
| 1447 |
+
nms_iou: 0.50
|
| 1448 |
+
max_candidates: 20
|
| 1449 |
+
confidence_threshold: 0.40
|
| 1450 |
+
benchmark_box_scale: 100.0
|
| 1451 |
+
coordinate_system: normalized_0_1
|
| 1452 |
+
encoder_projected_dim: 512
|
| 1453 |
+
```
|
| 1454 |
+
|
| 1455 |
+
### 11.1 The pre-registered rule
|
| 1456 |
+
|
| 1457 |
+
`docs/PHASE7_RESOLUTION_DECISION.md` records the rule, **fixed before the result was seen**:
|
| 1458 |
+
|
| 1459 |
+
```
|
| 1460 |
+
448 WINS if Recall@0.5 improves by >= 0.05 absolute
|
| 1461 |
+
OR mean best IoU improves by >= 0.05 absolute
|
| 1462 |
+
224 WINS otherwise
|
| 1463 |
+
INCONCLUSIVE if fewer than 30 samples were scored
|
| 1464 |
+
```
|
| 1465 |
+
|
| 1466 |
+
The artifact records `rule_changed_since_preregistration: false`.
|
| 1467 |
+
|
| 1468 |
+
### 11.2 The result
|
| 1469 |
+
|
| 1470 |
+
```
|
| 1471 |
+
224 WINS
|
| 1472 |
+
recall@0.5 gain 448/224 : -0.0022
|
| 1473 |
+
bestIoU gain 448/224 : -0.0147
|
| 1474 |
+
latency ratio : 1.59x
|
| 1475 |
+
```
|
| 1476 |
+
|
| 1477 |
+
Neither component came close to the +0.05 margin; both were **negative**. The measured detail:
|
| 1478 |
+
|
| 1479 |
+
| Metric | 224 | 448 | Delta |
|
| 1480 |
+
|---|---|---|---|
|
| 1481 |
+
| Token grid | 7 Γ 7 = 49 | 14 Γ 14 = 196 | 4.0Γ tokens |
|
| 1482 |
+
| Mean best IoU | **0.0972** | 0.0825 | **β0.0147** |
|
| 1483 |
+
| Recall@0.10 | **0.3298** | 0.2599 | **β0.0699** |
|
| 1484 |
+
| Recall@0.25 | **0.1187** | 0.0944 | **β0.0243** |
|
| 1485 |
+
| Recall@0.50 | **0.0234** | 0.0212 | **β0.0022** |
|
| 1486 |
+
| Latency mean | **20.0 ms** | 31.8 ms | 1.59Γ |
|
| 1487 |
+
| Latency p90 | **20.9 ms** | 32.9 ms | 1.57Γ |
|
| 1488 |
+
| Peak VRAM | **592.1 MB** | 599.8 MB | +7.7 MB |
|
| 1489 |
+
|
| 1490 |
+
All 16,159 / 16,159 VRSBench eval records were scored at both resolutions on a Tesla T4.
|
| 1491 |
+
|
| 1492 |
+
### 11.3 The paired test
|
| 1493 |
+
|
| 1494 |
+
Because both resolutions scored the **same 16,159 samples**, the paired test is the stronger statistic:
|
| 1495 |
+
|
| 1496 |
+
```
|
| 1497 |
+
paired samples : 16159
|
| 1498 |
+
mean 224 : 0.0972
|
| 1499 |
+
mean 448 : 0.0825
|
| 1500 |
+
mean paired diff : -0.0147 (95% CI -0.0160 .. -0.0134)
|
| 1501 |
+
t statistic : -22.63
|
| 1502 |
+
CI excludes zero : True
|
| 1503 |
+
|
| 1504 |
+
448 better on : 1371/16159 ( 8.5%)
|
| 1505 |
+
448 worse on : 3372/16159 (20.9%)
|
| 1506 |
+
identical : 11416/16159 (70.6%)
|
| 1507 |
+
```
|
| 1508 |
+
|
| 1509 |
+
The paired test and the pre-registered rule **agree**, so there is *"no rule-versus-evidence disagreement
|
| 1510 |
+
to escalate."* The win/loss split is informative on its own: 448 wins on only 8.5% and loses on 20.9% β
|
| 1511 |
+
*"the finer grid is not merely neutral, it is actively harmful on a fifth of the corpus."*
|
| 1512 |
+
|
| 1513 |
+
The recall ladder narrows as the threshold rises (β0.0699 at 0.10, β0.0243 at 0.25, β0.0022 at 0.50),
|
| 1514 |
+
which the document reads as *"the signature of a method that cannot reach high IoU either way."*
|
| 1515 |
+
|
| 1516 |
+
### 11.4 Why 448 did not help
|
| 1517 |
+
|
| 1518 |
+
The document's honest reading: the zero-shot method selects a patch by text similarity and returns that
|
| 1519 |
+
patch's box. At 224 a box is 1/7 of the image; at 448 it is 1/14. A finer grid is only better if the
|
| 1520 |
+
target is small **and** the similarity peak lands on the correct fine cell. Two things work against that:
|
| 1521 |
+
the peak is not sharper at 448 (the similarity field on frozen features is smooth, so the argmax moves
|
| 1522 |
+
around), and Recall@0.10 β the loosest threshold β degrades *most*, which means *"the fine grid is adding
|
| 1523 |
+
positional noise rather than positional precision."*
|
| 1524 |
+
|
| 1525 |
+
The document is careful to attribute this to the zero-shot baseline, not to RemoteCLIP: *"A **learned**
|
| 1526 |
+
head trained to regress boxes from these features may respond differently to resolution."*
|
| 1527 |
+
|
| 1528 |
+
### 11.5 What the decision does not establish
|
| 1529 |
+
|
| 1530 |
+
- **Whether the zero-shot baseline is good.** It is not: mean best IoU 0.0972 and Recall@0.5 0.0234 are
|
| 1531 |
+
*"weak"*. It is an ablation floor for the Phase 8 head.
|
| 1532 |
+
- **Whether a trained head has the same resolution sensitivity.** Re-opening the question is legitimate
|
| 1533 |
+
**if** the head's validation curve suggests it, and would be *"a new pre-registered experiment, not a
|
| 1534 |
+
silent retune."*
|
| 1535 |
+
- **Anything about hidden ISRO/SAC imagery.** *"VRSBench is overhead optical. The hidden set is
|
| 1536 |
+
Cartosat-2S + RISAT, a different distribution entirely."*
|
| 1537 |
+
|
| 1538 |
+
### 11.6 Downstream consequences of the frozen resolution
|
| 1539 |
+
|
| 1540 |
+
The frozen resolution propagates into concrete serving parameters:
|
| 1541 |
+
|
| 1542 |
+
- **The decode grid.** `build_grounding_specialist` sets `grid = resolution // 32`, so at 224 the grid is
|
| 1543 |
+
**7Γ7 = 49 cells**. The head emits `(1, 49, 5)` β `[tx, ty, tw, th, objectness]` per cell.
|
| 1544 |
+
- **The projected dimension.** `grounding.encoder_projected_dim: 512` is the measured `visual.proj`
|
| 1545 |
+
output of RemoteCLIP ViT-B/32 (the transformer width is 768; `visual.proj` maps to 512). The
|
| 1546 |
+
per-cell feature is `concat([patch, text, patch*text, global_pool]) = 4 Γ 512 = 2048`, which
|
| 1547 |
+
`grounding_head.feature_dim: 2048` must equal β enforced at load time by `core/config.py`, which
|
| 1548 |
+
rejects any other value, because *"a mismatch here is a SILENT shape error"* that torch raises only at
|
| 1549 |
+
the similarity step, *"after the patch features are already cached."*
|
| 1550 |
+
- **NMS.** `grounding.nms_iou: 0.50` is passed to `nms(boxes, scores, iou_threshold)`
|
| 1551 |
+
(`specialists/grounding/head.py:294`).
|
| 1552 |
+
- **The serving candidate budget.** `GroundingSpecialist` defaults `max_candidates=6`, **not** the
|
| 1553 |
+
config's `20`. The class docstring records why: *"The Phase 8 decision measured that a 20-box budget
|
| 1554 |
+
inflates mean best IoU relative to the 5.99-box baseline; 6 is the decode-matched setting."* The builder
|
| 1555 |
+
reads `grounding.serving_max_candidates` with a default of 6.
|
| 1556 |
+
- **The degenerate-box guard.** `DEGENERATE_AREA_FRACTION = 0.9` drops any zero-shot candidate covering
|
| 1557 |
+
more than 90% of the frame, *"the shape a failed localisation takes once it has been smoothed into a
|
| 1558 |
+
rectangle."* This guard applies to the **zero-shot fallback path only** β *"the trained-head decode is
|
| 1559 |
+
a regressed box with its own learned prior; filtering it would change the frozen Phase 8 benchmark
|
| 1560 |
+
numbers."*
|
| 1561 |
+
- **The decode function.** `decode_cell_relative` (`specialists/grounding/head.py:204`) turns
|
| 1562 |
+
`(B, N, 4)` cell-relative parameters into `(B, N, 4)` normalised xyxy, clamped to `[0, 1]` and
|
| 1563 |
+
order-corrected by `enforce_order` (which swaps inverted corners rather than letting IoU go silently to
|
| 1564 |
+
zero). The decode is *"differentiable: the output feeds the box and GIoU losses directly, so the loss is
|
| 1565 |
+
computed on exactly the coordinates the metric measures."*
|
| 1566 |
+
|
| 1567 |
+
---
|
| 1568 |
+
|
| 1569 |
+
## 12. Worked examples
|
| 1570 |
+
|
| 1571 |
+
### 12.1 A 4-band Cartosat-2S optical scene
|
| 1572 |
+
|
| 1573 |
+
1. **Inspect.** `inspect_raster` reads `band_count=4`, no CRS or with CRS. `infer_modality(4)` returns
|
| 1574 |
+
`OPTICAL` (4 is in `_OPTICAL_BAND_COUNTS`).
|
| 1575 |
+
2. **Descriptor.** With no `SensorDescriptor`, `_descriptor_for` maps the 4 bands positionally onto
|
| 1576 |
+
`OPTICAL_CANONICAL[:4]` = `B01, B02, B03, B04` and appends a warning naming the assumption. With a
|
| 1577 |
+
descriptor built by `build_optical_adapter("cartosat_2s")`, the bands are `["B1","B2","B3","B4"]`
|
| 1578 |
+
canonicalised to `B01..B04`, and `availability = [True, True, True, True, False Γ 8]`.
|
| 1579 |
+
3. **Canonicalise.** `adapt_optical` returns a `(12, H, W)` tensor with channels 0β3 carrying the real
|
| 1580 |
+
bands and channels 4β11 exactly zero, plus a 12-element mask with four `True`.
|
| 1581 |
+
4. **Resize.** `resize_to_canonical` bilinearly resizes to `120Γ120`. The zero channels stay zero.
|
| 1582 |
+
5. **Encode.** `encoder.encode(optical_tensor[None], sar_tensor[None])` β **no mask** (C-1).
|
| 1583 |
+
6. **Assemble.** `assemble_fusion_input` concatenates the three 768-d GAPs with the 12-element optical
|
| 1584 |
+
mask and the 2-element SAR mask β `(1, 2318)`.
|
| 1585 |
+
7. **Head.** With a trained head, the classifier emits 19 logits β softmax β a label and a margin. With
|
| 1586 |
+
no head, `probabilities` stays `None`, `degraded=True`, and no label is invented.
|
| 1587 |
+
8. **Evidence.** An `AVAILABILITY_MASK` item per modality records the sensor, the 4 real bands, the band
|
| 1588 |
+
map, the mask (`[true Γ 4, false Γ 8]`), and `missing_channels: [4,5,6,7,8,9,10,11]`. A
|
| 1589 |
+
`JOINT_FEATURE_REGION` item records `fusion_input_dim: 2318`, `mask_consumed_by: "fusion_head"`,
|
| 1590 |
+
`croma_received_mask: false`.
|
| 1591 |
+
9. **Confidence.** `optical_confidence = 4/12 = 0.3333`; the SAR side contributes its own fraction; the
|
| 1592 |
+
fusion margin (if a head exists) is the largest weight. If no head, the raw confidence is **0.0** with
|
| 1593 |
+
`no_prediction: 1.0` recorded.
|
| 1594 |
+
|
| 1595 |
+
### 12.2 A RISAT SAR scene with no band labels
|
| 1596 |
+
|
| 1597 |
+
1. **Descriptor.** `positional_fallback_descriptor(stem, 2)` selects `("VV", "VH")` from
|
| 1598 |
+
`_FALLBACK_SAR_ORDER` and records the fallback in the `sensor` field. The specialist appends: *"The
|
| 1599 |
+
polarisation IDENTITY of each channel is NOT known from a band count alone."*
|
| 1600 |
+
2. **Canonicalise.** `adapt_sar` returns `(2, H, W)` with both slots filled, mask `[True, True]`.
|
| 1601 |
+
3. **Trace.** `available_bands = ["VV", "VH"]` is the **assumption**, not a measurement. A caller with
|
| 1602 |
+
real polarisation labels passes `available_bands=["HH","HV"]`, in which case the descriptor's own order
|
| 1603 |
+
becomes the slot order and the mask reflects it β no relabelling, no discarded data.
|
| 1604 |
+
|
| 1605 |
+
### 12.3 A change pair with differing dimensions
|
| 1606 |
+
|
| 1607 |
+
1. **Validate.** Exactly two assets, distinct paths/SHA-256, CRS compared.
|
| 1608 |
+
2. **Register.** `measure_registration` collapses both to grayscale float32 and runs `phaseCorrelate`.
|
| 1609 |
+
3. **Shape check.** If the two images differ in size, `STANetStyleChangeDetector.forward` raises
|
| 1610 |
+
`SpecialistError("T1 and T2 must have the same shape; got ...")` β the 726Β²/736Β² case (Β§9.3). No resize
|
| 1611 |
+
step exists to reconcile them.
|
| 1612 |
+
4. **If shapes match but registration is poor:** the maps are still computed, the regions are withheld,
|
| 1613 |
+
the transform is dropped from the `geospatial` block, and the confidence is floored at 0.0 with
|
| 1614 |
+
`suppressed_by_registration: 1.0` recorded.
|
| 1615 |
+
|
| 1616 |
+
---
|
| 1617 |
+
|
| 1618 |
+
## 13. What is NOT implemented β exhaustive
|
| 1619 |
+
|
| 1620 |
+
Each item below was checked against the code. Where a capability is absent, it is named rather than
|
| 1621 |
+
implied; where the absence cannot be confirmed from the files read, it is marked `UNKNOWN`.
|
| 1622 |
+
|
| 1623 |
+
| Capability | Status | Evidence |
|
| 1624 |
+
|---|---|---|
|
| 1625 |
+
| **Reprojection pipeline** | **NOT IMPLEMENTED** | `geospatial/crs.py::compare_crs` *reports* `requires_reprojection`; no code performs a reprojection. `preprocessing/imagery.py` explicitly does not reproject. |
|
| 1626 |
+
| **Cloud masking** | **NOT IMPLEMENTED** | No cloud, cirrus, QA or `SCL` handling anywhere in `preprocessing/` or the specialists. `UNKNOWN β not established from the available evidence` whether any dataset adapter applies a cloud mask. |
|
| 1627 |
+
| **Atmospheric correction** | **NOT IMPLEMENTED** | No atmospheric, surface-reflectance or `L2A`-processing code. Rasters are consumed as delivered. |
|
| 1628 |
+
| **Mosaicking** | **NOT IMPLEMENTED** | No mosaicking, seamline or composite code exists. |
|
| 1629 |
+
| **Tiling / top-K tile selection** | **DECLARED (not read)** | `image.tile_size`, `image.tile_overlap`, `image.max_tiles`, `image.top_k_tiles` are read only by `core/config.py`'s `top_k_tiles <= max_tiles` check. No serving-path tiling. |
|
| 1630 |
+
| **Optical percentile 2/98 conditioning** | **DECLARED (not read)** | `optical.normalization`, `optical.lower_percentile`, `optical.upper_percentile` are read by no code (`radiometry.py` docstring; PHASE14 Β§1). |
|
| 1631 |
+
| **SAR dB clip β30/+5 conditioning** | **DECLARED (not read)** | `sar.representation`, `sar.clip_min_db`, `sar.clip_max_db` are read by no code. |
|
| 1632 |
+
| **Resampling / cropping on the display path** | **DELIBERATELY ABSENT** | `preprocessing/imagery.py` docstring: *"It does not resample, crop, or reproject."* |
|
| 1633 |
+
| **Resize to reconcile differing pair dimensions** | **NOT IMPLEMENTED** | The change path requires equal shapes and has no reconciling resize (Β§9.3). |
|
| 1634 |
+
| **Nodata filling** | **NOT IMPLEMENTED** | `quality.py` refuses any non-finite input outright; its comment records: *"Filling nodata from the raster profile belongs in the tiling work."* |
|
| 1635 |
+
| **Area computation for geographic CRS** | **NOT IMPLEMENTED** | `pixel_area_m2` returns `None` for a non-metric CRS rather than approximating. |
|
| 1636 |
+
| **GeoJSON / WKT / KML output** | **NOT IMPLEMENTED** | Coordinates are emitted as lists; no geometry serialisation exists. |
|
| 1637 |
+
| **Multi-band raster β RGB band selection policy** | **PARTIAL** | `load_image_array` takes the first three bands as RGB (or repeats a single band); there is no configurable band combination. |
|
| 1638 |
+
| **Sensor adapter for sensors beyond the known list** | **REFUSED BY DESIGN** | An unknown optical sensor with no explicit band list raises `UnsupportedBandsError`; `_known_sar` returns `[]` for an unknown SAR sensor. |
|
| 1639 |
+
|
| 1640 |
+
**A note on the "not implemented" list.** None of these omissions is hidden. The two most consequential β
|
| 1641 |
+
the tiling policy and the radiometric conditioning β are declared in the frozen config and read by no
|
| 1642 |
+
code, and both the config comments and the DEV-2 ruling name that state explicitly rather than letting the
|
| 1643 |
+
key read as a satisfied requirement.
|
| 1644 |
+
|
| 1645 |
+
---
|
| 1646 |
+
|
| 1647 |
+
## 14. What is NOT RUN / OPEN / BLOCKED for this topic
|
| 1648 |
+
|
| 1649 |
+
### NOT RUN
|
| 1650 |
+
|
| 1651 |
+
- **The optical-SAR forward pass at serving time with a trained head.** `CROMA_base.pt` is not present
|
| 1652 |
+
locally and `use_croma.py` must be vendored; the specialist runs sensor-only and produces no label.
|
| 1653 |
+
- **A system-level geospatial accuracy benchmark.** No end-to-end benchmark exists; no system-level
|
| 1654 |
+
accuracy is claimed anywhere.
|
| 1655 |
+
- **Reprojection of a genuinely mismatched-CRS pair.** The code path that would consume a reprojection is
|
| 1656 |
+
absent, so the case has never been exercised.
|
| 1657 |
+
- **Change on a bundled pair with equal dimensions.** The only bundled demo pair is 726Β² vs 736Β² and
|
| 1658 |
+
errors; a same-shape pair has not been run as a demo.
|
| 1659 |
+
|
| 1660 |
+
### OPEN
|
| 1661 |
+
|
| 1662 |
+
- **The `optical.*` / `sar.*` conditioning stage.** `OPEN` β either it gets an implementation or the keys
|
| 1663 |
+
are retired (`docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§6 item 3).
|
| 1664 |
+
- **The arm-constructibility question.** The Β§4-registered arm B (*"percentile/dB only, no encoder-input
|
| 1665 |
+
stretch"*) is **not constructible** while stage 1 is unimplemented; the arm now named B corresponds to
|
| 1666 |
+
the registered arm C. The ruling records this as *"a human decision and deliberately not resolved."*
|
| 1667 |
+
- **The optical-SAR ruling.** Accuracy **0.931** with macro-F1 **0.434161**; ruling **OPEN**. Never quote
|
| 1668 |
+
the accuracy without the macro-F1.
|
| 1669 |
+
- **The grounding resolution question after a trained head.** Re-opening 224 is legitimate only as a new
|
| 1670 |
+
pre-registered experiment.
|
| 1671 |
+
- **The change-VQA metric ruling.** `OPEN` and owner-gated.
|
| 1672 |
+
- **Whether any call site builds a `Box` outside the two conversion functions** (Β§5.4) β `UNKNOWN`.
|
| 1673 |
+
|
| 1674 |
+
### BLOCKED
|
| 1675 |
+
|
| 1676 |
+
- **A clean change demo.** Blocked on a same-shape pair or a resize step.
|
| 1677 |
+
- **The CROMA normalisation experiment (option c).** Blocked on arm B being non-constructible and on the
|
| 1678 |
+
Arm-B feature cache, which does not exist (`artifacts/optical_sar/fusion_features/` holds only the
|
| 1679 |
+
Arm-A cache).
|
| 1680 |
+
- **The gated experiment's sample size.** Cannot be specified until a variance-estimation pass is run;
|
| 1681 |
+
the ruling deliberately declines to invent a number.
|
| 1682 |
+
|
| 1683 |
+
### REJECTED
|
| 1684 |
+
|
| 1685 |
+
- **448 px grounding resolution.** `REJECTED` by the pre-registered rule and the paired test.
|
| 1686 |
+
- **Option (b) β match `base.yaml` alone, accept the shift.** `REJECTED` as a standalone path
|
| 1687 |
+
(`docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` Β§2.1).
|
| 1688 |
+
- **Reusing `preprocessing/imagery.py`'s stretch for CROMA.** `REJECTED` as `forbidden` (Β§3.3 of the
|
| 1689 |
+
ruling).
|
| 1690 |
+
- **Coherence- and texture-based registration gates.** One `REJECTED`, one `CONFIRMED but INSUFFICIENT`.
|
| 1691 |
+
|
| 1692 |
+
---
|
| 1693 |
+
|
| 1694 |
+
## 15. Where the evidence lives
|
| 1695 |
+
|
| 1696 |
+
**Source modules (the code this chapter describes):**
|
| 1697 |
+
|
| 1698 |
+
| Path | What it defines |
|
| 1699 |
+
|---|---|
|
| 1700 |
+
| `preprocessing/raster.py` | The validation chain, `inspect_raster`, `read_bands`, `write_raster`, `infer_modality`, `file_sha256`, `pixel_area_m2` |
|
| 1701 |
+
| `preprocessing/imagery.py` | `load_image_array`, `load_pil_image`, the per-scene 2/98 display stretch |
|
| 1702 |
+
| `preprocessing/quality.py` | The deterministic quality gate, `MIN_AUTOCORRELATION=0.10`, `NOISE_ENTROPY_BITS=7.8` |
|
| 1703 |
+
| `geospatial/crs.py` | `parse_crs`, `describe_crs`, `compare_crs`, `is_metric`, `CRSCompatibility` |
|
| 1704 |
+
| `geospatial/transform.py` | `PixelWindow`, `GeoBounds`, `to_normalized`/`to_pixel`/`to_geo`, `intersect`, `iou`, `benchmark_boxes_to_normalized`, `affine_from_metadata` |
|
| 1705 |
+
| `core/schemas.py` | `CoordinateSystem`, `GeoMetadata`, `SensorDescriptor`, `AssetMetadata`, `Box`, `Region`, `ChangeRegion`, `Evidence._spatial_needs_crs` |
|
| 1706 |
+
| `core/config.py` | Config loading, validation, `Config.hash`, `device_preference` |
|
| 1707 |
+
| `specialists/optical_sar/sensor_adapter.py` | `OPTICAL_CANONICAL`, `SAR_CANONICAL`, `build_optical_adapter`, `describe_sar`, `_apply_canonical`, `positional_fallback_descriptor` |
|
| 1708 |
+
| `specialists/optical_sar/radiometry.py` | `TRANSFORM_NAME`, `normalise_for_croma`, `resolve_use_8_bit`, the zero-channel rule |
|
| 1709 |
+
| `specialists/optical_sar/fusion_head.py` | `expected_fusion_dim`, `assemble_fusion_input`, `channel_dropout`, `build_fusion_head` |
|
| 1710 |
+
| `specialists/optical_sar/inference.py` | `run_pipeline`, `resize_to_canonical`, `cross_modal_agreement`, `build_head_from_config` |
|
| 1711 |
+
| `specialists/optical_sar/specialist.py` | `_assess_pair`, `_descriptor_for`, `_confidence_components`, `_geospatial` |
|
| 1712 |
+
| `specialists/change/postprocess.py` | `RegistrationQuality`, `measure_registration`, `to_grayscale_float`, `regions_to_schema` |
|
| 1713 |
+
| `specialists/change/specialist.py` | `_assess_pair`, `_predict`, `_geospatial`, the suppression path |
|
| 1714 |
+
| `specialists/change/stanet.py:415` | The equal-shape guard |
|
| 1715 |
+
| `specialists/grounding/specialist.py` | `_to_geo_boxes`, the `NORMALIZED_0_1` box construction, `DEGENERATE_AREA_FRACTION` |
|
| 1716 |
+
| `specialists/grounding/head.py` | `decode_cell_relative`, `enforce_order`, `nms` |
|
| 1717 |
+
| `specialists/grounding/inference.py` | `decode_candidates_from_features`, `ground_phrase` |
|
| 1718 |
+
| `evaluation/metrics/grounding.py` | `benchmark_to_normalized(box, scale=100.0)` |
|
| 1719 |
+
|
| 1720 |
+
**Configuration:**
|
| 1721 |
+
|
| 1722 |
+
- `configs/base.yaml` β `image.*` (tiling), `optical.*`, `sar.*`, `croma.*`, `fusion.*`, `grounding.*`,
|
| 1723 |
+
`grounding_head.*`, `change.*`.
|
| 1724 |
+
|
| 1725 |
+
**Documents:**
|
| 1726 |
+
|
| 1727 |
+
- `docs/PHASE14_CROMA_NORMALISATION_CHANGE.md` β the DEV-2 ruling, the two ordered stages, the Gate F arm
|
| 1728 |
+
amendment.
|
| 1729 |
+
- `docs/CROMA_NORMALISATION_UPSTREAM_EVIDENCE.md` β what is and is not established about CROMA's
|
| 1730 |
+
pretraining distribution.
|
| 1731 |
+
- `docs/PHASE7_RESOLUTION_DECISION.md` β the 224/448 decision in full.
|
| 1732 |
+
- `docs/CHANGE_REGISTRATION_GATE_COHERENCE_TEST.md`, `docs/CHANGE_REGISTRATION_GATE_TEXTURE_TEST.md` β the
|
| 1733 |
+
rejected and insufficient registration gates.
|
| 1734 |
+
- `docs/API_CONTRACT.md` Β§2.5 (asset allowlist), Β§3.2 (coordinate systems).
|
| 1735 |
+
- `docs/FINAL_DELIVERY_REPORT.md` Β§6 β the 726Β²/736Β² change demo recorded as DEGRADED.
|
| 1736 |
+
- `docs/PHASE10_CDVQA_DATA_STATUS.md` β the CDVQA/SECOND pair layout and temporal semantics.
|
| 1737 |
+
- `docs/PHASE11_CROMA_STATUS.md` β the reproduced CROMA forward pass.
|
| 1738 |
+
|
| 1739 |
+
**Sibling release chapters:**
|
| 1740 |
+
|
| 1741 |
+
- [`DATA_PIPELINE.md`](DATA_PIPELINE.md) β the end-to-end data path, caches and artefacts.
|
| 1742 |
+
- [`architecture/03-request-lifecycle.md`](architecture/03-request-lifecycle.md) β the nine controller
|
| 1743 |
+
states, including VALIDATE and PREPROCESS.
|
| 1744 |
+
- [`architecture/05-specialists.md`](architecture/05-specialists.md) β the six specialists.
|
| 1745 |
+
- [`architecture/07-configuration-freeze.md`](architecture/07-configuration-freeze.md) β the frozen hash
|
| 1746 |
+
and every config key.
|
| 1747 |
+
- [`DATASETS.md`](DATASETS.md) β the corpora and their measured shapes.
|