Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
release: add docs/DEVELOPMENT.md
Browse files- docs/DEVELOPMENT.md +1330 -0
docs/DEVELOPMENT.md
ADDED
|
@@ -0,0 +1,1330 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Development Guide
|
| 2 |
+
|
| 3 |
+
**Status tags:** `IMPLEMENTED` Β· `VERIFIED` Β· `MEASURED` Β· `ATTEMPTED` Β· `NOT RUN` Β· `BLOCKED` Β·
|
| 4 |
+
`DEFERRED` Β· `OPEN` Β· `RESOLVED` Β· `BY DESIGN`.
|
| 5 |
+
|
| 6 |
+
This is the contributor-facing guide to the SatQuery AI codebase. It answers: *what do I need to
|
| 7 |
+
install*; *how do I run it*; *what is each directory for*; *what rules will break the build if I break
|
| 8 |
+
them*; *how do I add a component*; *how do I test*; and *which environment traps will cost me an hour if
|
| 9 |
+
nobody tells me about them*.
|
| 10 |
+
|
| 11 |
+
Every command, path, timeout and rule below comes from a file that was read. Where the evidence does not
|
| 12 |
+
settle a question, the text says `UNKNOWN β not established from the available evidence` rather than
|
| 13 |
+
guessing.
|
| 14 |
+
|
| 15 |
+
> **Read this first.** `configs/base.yaml` is the single registry, and its hash β **`78f1e3700da15aa1`**
|
| 16 |
+
> β is frozen. **Editing the config moves the hash and invalidates every artifact keyed to it.** This is
|
| 17 |
+
> the one rule in the repository whose violation is not locally reversible (Β§7).
|
| 18 |
+
|
| 19 |
+
---
|
| 20 |
+
|
| 21 |
+
## Table of contents
|
| 22 |
+
|
| 23 |
+
**Part I β Prerequisites and setup**
|
| 24 |
+
1. Prerequisites
|
| 25 |
+
2. Local setup
|
| 26 |
+
3. Running the service locally
|
| 27 |
+
|
| 28 |
+
**Part II β Repository layout**
|
| 29 |
+
4. The top-level directory map
|
| 30 |
+
5. Where the contracts live
|
| 31 |
+
6. The package layout inside `core/`, `specialists/`, `app/`
|
| 32 |
+
|
| 33 |
+
**Part III β The configuration discipline**
|
| 34 |
+
7. "No magic numbers in Python"
|
| 35 |
+
8. What happens if you break the rule
|
| 36 |
+
9. The hash-exempt path: environment overrides
|
| 37 |
+
10. The invariants the loader enforces
|
| 38 |
+
|
| 39 |
+
**Part IV β Contract-first development**
|
| 40 |
+
11. `core/schemas.py` is the binding contract
|
| 41 |
+
12. The specialist interface
|
| 42 |
+
13. How to add a new specialist
|
| 43 |
+
14. How evidence is emitted
|
| 44 |
+
|
| 45 |
+
**Part V β The test workflow**
|
| 46 |
+
15. Running tests
|
| 47 |
+
16. The evidence-class markers
|
| 48 |
+
17. The known failures, and why they are not regressions
|
| 49 |
+
|
| 50 |
+
**Part VI β Scripts and notebooks**
|
| 51 |
+
18. The class of scripts under `scripts/`
|
| 52 |
+
19. `notebooks/`
|
| 53 |
+
|
| 54 |
+
**Part VII β Training entry points**
|
| 55 |
+
20. Local training (router, grounding, change, optical-SAR)
|
| 56 |
+
21. External GPU training (change-VQA, VLM LoRA)
|
| 57 |
+
22. Calibration
|
| 58 |
+
|
| 59 |
+
**Part VIII β Coding conventions**
|
| 60 |
+
23. What the code actually does
|
| 61 |
+
|
| 62 |
+
**Part IX β Known development traps**
|
| 63 |
+
24. The traps, in one table
|
| 64 |
+
25. Trap 1 β the stale untracked `deploy/`
|
| 65 |
+
26. Trap 2 β the dead sandbox proxy
|
| 66 |
+
27. Trap 3 β pytest only in the venv
|
| 67 |
+
28. Trap 4 β Chrome drops synthetic CDP key events
|
| 68 |
+
29. Trap 5 β Cloudflare `_headers` concatenate
|
| 69 |
+
30. Trap 6 β the annotation-scope trap
|
| 70 |
+
31. Trap 7 β the full-suite bulk-delete guard
|
| 71 |
+
32. Trap 8 β the stale serve process
|
| 72 |
+
|
| 73 |
+
**Part X β Status and evidence**
|
| 74 |
+
33. `NOT RUN` / `OPEN` / `BLOCKED` / `UNKNOWN` for development
|
| 75 |
+
34. Where the evidence lives
|
| 76 |
+
|
| 77 |
+
---
|
| 78 |
+
|
| 79 |
+
# Part I β Prerequisites and setup
|
| 80 |
+
|
| 81 |
+
## 1. Prerequisites
|
| 82 |
+
|
| 83 |
+
| Requirement | Value | Source |
|
| 84 |
+
|---|---|---|
|
| 85 |
+
| Python | **3.11+** | `release/repo/README.md` Β§Installation |
|
| 86 |
+
| Device | **a CPU is sufficient**; no CUDA requirement | `release/repo/README.md` Β§Installation |
|
| 87 |
+
| GPU | not required | `docs/DEPLOYMENT_DECISION.md` Β§5 |
|
| 88 |
+
| Git | needed to clone | `release/repo/README.md` Β§Installation |
|
| 89 |
+
|
| 90 |
+
The codebase is CPU-first and the deployment is CPU-only. This is not a fallback β it is the design:
|
| 91 |
+
|
| 92 |
+
- `core/config.py:86-91` β `device_preference` honours the `SATQUERY_DEVICE` env override, else
|
| 93 |
+
`"cuda" if torch.cuda.is_available() else "cpu"`.
|
| 94 |
+
- Every specialist defaults to `device="cpu"` (`docs/DEPLOYMENT_DECISION.md` Β§5).
|
| 95 |
+
- **No `.cuda()` call exists anywhere**; all placement is `.to(device)` (`docs/DEPLOYMENT_DECISION.md`
|
| 96 |
+
Β§5).
|
| 97 |
+
- `configs/base.yaml:293` β `cpu_mode_required: true`.
|
| 98 |
+
|
| 99 |
+
The development container that runs the inference tier pins Python 3.12
|
| 100 |
+
(`.devcontainer/devcontainer.json:3`, `image: mcr.microsoft.com/devcontainers/python:3.12`). The
|
| 101 |
+
repository's own guidance is 3.11+; a 3.11 or 3.12 interpreter both work.
|
| 102 |
+
|
| 103 |
+
**Not implemented** (noted, not needed): no thread capping (`torch.set_num_threads`) and no
|
| 104 |
+
quantisation. `lazy_load` and `cache_max_models` are only *reported* (`app/deployment.py:609-610`,
|
| 105 |
+
`:896-897`); **UNVERIFIED** whether any code enforces a one-model cache (`docs/DEPLOYMENT_DECISION.md`
|
| 106 |
+
Β§5).
|
| 107 |
+
|
| 108 |
+
## 2. Local setup
|
| 109 |
+
|
| 110 |
+
```bash
|
| 111 |
+
git clone https://github.com/Anish-lab-blip/SatQuery-AI
|
| 112 |
+
cd SatQuery-AI
|
| 113 |
+
python -m venv .venv
|
| 114 |
+
source .venv/Scripts/activate # Windows git-bash; use .venv/bin/activate on Linux/macOS
|
| 115 |
+
pip install -r requirements.txt
|
| 116 |
+
```
|
| 117 |
+
|
| 118 |
+
(`release/repo/README.md` Β§Installation)
|
| 119 |
+
|
| 120 |
+
### 2.1 The dependency manifest and its two profiles
|
| 121 |
+
|
| 122 |
+
`requirements.txt` is **frozen at architecture v1.0** and declares two install profiles in its header
|
| 123 |
+
comment (`requirements.txt:1-8`):
|
| 124 |
+
|
| 125 |
+
```
|
| 126 |
+
# Two install profiles:
|
| 127 |
+
# CPU (local dev / schema / geo / unit tests):
|
| 128 |
+
# pip install -r requirements.txt
|
| 129 |
+
# GPU (Kaggle T4x2 / HF ZeroGPU): torch is preinstalled on both.
|
| 130 |
+
# Do NOT pin torch here β Kaggle and HF ship their own builds.
|
| 131 |
+
```
|
| 132 |
+
|
| 133 |
+
**torch is deliberately not pinned** (`requirements.txt:21`). The platform supplies it. This is why the
|
| 134 |
+
local CPU install works on a machine with no CUDA and why the Kaggle/ZeroGPU targets get their own
|
| 135 |
+
builds.
|
| 136 |
+
|
| 137 |
+
The manifest's sections, with the contract notes the file itself carries:
|
| 138 |
+
|
| 139 |
+
| Section | Packages | Note |
|
| 140 |
+
|---|---|---|
|
| 141 |
+
| core | `numpy`, `pyyaml`, `pydantic` | β |
|
| 142 |
+
| geospatial | `rasterio`, `pyproj`, `opencv-python-headless` | β |
|
| 143 |
+
| models | `transformers>=4.52`, `open-clip-torch>=2.24`, `sentence-transformers>=2.7`, `peft>=0.10`, `huggingface_hub>=0.23`, `safetensors>=0.4`, `einops>=0.7` | see below |
|
| 144 |
+
| ui + reporting | `gradio>=4.44`, `reportlab>=4.1` | β |
|
| 145 |
+
| dev | `pytest>=8.0`, `pytest-cov>=5.0` | β |
|
| 146 |
+
|
| 147 |
+
Three contract notes in the file are load-bearing (`requirements.txt:43-53`):
|
| 148 |
+
|
| 149 |
+
- `transformers>=4.52` β SmolVLM via `AutoModelForImageTextToText`; `AutoModelForVision2Seq` is
|
| 150 |
+
deprecated (finding C-2).
|
| 151 |
+
- `open-clip-torch` β the RemoteCLIP checkpoint is loaded via `pretrained=<path>` so
|
| 152 |
+
`load_checkpoint()` runs its state-dict fixups (C-4).
|
| 153 |
+
- `huggingface_hub` β used to pin the RemoteCLIP / SmolVLM / CROMA revisions.
|
| 154 |
+
|
| 155 |
+
> **The `einops` lesson.** `einops>=0.7` is **not optional**. It is required by the vendored
|
| 156 |
+
> `specialists/optical_sar/vendor/use_croma.py` (`from einops import rearrange`). It was found missing on
|
| 157 |
+
> 2026-09-18 by executing the vendored module: it raised `ModuleNotFoundError`, so the CROMA path was
|
| 158 |
+
> blocked on a dependency that no document declared (`requirements.txt:28-33`).
|
| 159 |
+
|
| 160 |
+
### 2.2 Device selection
|
| 161 |
+
|
| 162 |
+
Set `SATQUERY_DEVICE` to force a device:
|
| 163 |
+
|
| 164 |
+
```bash
|
| 165 |
+
export SATQUERY_DEVICE=cpu # cpu | cuda | mps | null
|
| 166 |
+
```
|
| 167 |
+
|
| 168 |
+
`device_preference` reads the env var first and only falls back to a torch probe if it is unset
|
| 169 |
+
(`core/config.py:87-91`). The value is read **without importing torch** on the metadata path, which is
|
| 170 |
+
what keeps the health route cheap (`docs/DEPLOYMENT_TOPOLOGY.md` Β§3.3).
|
| 171 |
+
|
| 172 |
+
## 3. Running the service locally
|
| 173 |
+
|
| 174 |
+
The inference service is served by the same launcher the Codespace runs
|
| 175 |
+
(`release/repo/README.md` Β§Local development):
|
| 176 |
+
|
| 177 |
+
```bash
|
| 178 |
+
# Inference service, CPU (this is the launcher the Codespace runs)
|
| 179 |
+
PORT=8000 python deploy/codespace/serve.py
|
| 180 |
+
|
| 181 |
+
# Health
|
| 182 |
+
curl localhost:8000/v1/health
|
| 183 |
+
```
|
| 184 |
+
|
| 185 |
+
`serve.py` builds the app through `build_space_app()` and binds it with uvicorn on `$PORT` (default 8000)
|
| 186 |
+
(`deploy/codespace/serve.py:16-25`). It imports cheaply β FastAPI is imported inside `build_space_app()`
|
| 187 |
+
and no model is loaded at module scope β so it stays import-safe on a CPU host with no GPU and no weights
|
| 188 |
+
present (`deploy/codespace/serve.py:1-13`).
|
| 189 |
+
|
| 190 |
+
The server answers four routes (`app/space_app.py`):
|
| 191 |
+
|
| 192 |
+
| Route | Line | Purpose |
|
| 193 |
+
|---|---|---|
|
| 194 |
+
| `GET /v1/health` | `app/space_app.py:521` | liveness, derived status, torch-free device probe |
|
| 195 |
+
| `GET /v1/capabilities` | `app/space_app.py:549` | the six declared capabilities |
|
| 196 |
+
| `POST /v1/assets` | `app/space_app.py:555` | handle-based upload (off unless enabled) |
|
| 197 |
+
| `POST /v1/analyze` | `app/space_app.py:661` | the analysis path |
|
| 198 |
+
|
| 199 |
+
`POST /v1/assets` is **off unless explicitly enabled** (`SATQUERY_ASSET_ENABLED` **and**
|
| 200 |
+
`SATQUERY_ASSET_DIR` must both be set; otherwise it answers `503` rather than defaulting to a temp
|
| 201 |
+
directory) (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` Β§3.1.1).
|
| 202 |
+
|
| 203 |
+
### 3.1 Serving the frontend locally
|
| 204 |
+
|
| 205 |
+
The frontend is fully static and needs no build step (`release/repo/README.md` Β§Local development):
|
| 206 |
+
|
| 207 |
+
```bash
|
| 208 |
+
python -m http.server 5500 --directory frontend
|
| 209 |
+
```
|
| 210 |
+
|
| 211 |
+
The gateway's dev-origin allowlist includes `localhost` and `127.0.0.1` on ports 3000, 5500, 5173, 8000
|
| 212 |
+
and 8080 (`deploy/render/main.py:139-143`), so a local dev server on any of those ports is accepted
|
| 213 |
+
without an env-var change.
|
| 214 |
+
|
| 215 |
+
### 3.2 Running the gateway locally
|
| 216 |
+
|
| 217 |
+
```bash
|
| 218 |
+
PORT=10000 \
|
| 219 |
+
GITHUB_TOKEN=<token> \
|
| 220 |
+
CODESPACE_NAME=<codespace> \
|
| 221 |
+
SATQUERY_ALLOWED_ORIGINS="https://satquery.pages.dev" \
|
| 222 |
+
uvicorn deploy.render.main:app --port 10000
|
| 223 |
+
```
|
| 224 |
+
|
| 225 |
+
(`deploy/render/README.md` Β§Running locally β the placeholder is the repository's own.)
|
| 226 |
+
|
| 227 |
+
Then exercise it:
|
| 228 |
+
|
| 229 |
+
```bash
|
| 230 |
+
curl http://localhost:10000/api/health
|
| 231 |
+
curl -X POST http://localhost:10000/api/infer -H 'content-type: application/json' -d '{"query":"..."}'
|
| 232 |
+
curl http://localhost:10000/api/capabilities
|
| 233 |
+
```
|
| 234 |
+
|
| 235 |
+
> **The monorepo `deploy/` is stale** (Β§25). Use it for local experimentation only; it is **not** the
|
| 236 |
+
> deployed source and it lacks the tunnel code the live service runs.
|
| 237 |
+
|
| 238 |
+
---
|
| 239 |
+
|
| 240 |
+
# Part II β Repository layout
|
| 241 |
+
|
| 242 |
+
## 4. The top-level directory map
|
| 243 |
+
|
| 244 |
+
| Directory / file | Purpose | Source |
|
| 245 |
+
|---|---|---|
|
| 246 |
+
| `app/` | the serving composition root and the FastAPI entrypoint | `app/serving.py`, `app/space_app.py`, `app/deployment.py` |
|
| 247 |
+
| `core/` | the config loader, the typed schemas, the controller, the planner, the registry, the error taxonomy | `core/config.py`, `core/schemas.py`, `core/controller.py`, `core/planner.py`, `core/registry.py`, `core/errors.py`, `core/code_revision.py` |
|
| 248 |
+
| `router/` | the MiniLM intent router: encoder, classifier, adapter, training, label space, lexical fallback | `router/encoder.py`, `router/classifier.py`, `router/adapter.py`, `router/train.py`, `router/dataset.py`, `router/fallback.py`, `router/label_space.py` |
|
| 249 |
+
| `specialists/` | the six specialist implementations behind one interface | `specialists/base.py`, `specialists/vqa/`, `specialists/grounding/`, `specialists/change/`, `specialists/optical_sar/` |
|
| 250 |
+
| `preprocessing/` | raster loading, imagery, quality checks | `preprocessing/raster.py`, `preprocessing/imagery.py`, `preprocessing/quality.py` |
|
| 251 |
+
| `geospatial/` | CRS handling and transforms | `geospatial/crs.py`, `geospatial/transform.py` |
|
| 252 |
+
| `evidence/` | the evidence engine and confidence | `evidence/engine.py`, `evidence/confidence.py` |
|
| 253 |
+
| `evaluation/` | manifests, leakage, metrics, normalisation, the runner, benchmark adapters, the frozen prompt set | `evaluation/manifests.py`, `evaluation/leakage.py`, `evaluation/metrics/`, `evaluation/normalize.py`, `evaluation/runner.py`, `evaluation/benchmark_adapters/`, `evaluation/prompt_freeze.json`, `evaluation/manifest_freeze.json` |
|
| 254 |
+
| `training/` | the training loops and data adapters, split by task | `training/router/`, `training/grounding/`, `training/change/`, `training/change_vqa/`, `training/fusion/`, `training/vlm/`, `training/calibration/`, `training/data/` |
|
| 255 |
+
| `configs/` | the frozen registry and the frozen deploy manifest | `configs/base.yaml`, `configs/deploy.yaml` |
|
| 256 |
+
| `gateway/` | the standalone gateway: policy (decisions) + app (HTTP plumbing) | `gateway/policy.py`, `gateway/app.py`, `gateway/assets.py` |
|
| 257 |
+
| `deploy/` | the deployment launchers β **stale and untracked in the monorepo** | `deploy/codespace/`, `deploy/render/` (Β§25) |
|
| 258 |
+
| `tests/` | the suites, split by concern and evidence class | `tests/unit/`, `tests/integration/`, `tests/routing/`, `tests/geospatial/`, `tests/leakage/`, `tests/model/`, `tests/e2e/` |
|
| 259 |
+
| `scripts/` | the flat set of executable helpers | Β§18 |
|
| 260 |
+
| `notebooks/` | the Kaggle notebooks | Β§19 |
|
| 261 |
+
| `artifacts/` | trained heads, checkpoints, caches, evidence archives | ~3.7 GB total (`release/CURRENT_RELEASE_STATE.md` Β§3) |
|
| 262 |
+
| `frontend/` | the static site | staged by `scripts/stage_pages.mjs` |
|
| 263 |
+
| `docs/` | the project's own engineering records | ~60 files |
|
| 264 |
+
| `hf/` | the Hugging Face project card / model cards | β |
|
| 265 |
+
| `demo/`, `benchmark/`, `reports/`, `data/`, `logs/` | supporting material | β |
|
| 266 |
+
|
| 267 |
+
> **`training/router/` is empty** in the working copy; the router's training lives in `router/train.py`
|
| 268 |
+
> and `router/dataset.py` and is driven by `scripts/train_router.py` (Β§20).
|
| 269 |
+
|
| 270 |
+
> **The repository has no `pyproject.toml` and no `setup.py`.** `app` is a plain package, so the repo root
|
| 271 |
+
> must be on `sys.path` for `from app.space_app import ...` to resolve
|
| 272 |
+
> (`deploy/codespace/launch.sh:31-34`; `pytest.ini` sets `pythonpath = .`).
|
| 273 |
+
|
| 274 |
+
## 5. Where the contracts live
|
| 275 |
+
|
| 276 |
+
Three files are authoritative, and the code β not a document β wins when they disagree:
|
| 277 |
+
|
| 278 |
+
| Contract | File | What it fixes |
|
| 279 |
+
|---|---|---|
|
| 280 |
+
| the config registry | `configs/base.yaml` | every tunable value; hashed (Β§7) |
|
| 281 |
+
| the wire/data schemas | `core/schemas.py` | the request, result, evidence, trace and health shapes |
|
| 282 |
+
| the specialist interface | `specialists/base.py` | the four-method contract every specialist implements |
|
| 283 |
+
| the capability table | `core/registry.py` | which capabilities exist and how to build them |
|
| 284 |
+
| the planner's mapping | `core/planner.py` | `TASK_CAPABILITY` and `CAPABILITY_ASSETS` |
|
| 285 |
+
|
| 286 |
+
When `docs/API_CONTRACT.md` and `core/schemas.py` disagree, the contract document's own authority clause
|
| 287 |
+
resolves it: *"the request/response shapes are not invented here. They are the existing, tested Pydantic
|
| 288 |
+
models in `core/schemas.py`. This document describes them; it does not declare new ones."*
|
| 289 |
+
(`docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§15). The known instance of this β the forward-compatibility
|
| 290 |
+
promise versus `extra="forbid"` β is recorded as **C-2** and left as a documented contradiction rather
|
| 291 |
+
than silently resolved (`docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§15).
|
| 292 |
+
|
| 293 |
+
## 6. The package layout inside `core/`, `specialists/`, `app/`
|
| 294 |
+
|
| 295 |
+
### 6.1 `core/`
|
| 296 |
+
|
| 297 |
+
| File | Role |
|
| 298 |
+
|---|---|
|
| 299 |
+
| `core/config.py` | the loader, the validator, the hash (Β§7, Β§10) |
|
| 300 |
+
| `core/schemas.py` | every Pydantic model (Β§11) |
|
| 301 |
+
| `core/controller.py` | the nine-state controller that runs a plan and produces the trace |
|
| 302 |
+
| `core/planner.py` | the policy layer: `TASK_CAPABILITY`, `CAPABILITY_ASSETS`, `plan()` |
|
| 303 |
+
| `core/registry.py` | the capability registry: spec table, lazy construction, degradation states |
|
| 304 |
+
| `core/errors.py` | the typed error taxonomy and `scrub_paths` |
|
| 305 |
+
| `core/code_revision.py` | the code revision recorded in a run |
|
| 306 |
+
|
| 307 |
+
The three-tier control split is deliberate: the **registry** knows *which* specialists exist and *how* to
|
| 308 |
+
construct them; the **planner** *decides* what runs; the **controller** executes. The registry's own
|
| 309 |
+
docstring states it: *"It never decides what runs β that is `core.planner`'s job alone"*
|
| 310 |
+
(`core/registry.py:6-8`).
|
| 311 |
+
|
| 312 |
+
### 6.2 `specialists/`
|
| 313 |
+
|
| 314 |
+
| Package | Files | Capability |
|
| 315 |
+
|---|---|---|
|
| 316 |
+
| `specialists/vqa/` | `inference.py`, `model.py`, `prompts.py` | `vqa`, `caption` |
|
| 317 |
+
| `specialists/grounding/` | `specialist.py`, `remoteclip.py`, `head.py`, `inference.py` | `grounding` |
|
| 318 |
+
| `specialists/change/` | `specialist.py`, `stanet.py`, `vqa_specialist.py`, `postprocess.py` | `change`, `change_vqa` |
|
| 319 |
+
| `specialists/optical_sar/` | `specialist.py`, `croma.py`, `fusion_head.py`, `sensor_adapter.py`, `radiometry.py`, `inference.py`, `prompts.py`, `vendor/` | `optical_sar` |
|
| 320 |
+
|
| 321 |
+
### 6.3 `app/`
|
| 322 |
+
|
| 323 |
+
| File | Role |
|
| 324 |
+
|---|---|
|
| 325 |
+
| `app/serving.py` | the composition root β wires checkpoints through the registry's `builders=` override |
|
| 326 |
+
| `app/space_app.py` | `build_space_app()` β the FastAPI app and the four routes |
|
| 327 |
+
| `app/deployment.py` | the deployment description and device resolution |
|
| 328 |
+
|
| 329 |
+
---
|
| 330 |
+
|
| 331 |
+
# Part III β The configuration discipline
|
| 332 |
+
|
| 333 |
+
## 7. "No magic numbers in Python"
|
| 334 |
+
|
| 335 |
+
`configs/base.yaml` opens with the rule, in the file itself:
|
| 336 |
+
|
| 337 |
+
```yaml
|
| 338 |
+
# RULE: no magic numbers anywhere in Python. Everything tunable lives here.
|
| 339 |
+
# Every value below is loaded, validated and hashed by core/config.py.
|
| 340 |
+
```
|
| 341 |
+
|
| 342 |
+
(`configs/base.yaml:4-5`)
|
| 343 |
+
|
| 344 |
+
The loader reinforces it: *"One config system. No duplicated constants. Every value in configs/base.yaml
|
| 345 |
+
is loaded, validated against the frozen architecture, and hashed so evaluation runs are reproducible."*
|
| 346 |
+
(`core/config.py:1-5`).
|
| 347 |
+
|
| 348 |
+
**Access pattern.** Never read the YAML directly; import the singleton:
|
| 349 |
+
|
| 350 |
+
```python
|
| 351 |
+
from core.config import get_config
|
| 352 |
+
|
| 353 |
+
cfg = get_config()
|
| 354 |
+
cfg.get("croma.image_resolution") # dotted-path access
|
| 355 |
+
cfg.require("change.encoder") # raises ConfigError if missing
|
| 356 |
+
cfg.seed # project.seed, default 42
|
| 357 |
+
```
|
| 358 |
+
|
| 359 |
+
`get_config()` is an `lru_cache(maxsize=1)` singleton β *"Import this, do not re-read YAML"*
|
| 360 |
+
(`core/config.py:270-273`).
|
| 361 |
+
|
| 362 |
+
## 8. What happens if you break the rule
|
| 363 |
+
|
| 364 |
+
Two distinct failures, and both are loud.
|
| 365 |
+
|
| 366 |
+
### 8.1 The loader fails startup
|
| 367 |
+
|
| 368 |
+
`Config.__init__` calls `_validate()`, which collects **every** violation and raises a single
|
| 369 |
+
`ConfigError` naming them all (`core/config.py:46-49,94-222`):
|
| 370 |
+
|
| 371 |
+
```python
|
| 372 |
+
if errors:
|
| 373 |
+
raise ConfigError(
|
| 374 |
+
"configuration failed frozen-architecture validation:\n - "
|
| 375 |
+
+ "\n - ".join(errors)
|
| 376 |
+
)
|
| 377 |
+
```
|
| 378 |
+
|
| 379 |
+
The header states the intent: *"if a config tries to violate a frozen decision (e.g. bf16 on T4,
|
| 380 |
+
torch.compile on ZeroGPU, a CROMA image_resolution that is not a multiple of 8), it fails loudly rather
|
| 381 |
+
than at runtime"* (`core/config.py:6-8`).
|
| 382 |
+
|
| 383 |
+
### 8.2 Editing the config MOVES the hash
|
| 384 |
+
|
| 385 |
+
The hash is a sha256 over the whole registry, truncated to 16 hex characters
|
| 386 |
+
(`core/config.py:76-80`):
|
| 387 |
+
|
| 388 |
+
```python
|
| 389 |
+
@property
|
| 390 |
+
def hash(self) -> str:
|
| 391 |
+
"""Stable hash of the whole registry. Recorded in every evaluation run."""
|
| 392 |
+
blob = json.dumps(self._data, sort_keys=True, default=str).encode()
|
| 393 |
+
return hashlib.sha256(blob).hexdigest()[:16]
|
| 394 |
+
```
|
| 395 |
+
|
| 396 |
+
Because it is computed over the entire registry, **any** edit to `configs/base.yaml` changes it. The
|
| 397 |
+
frozen value is **`78f1e3700da15aa1`** (`release/repo/README.md` Β§Reproducibility;
|
| 398 |
+
`docs/PHASE18_DEPLOYMENT_PACKAGING.md` Β§3). Every artifact records the hash it was produced against, so a
|
| 399 |
+
hash change **invalidates every artifact keyed to it**.
|
| 400 |
+
|
| 401 |
+
> **This is the one non-reversible action in the repository.** `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` Β§6.1
|
| 402 |
+
> lists it as the single row in the rollback table whose answer to "Reversible?" is **no**: *"The recorded
|
| 403 |
+
> benchmark hash is gone; the frozen benchmark no longer matches."*
|
| 404 |
+
|
| 405 |
+
**Verify the hash after any config-adjacent change:**
|
| 406 |
+
|
| 407 |
+
```bash
|
| 408 |
+
$PY -c "from core.config import get_config; print(get_config().hash)"
|
| 409 |
+
# expected: 78f1e3700da15aa1
|
| 410 |
+
```
|
| 411 |
+
|
| 412 |
+
(`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` Β§1.1, where `$PY` is `$REPO/.venv/Scripts/python.exe` on Windows.)
|
| 413 |
+
|
| 414 |
+
**Two things that do *not* move the hash** β both verified:
|
| 415 |
+
|
| 416 |
+
1. **`configs/deploy.yaml` is inert.** It carries a `registry: false` marker and is never merged into the
|
| 417 |
+
registry, so editing it cannot move the hash. Verify with `$PY scripts/validate_deploy_config.py` β
|
| 418 |
+
exit 0, `"deploy.yaml is inert (not in the registry) and C-8-consistent."`
|
| 419 |
+
(`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` Β§2.4). **But** `scripts/validate_deploy_config.py` hard-fails if
|
| 420 |
+
the `deployment:` block in `deploy.yaml` differs key-for-key from `base.yaml`'s, so a change to one
|
| 421 |
+
alone fails the validator (`docs/DEPLOYMENT_DECISION.md` Β§4).
|
| 422 |
+
2. **The serving wiring uses `builders=`.** `app/serving.py` wires the change and change-VQA heads through
|
| 423 |
+
the registry's `builders=` override instead of config, which is *"the seam that keeps `Config.hash`
|
| 424 |
+
unchanged while still pointing serving at the trained artifacts"* (`docs/PHASE19_FINAL_HARDENING.md`
|
| 425 |
+
Β§3.4).
|
| 426 |
+
|
| 427 |
+
## 9. The hash-exempt path: environment overrides
|
| 428 |
+
|
| 429 |
+
Two registry values can be overridden from the environment without editing the YAML
|
| 430 |
+
(`core/config.py:261-265`):
|
| 431 |
+
|
| 432 |
+
| Variable | Effect |
|
| 433 |
+
|---|---|
|
| 434 |
+
| `SATQUERY_PRECISION` | overrides `training.precision` |
|
| 435 |
+
| `SATQUERY_TORCH_COMPILE` | overrides `deployment.torch_compile` (`"true"` β `True`) |
|
| 436 |
+
|
| 437 |
+
**Both are still validated.** Setting `SATQUERY_TORCH_COMPILE=true` **fails startup** because finding C-8
|
| 438 |
+
forbids `torch.compile` (`release/repo/docs/DEPLOYMENT.md` Β§6.3; `core/config.py:114-119`).
|
| 439 |
+
|
| 440 |
+
This is the general pattern for deployment state that must not move the hash: read it from the
|
| 441 |
+
environment. The asset-store capacity, TTL and per-file cap follow the same rule
|
| 442 |
+
(`release/repo/README.md` Β§Installation).
|
| 443 |
+
|
| 444 |
+
## 10. The invariants the loader enforces
|
| 445 |
+
|
| 446 |
+
`core/config.py::_validate` is not documentation β it is a check that raises `ConfigError`. Each invariant
|
| 447 |
+
exists because a specific finding proved the failure mode.
|
| 448 |
+
|
| 449 |
+
| Invariant | Why it exists | Finding |
|
| 450 |
+
|---|---|---|
|
| 451 |
+
| `croma.image_resolution % 8 == 0` | CROMA asserts this; native 120 β 225 patches | C-7 |
|
| 452 |
+
| `training.precision β {fp16, bf16, fp32}` | the T4 is SM 7.5, so bf16 is unavailable | C-6 |
|
| 453 |
+
| `deployment.torch_compile is not true` | ZeroGPU does not support `torch.compile` | C-8 |
|
| 454 |
+
| `vlm.processor_longest_edge β€ image.tile_size` | the processor's default `longest_edge` is 2048, which upscales a 512 px tile 4Γ and then splits it into **17** sub-images β a ~17Γ overrun, not the 4Γ the plan estimated | F5-2 |
|
| 455 |
+
| `vlm.prompt_must_use_chat_template is true` | SmolVLM raises `ValueError` on prompts lacking one `<image>` token per image | F5-3 |
|
| 456 |
+
| `fusion.input_dim == 3*encoder_dim + optical_channels + sar_channels` (= 2318) | CROMA emits optical/SAR/joint GAP vectors; the availability mask is consumed by the head | C-1 |
|
| 457 |
+
| `croma.optical_channels == 12` and `croma.sar_channels == 2` | CROMA's `s2_channels` / `s1_channels` are fixed | β |
|
| 458 |
+
| `grounding_head.feature_dim == 4 * grounding.encoder_projected_dim` (= 2048) | a mismatch is a **silent** shape error β torch raises only at the similarity step, after patch features are already cached | P7-1 |
|
| 459 |
+
| `router.tasks` includes `unsupported` and `router.num_tasks == len(router.tasks)` | the ontology and its declared size cannot drift apart | β |
|
| 460 |
+
| `change.sa_mode β {BAM, PAM}` and `change.encoder` is set | the change architecture is not implicit | C-9 |
|
| 461 |
+
| `image.top_k_tiles β€ image.max_tiles` | the dispatch ceiling cannot exceed the examination ceiling | β |
|
| 462 |
+
|
| 463 |
+
(`core/config.py:94-222`; `release/repo/README.md` Β§Installation.)
|
| 464 |
+
|
| 465 |
+
> **Why the grounding-head guard matters most.** A mismatch there is a *silent* shape error: torch raises
|
| 466 |
+
> only at the similarity step, by which point the patch features have already been computed and cached β
|
| 467 |
+
> *"the failure surfaces far from its cause"* (`core/config.py:170-196`).
|
| 468 |
+
|
| 469 |
+
---
|
| 470 |
+
|
| 471 |
+
# Part IV β Contract-first development
|
| 472 |
+
|
| 473 |
+
## 11. `core/schemas.py` is the binding contract
|
| 474 |
+
|
| 475 |
+
`core/schemas.py` (462 lines) holds every model the system exchanges. The rule is simple: **no specialist
|
| 476 |
+
may invent its own result shape.** Every specialist returns exactly `SpecialistResult`
|
| 477 |
+
(`core/schemas.py:325-406`), whose docstring says so: *"Every specialist returns exactly this. No
|
| 478 |
+
exceptions."*
|
| 479 |
+
|
| 480 |
+
The models, in order:
|
| 481 |
+
|
| 482 |
+
| Model | Line | Role |
|
| 483 |
+
|---|---|---|
|
| 484 |
+
| `Task` (enum) | 35 | the six-task ontology (`vqa`, `caption`, `grounding`, `change`, `optical_sar`, `change_vqa`) + `unsupported` |
|
| 485 |
+
| `Modality` (enum) | 50 | `optical` / `sar` / `joint` |
|
| 486 |
+
| `CoordinateSystem` (enum) | 57 | `normalized_0_1` / `geo` |
|
| 487 |
+
| `EvidenceType` (enum) | 65 | the evidence classes |
|
| 488 |
+
| `ControllerState` (enum) | 79 | the nine-state spine |
|
| 489 |
+
| `Intent` | 94 | the router's output |
|
| 490 |
+
| `GeoMetadata` | 120 | CRS / transform metadata |
|
| 491 |
+
| `SensorDescriptor` | 138 | sensor identity |
|
| 492 |
+
| `AssetMetadata` | 153 | one input asset |
|
| 493 |
+
| `Box` | 171 | a flat box (see the flat-vs-nested trap below) |
|
| 494 |
+
| `Region` / `ChangeRegion` | 189 / 202 | spatial outputs |
|
| 495 |
+
| `Evidence` | 216 | one observable artefact |
|
| 496 |
+
| `ConfidenceBreakdown` | 259 | measurable confidence |
|
| 497 |
+
| `TraceStep` / `ModelRef` / `ExecutionTrace` | 279 / 288 / 296 | the trace |
|
| 498 |
+
| `SpecialistResult` | 325 | the master contract |
|
| 499 |
+
| `AnalysisRequest` | 412 | the request |
|
| 500 |
+
| `ResultEnvelope` | 421 | the response wrapper |
|
| 501 |
+
| `HealthStatus` | 430 | the health shape |
|
| 502 |
+
|
| 503 |
+
**Every model sets `model_config = ConfigDict(extra="forbid")`.** This is deliberate and has a documented
|
| 504 |
+
consequence (C-2, Β§5): the models reject a body carrying an unknown key, which contradicts
|
| 505 |
+
`API_CONTRACT.md` Β§1.1's forward-compatibility promise. The code is authoritative; the document is the
|
| 506 |
+
inaccurate half (`docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§15).
|
| 507 |
+
|
| 508 |
+
> **The flat-vs-nested trap.** `Box` is **flat**, not nested. An earlier documentation draft described it
|
| 509 |
+
> with nested geometry, and *"every box the frontend drew would have been at the origin"*
|
| 510 |
+
> (`docs/PHASE19_FINAL_HARDENING.md` Β§3.6). Documentation is validated against the real models by tests
|
| 511 |
+
> (`test_api_contract_doc.py`, `test_frontend_guide_doc.py`), which is what caught it.
|
| 512 |
+
|
| 513 |
+
### 11.1 The two cross-field validators you must not break
|
| 514 |
+
|
| 515 |
+
`SpecialistResult` carries a `_task_output_consistency` validator (`core/schemas.py:353-406`) with two
|
| 516 |
+
rules:
|
| 517 |
+
|
| 518 |
+
- **Grounding with no localisation is degraded, not a crash.** If `task == GROUNDING` and there are no
|
| 519 |
+
boxes or regions, the validator appends a warning and sets `degraded = True`.
|
| 520 |
+
- **Change-VQA with no answer text is degraded.** If `task == CHANGE_VQA` and the answer is blank, the
|
| 521 |
+
validator marks it degraded β but only if the specialist has not already set `degraded` itself.
|
| 522 |
+
|
| 523 |
+
The `CHANGE` clause that once lived there was **removed** (F-16c): it had collapsed to `not regions`, and
|
| 524 |
+
a *successful* no-change analysis began reporting `degraded: true`. **CHANGE is the one task whose
|
| 525 |
+
`degraded` flag is now set entirely by its specialist** (`core/schemas.py:362-395`).
|
| 526 |
+
|
| 527 |
+
## 12. The specialist interface
|
| 528 |
+
|
| 529 |
+
Every specialist implements exactly one abstract base class, `Specialist`, in `specialists/base.py`. The
|
| 530 |
+
interface is **four methods**, and the split is deliberate (`specialists/base.py:1-17`):
|
| 531 |
+
|
| 532 |
+
```
|
| 533 |
+
validate_request -> can this specialist serve this request at all?
|
| 534 |
+
execute -> do the work, return a SpecialistResult
|
| 535 |
+
produce_evidence -> what observable artefacts support the result?
|
| 536 |
+
estimate_confidence -> what measurable signals support the score?
|
| 537 |
+
```
|
| 538 |
+
|
| 539 |
+
The class attributes a specialist must set (`specialists/base.py:48-58`):
|
| 540 |
+
|
| 541 |
+
| Attribute | Meaning |
|
| 542 |
+
|---|---|
|
| 543 |
+
| `name` | stable machine-readable name, used in traces and evidence sources |
|
| 544 |
+
| `version` | semantic version; *"bump when behaviour changes, not when code moves"* |
|
| 545 |
+
| `capabilities` | the capability strings this specialist serves, e.g. `("vqa", "caption")` |
|
| 546 |
+
|
| 547 |
+
The four abstract methods (`specialists/base.py:62-96`):
|
| 548 |
+
|
| 549 |
+
| Method | Contract |
|
| 550 |
+
|---|---|
|
| 551 |
+
| `validate_request(request)` | raise the **most specific** typed error available (`InvalidRequestError`, `PairMisalignmentError`, β¦), never a bare `Exception` |
|
| 552 |
+
| `execute(request)` | return a normalised `SpecialistResult`; *"Never returns None."* |
|
| 553 |
+
| `produce_evidence(result)` | return `list[Evidence]`; *"Must not invent anything the specialist did not actually compute."* |
|
| 554 |
+
| `estimate_confidence(result)` | return a `ConfidenceBreakdown`; *"Never an LLM utterance."* |
|
| 555 |
+
|
| 556 |
+
The base class also provides helpers a specialist should use rather than re-implement
|
| 557 |
+
(`specialists/base.py:98-134`): `supports()`, `require_assets()` (asserts an exact asset count and raises
|
| 558 |
+
a typed error), `require_capability()`, `model_refs()` (for the trace), and `describe()`.
|
| 559 |
+
|
| 560 |
+
`SpecialistRequest` (`specialists/base.py:34-45`) is the input: `assets`, `query`, `params`, `run_id`.
|
| 561 |
+
Its docstring states the isolation rule: *"Everything a specialist is given. No specialist reads global
|
| 562 |
+
state."*
|
| 563 |
+
|
| 564 |
+
## 13. How to add a new specialist
|
| 565 |
+
|
| 566 |
+
This is the verified procedure, read from the code. There are **five** steps, and skipping any one fails
|
| 567 |
+
loudly (which is the design).
|
| 568 |
+
|
| 569 |
+
### Step 1 β implement the `Specialist` ABC
|
| 570 |
+
|
| 571 |
+
Subclass `specialists.base.Specialist`, set `name`, `version` and `capabilities`, and implement the four
|
| 572 |
+
methods. The builder for the class is a module-level function
|
| 573 |
+
`build_<x>_specialist(config, *, device, **kwargs)` β the same shape as the four existing builders
|
| 574 |
+
(`core/registry.py:185-259` names them: `build_vqa_specialist`, `build_grounding_specialist`,
|
| 575 |
+
`build_change_specialist`, `build_change_vqa_specialist`, `build_optical_sar_specialist`).
|
| 576 |
+
|
| 577 |
+
### Step 2 β add a row to the spec table
|
| 578 |
+
|
| 579 |
+
Add a `SpecialistSpec` to `default_specs()` in `core/registry.py:171-260`. The spec's fields
|
| 580 |
+
(`core/registry.py:121-165`):
|
| 581 |
+
|
| 582 |
+
| Field | Meaning |
|
| 583 |
+
|---|---|
|
| 584 |
+
| `name` | the registry key β **the capability string**, never the specialist's own `name` |
|
| 585 |
+
| `capabilities` | the tuple the built object must declare; **asserted** after construction |
|
| 586 |
+
| `module` | dotted module path containing the builder |
|
| 587 |
+
| `builder` | builder function name inside `module` |
|
| 588 |
+
| `requires_assets` | exact asset count, or `None` for "any" |
|
| 589 |
+
| `config_keys` | `{builder_kwarg: dotted.config.key}` β only keys that resolve are passed |
|
| 590 |
+
| `optional_config_keys` | as above, but a missing key contributes nothing |
|
| 591 |
+
| `failure_states` | registry state to use when construction raises, keyed by exception class name |
|
| 592 |
+
|
| 593 |
+
The `SpecialistSpec.__post_init__` refuses a spec whose `name` is not in its own `capabilities`
|
| 594 |
+
(`core/registry.py:158-165`): *"the registry keys on capability, so this spec would be unreachable."*
|
| 595 |
+
|
| 596 |
+
### Step 3 β understand the capability-vs-name asymmetry
|
| 597 |
+
|
| 598 |
+
**This is the trap the registry exists to encode** (`core/registry.py:30-45`). The VQA specialist declares:
|
| 599 |
+
|
| 600 |
+
```python
|
| 601 |
+
name = "vlm" # specialists/vqa/inference.py:114
|
| 602 |
+
capabilities = ("vqa", "caption") # specialists/vqa/inference.py:116
|
| 603 |
+
```
|
| 604 |
+
|
| 605 |
+
Every other specialist's `name` equals its single capability. A registry keyed on `name` would make `vqa`
|
| 606 |
+
permanently unfindable while every other specialist kept working. So the registry keys on **capability**
|
| 607 |
+
and **asserts** the capability tuple against the constructed object
|
| 608 |
+
(`core/registry.py:497-515`). If your spec's `capabilities` disagrees with what your specialist declares,
|
| 609 |
+
construction fails with a `SpecialistError` naming both tuples.
|
| 610 |
+
|
| 611 |
+
### Step 4 β register the task (only if it is a *new* task)
|
| 612 |
+
|
| 613 |
+
If the new specialist serves an existing task, nothing else is needed. If it is a new task, add it to:
|
| 614 |
+
|
| 615 |
+
- `Task` in `core/schemas.py:35-49`, and
|
| 616 |
+
- `TASK_CAPABILITY` in `core/planner.py:121-128`, and
|
| 617 |
+
- `CAPABILITY_ASSETS` in `core/planner.py:133-142`.
|
| 618 |
+
|
| 619 |
+
`CAPABILITY_ASSETS` mirrors each specialist's own `validate_request`, which **stays authoritative** β the
|
| 620 |
+
planner's copy is a *"cheap precondition so it can refuse before construction is attempted"*
|
| 621 |
+
(`core/planner.py:130-132`). A test asserts `CAPABILITY_ASSETS` equals `SpecialistSpec.requires_assets`
|
| 622 |
+
in both directions, and asserts the gateway carries **no third copy** of the asset-count table
|
| 623 |
+
(`docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§3).
|
| 624 |
+
|
| 625 |
+
> **Adding a task touches `router.tasks`.** `router.num_tasks` must equal `len(router.tasks)`
|
| 626 |
+
> (`core/config.py:199-206`), so a new task means a config edit β **which moves the hash** (Β§8.2). This is
|
| 627 |
+
> the one part of "add a specialist" that has a global consequence.
|
| 628 |
+
|
| 629 |
+
### Step 5 β return the right shapes
|
| 630 |
+
|
| 631 |
+
`execute` must return a `SpecialistResult`; `produce_evidence` must return `list[Evidence]`;
|
| 632 |
+
`estimate_confidence` must return a `ConfidenceBreakdown`. The registry will construct your specialist
|
| 633 |
+
lazily and record its state.
|
| 634 |
+
|
| 635 |
+
### What the registry does with your specialist
|
| 636 |
+
|
| 637 |
+
| Behaviour | Mechanism |
|
| 638 |
+
|---|---|
|
| 639 |
+
| lazy construction | the builder is imported via `importlib.import_module` **inside** `build()` β no module-level specialist import (`core/registry.py:15-25,462-475`) |
|
| 640 |
+
| memoisation | the second `build(cap)` returns the cached entry (`core/registry.py:432-434`) |
|
| 641 |
+
| three states | `AVAILABLE` / `DEGRADED` / `UNAVAILABLE` (`core/registry.py:102-107`) |
|
| 642 |
+
| degradation detection | duck-typed on `has_checkpoint` / `has_head` / `has_encoder` / `model is None` (`core/registry.py:527-548`) |
|
| 643 |
+
| failure is retained | a construction failure returns an `UNAVAILABLE` entry rather than raising; **only an unknown capability raises** (`core/registry.py:418-450`) |
|
| 644 |
+
| corrupt β missing | `ModelLoadError` / `ModelUnavailableError` map to `UNAVAILABLE`, never retried as `DEGRADED` (`core/registry.py:550-612`) |
|
| 645 |
+
| path scrubbing | the client-visible `detail` is `scrub_paths(...)`; the raw string goes to the log only (`core/registry.py:560-598`) |
|
| 646 |
+
|
| 647 |
+
> **The corrupt-vs-missing rule is load-bearing.** *"silently running an untrained model because a real
|
| 648 |
+
> checkpoint failed to load would be the worst outcome"* (`specialists/change/specialist.py:834-838`,
|
| 649 |
+
> quoted in `core/registry.py:63-75`). Do not add a fallback that re-adds a degradation the builder
|
| 650 |
+
> refused.
|
| 651 |
+
|
| 652 |
+
## 14. How evidence is emitted
|
| 653 |
+
|
| 654 |
+
Evidence is produced by `Specialist.produce_evidence(result)` and returned as `list[Evidence]`. The
|
| 655 |
+
`Evidence` model (`core/schemas.py:216-256`):
|
| 656 |
+
|
| 657 |
+
| Field | Type | Note |
|
| 658 |
+
|---|---|---|
|
| 659 |
+
| `evidence_id` | `str` | auto-generated `ev_<hex>`; **must be unique within a result** |
|
| 660 |
+
| `type` | `EvidenceType` | the evidence class |
|
| 661 |
+
| `source_specialist` | `str` | which specialist computed it |
|
| 662 |
+
| `coordinate_system` | `CoordinateSystem \| None` | required for spatial evidence |
|
| 663 |
+
| `coordinates` | `list[float] \| None` | the geometry |
|
| 664 |
+
| `score` | `float \| None` | `0.0 β€ score β€ 1.0` |
|
| 665 |
+
| `artifact_ref` | `str \| None` | **never a filesystem path** (F-16) |
|
| 666 |
+
| `payload` | `dict[str, Any]` | structured detail |
|
| 667 |
+
|
| 668 |
+
Two validators enforce correctness:
|
| 669 |
+
|
| 670 |
+
- `SpecialistResult._unique_evidence_ids` rejects a result with duplicate `evidence_id` values
|
| 671 |
+
(`core/schemas.py:345-351`).
|
| 672 |
+
- `Evidence._spatial_needs_crs` rejects spatial evidence (`BOUNDING_BOX`, `MASK`, `CHANGE_MAP`, `TILE`,
|
| 673 |
+
`IMAGE_CROP`, `JOINT_FEATURE_REGION`) that carries coordinates but **no** `coordinate_system`
|
| 674 |
+
(`core/schemas.py:241-255`).
|
| 675 |
+
|
| 676 |
+
> **`artifact_ref` is never a filesystem path.** v1 exposes no artifact-serving endpoint, so it is null
|
| 677 |
+
> unless a deployment supplies a client-fetchable reference. *"An artifact may still be written
|
| 678 |
+
> server-side where configured; being written is not the same as being retrievable."*
|
| 679 |
+
> (`core/schemas.py:227-238`.)
|
| 680 |
+
|
| 681 |
+
---
|
| 682 |
+
|
| 683 |
+
# Part V β The test workflow
|
| 684 |
+
|
| 685 |
+
## 15. Running tests
|
| 686 |
+
|
| 687 |
+
### 15.1 Always use the repository venv interpreter
|
| 688 |
+
|
| 689 |
+
pytest is installed **only** in the repository virtualenv. Invoking the system `pytest` fails or resolves
|
| 690 |
+
to a different interpreter (`release/repo/docs/REPRODUCIBILITY.md` Β§10.2):
|
| 691 |
+
|
| 692 |
+
```bash
|
| 693 |
+
.venv/Scripts/python.exe -m pytest tests/unit/test_frontend_live_wiring.py -q
|
| 694 |
+
```
|
| 695 |
+
|
| 696 |
+
On Windows, pytest must be given `-p no:cacheprovider` because the sandbox refuses `.pytest_cache` writes
|
| 697 |
+
(`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` Β§0).
|
| 698 |
+
|
| 699 |
+
### 15.2 Run targeted files, not the whole tree
|
| 700 |
+
|
| 701 |
+
**A full `tests/unit` run trips the sandbox's bulk-delete guard** (4Γ `test_safe_delete_shim` failures)
|
| 702 |
+
(`release/repo/docs/REPRODUCIBILITY.md` Β§10.3). The workaround is to run the targeted suites:
|
| 703 |
+
|
| 704 |
+
```bash
|
| 705 |
+
.venv/Scripts/python.exe -m pytest tests/unit/test_frontend_live_wiring.py -q # 106 passed
|
| 706 |
+
```
|
| 707 |
+
|
| 708 |
+
The runbook notes that **multi-suite invocations in one command have been refused by the environment
|
| 709 |
+
before; single suites are reliable** (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` Β§2.6):
|
| 710 |
+
|
| 711 |
+
```bash
|
| 712 |
+
$PY -m pytest tests/unit/test_deploy_config.py -p no:cacheprovider -q
|
| 713 |
+
$PY -m pytest tests/unit/test_app_serving.py -p no:cacheprovider -q
|
| 714 |
+
$PY -m pytest tests/unit/test_api_contract_doc.py tests/unit/test_frontend_guide_doc.py -p no:cacheprovider -q
|
| 715 |
+
```
|
| 716 |
+
|
| 717 |
+
> **Do not pipe pytest through `grep`** in the authoring sandbox β output is block-buffered and a killed
|
| 718 |
+
> pipeline swallows it. Redirect to a file instead (`release/repo/docs/REPRODUCIBILITY.md` Β§5.6).
|
| 719 |
+
|
| 720 |
+
### 15.3 `pytest.ini`
|
| 721 |
+
|
| 722 |
+
```ini
|
| 723 |
+
[pytest]
|
| 724 |
+
testpaths = tests
|
| 725 |
+
pythonpath = .
|
| 726 |
+
addopts = -q --tb=short
|
| 727 |
+
```
|
| 728 |
+
|
| 729 |
+
(`pytest.ini:1-4`)
|
| 730 |
+
|
| 731 |
+
`pythonpath = .` is what makes `from core.config import ...` resolve from the repo root without an
|
| 732 |
+
installed package.
|
| 733 |
+
|
| 734 |
+
## 16. The evidence-class markers
|
| 735 |
+
|
| 736 |
+
`pytest.ini` registers four markers whose purpose is to make a test's **evidence class** explicit
|
| 737 |
+
(`pytest.ini:23-27`):
|
| 738 |
+
|
| 739 |
+
| Marker | Meaning |
|
| 740 |
+
|---|---|
|
| 741 |
+
| `unit` | fast, no I/O, no server, no network. Evidence about a component in isolation. |
|
| 742 |
+
| `integration` | exercises two or more real components wired together in-process. |
|
| 743 |
+
| `smoke` | a minimal end-to-end path run against real local artifacts, not a mock. |
|
| 744 |
+
| `real_inference` | ran the real model on real inputs in this environment. |
|
| 745 |
+
|
| 746 |
+
The file's comment states the discipline:
|
| 747 |
+
|
| 748 |
+
> *"a unit test never claims a deployment was exercised, and no test may be reported as 'real inference'
|
| 749 |
+
> unless it ran the real model on real data in this environment."* (`pytest.ini:13-17`)
|
| 750 |
+
|
| 751 |
+
**`environment_blocked` is deliberately not a marker.** *"a blocked path is reported in the STEP 7 report
|
| 752 |
+
rather than encoded as a permanently-skipped test, because a skip can be mistaken for coverage"*
|
| 753 |
+
(`pytest.ini:19-21`). An unregistered mark is an **error** rather than a silent no-op (`pytest.ini:14-15`).
|
| 754 |
+
|
| 755 |
+
## 17. The known failures, and why they are not regressions
|
| 756 |
+
|
| 757 |
+
Running the entire unit tree trips **5β6** failures, classified by cause
|
| 758 |
+
(`release/repo/docs/REPRODUCIBILITY.md` Β§5.4):
|
| 759 |
+
|
| 760 |
+
| # | Failure | Cause | Regression? |
|
| 761 |
+
|---|---|---|---|
|
| 762 |
+
| 1β4 | `test_safe_delete_shim` (Γ4) | the sandbox's bulk-**delete guard** (Windows verbatim-path behaviour) | **No** |
|
| 763 |
+
| 5 | one ordering flake in the router route test | passes in isolation; order/collection-dependent | **No** |
|
| 764 |
+
| 6 | one stale adapter test | asserts `optical_sar` absent when CROMA is *unshipped* β **CROMA is now shipped** | **No** |
|
| 765 |
+
|
| 766 |
+
Re-running the affected files together passes **137** tests, which is what isolates them as environmental
|
| 767 |
+
rather than behavioural (`release/repo/docs/REPRODUCIBILITY.md` Β§5.4).
|
| 768 |
+
|
| 769 |
+
**Known per-suite results:**
|
| 770 |
+
|
| 771 |
+
| Suite | Command | Expected |
|
| 772 |
+
|---|---|---|
|
| 773 |
+
| Frontend live-wiring | `pytest tests/unit/test_frontend_live_wiring.py` | **106 passed** |
|
| 774 |
+
| Doc/frontend suite | `pytest` on the 5 doc/frontend files | **183 passed** |
|
| 775 |
+
| Full unit suite | `pytest tests/unit` | 5β6 **environmental** failures |
|
| 776 |
+
| Gateway policy | `pytest tests/unit/test_gateway_policy.py` | **51 passed** |
|
| 777 |
+
| Evidence engine | `pytest tests/unit/test_evidence_engine.py` | **73 passed** |
|
| 778 |
+
|
| 779 |
+
(`release/repo/README.md` Β§The test suites; `docs/PHASE19_FINAL_HARDENING.md` Β§6.)
|
| 780 |
+
|
| 781 |
+
> **The precise full-suite *collected* count is `UNKNOWN β not established from the available evidence`.**
|
| 782 |
+
> Per-suite counts are known; the single collected total is not (`release/repo/docs/REPRODUCIBILITY.md`
|
| 783 |
+
> Β§5.4).
|
| 784 |
+
|
| 785 |
+
---
|
| 786 |
+
|
| 787 |
+
# Part VI β Scripts and notebooks
|
| 788 |
+
|
| 789 |
+
## 18. The class of scripts under `scripts/`
|
| 790 |
+
|
| 791 |
+
`scripts/` is a **flat** directory of executable helpers β no sub-packages. The listing holds **51 script
|
| 792 |
+
files** (47 `.py`, 3 `.ps1`, 1 `.mjs`) plus `__init__.py`. They fall into clear classes:
|
| 793 |
+
|
| 794 |
+
| Class | Examples | Purpose |
|
| 795 |
+
|---|---|---|
|
| 796 |
+
| **Training entry points** | `train_router.py`, `train_grounding.py`, `train_change.py`, `train_fusion.py`, `train_change_vqa.py` | thin CLIs over the `training/` modules (Β§20) |
|
| 797 |
+
| **Evaluation** | `eval_change.py`, `eval_fusion_115.py`, `eval_grounding_head.py`, `evaluate_change_vqa.py` | per-task evaluation |
|
| 798 |
+
| **Data preparation** | `prepare_bigearthnet.py`, `prepare_change_vqa.py`, `select_bigearthnet_slice.py` | build corpora and selection manifests |
|
| 799 |
+
| **Contract probes** | `probe_grounding_head_contract.py`, `probe_remoteclip_contract.py`, `probe_vlm_contract.py` | measure a real model's contract before relying on it |
|
| 800 |
+
| **Smoke tests** | `smoke_test_change.py`, `smoke_test_grounding_head.py`, `smoke_test_vlm.py` | a minimal real path per specialist |
|
| 801 |
+
| **Verification / gates** | `verify_gate1.py`, `verify_croma_forward.py`, `verify_grounding_e2e.py`, `verify_levir_real.py`, `verify_cdvqa_imagery.py` | prove a property end to end |
|
| 802 |
+
| **Threshold / hyperparameter sweeps** | `sweep_change_threshold.py`, `sweep_router_threshold.py`, `fusion_seed_variance.py` | sweep a frozen knob |
|
| 803 |
+
| **Phase-6 (VLM) tooling** | `phase6_train_vlm.py`, `phase6_adjudicate_test.py`, `phase6_close.py`, `phase6_rerule.py`, `phase6_recover_baseline_test.py` | the VLM acceptance workflow |
|
| 804 |
+
| **Phase-12 (fusion) tooling** | `p12_integrity_verify.py`, `p12_preflight_verify.py`, `run_phase12_extraction.ps1`, `run_phase12_resume_ref.ps1`, `run_phase12_resume_s1_ref.ps1` | the fusion feature-extraction workflow (PowerShell drivers) |
|
| 805 |
+
| **Calibration** | `fit_calibration.py` | fit the temperature-scaling artifact on validation (Β§22) |
|
| 806 |
+
| **Kaggle packaging / rehearsal** | `package_kaggle_code.py`, `rehearse_kaggle_notebook.py`, `rehearse_change_notebook.py` | package and dry-run the notebooks |
|
| 807 |
+
| **Deployment validation** | `validate_deploy_config.py` | prove `configs/deploy.yaml` is inert |
|
| 808 |
+
| **Frontend staging** | `stage_pages.mjs` | build the Cloudflare Pages bundle |
|
| 809 |
+
| **Environment / diagnostics** | `check_env.py`, `diagnose_feature_cache.py`, `croma_normalisation_arm_probe.py`, `analyze_grounding_resolution.py`, `analyze_reben_labels.py`, `check_cdvqa_second_overlap.py`, `establish_cdvqa_temporal_order.py`, `exp_grounding_resolution.py`, `extract_fusion_features.py` | environment and dataset diagnostics |
|
| 810 |
+
|
| 811 |
+
**The training entry points share one convention**, stated in their docstrings: *"A THIN CLI over
|
| 812 |
+
`training/<task>/train.py` β¦ Everything that computes lives in the module; this file parses arguments,
|
| 813 |
+
reports the environment, prints the accounting, and returns an exit code."* (`scripts/train_change.py:1-6`,
|
| 814 |
+
`scripts/train_fusion.py:1-6`). The exit codes are a contract (`scripts/train_change.py:17-24`):
|
| 815 |
+
|
| 816 |
+
```
|
| 817 |
+
0 training completed (or --dry-run validated a present dataset)
|
| 818 |
+
2 the dataset is missing, empty, or cannot be split -- NOT a crash
|
| 819 |
+
3 the run started and the trainer raised a typed error
|
| 820 |
+
```
|
| 821 |
+
|
| 822 |
+
> **Re-exported names are part of the contract.** Several scripts re-export their trainer's symbols at
|
| 823 |
+
> module scope, and a test asserts it. E.g. `tests/unit/test_change_train_script_contract.py` asserts
|
| 824 |
+
> `change_loss`, `train_change_head` and `evaluate` are reachable as `scripts.train_change.*`
|
| 825 |
+
> (`scripts/train_change.py:26-30`).
|
| 826 |
+
|
| 827 |
+
## 19. `notebooks/`
|
| 828 |
+
|
| 829 |
+
| Notebook | Purpose |
|
| 830 |
+
|---|---|
|
| 831 |
+
| `kaggle_change_vqa_train.ipynb` | R-02 change-VQA training (Β§21) |
|
| 832 |
+
| `kaggle_phase6_vlm_lora.ipynb` | Phase-6 SmolVLM LoRA training (Β§21) |
|
| 833 |
+
| `kaggle_change_training.ipynb` | the change **detector** β *"frozen and out of scope"* (`docs/R02_KAGGLE_TRAINING_GUIDE.md` Β§4) |
|
| 834 |
+
| `kaggle_grounding_resolution.ipynb` | Phase-7 grounding resolution β unrelated to R-02 |
|
| 835 |
+
|
| 836 |
+
> **Do not confuse the two change notebooks.** For change-VQA use `kaggle_change_vqa_train.ipynb`; do
|
| 837 |
+
> **not** use `kaggle_change_training.ipynb` (that one trains the change *detector*) or
|
| 838 |
+
> `kaggle_grounding_resolution.ipynb` (Phase 7) (`docs/R02_KAGGLE_TRAINING_GUIDE.md` Β§4).
|
| 839 |
+
|
| 840 |
+
Two `.bak-*` files sit beside the Phase-6 notebook; they are editor backups, not runnable notebooks.
|
| 841 |
+
|
| 842 |
+
---
|
| 843 |
+
|
| 844 |
+
# Part VII β Training entry points
|
| 845 |
+
|
| 846 |
+
The six artifacts train in two places: four locally, two on an external GPU. This is a **deliberate
|
| 847 |
+
boundary** β the release ships *"frozen artifacts with provenance, not a retraining harness"*
|
| 848 |
+
(`release/repo/docs/REPRODUCIBILITY.md` Β§8.5).
|
| 849 |
+
|
| 850 |
+
| Artifact | Where it trains | Reproducible from this release? |
|
| 851 |
+
|---|---|---|
|
| 852 |
+
| router adapter | local CPU | **yes** β `configs/base.yaml` Β§`router.training` |
|
| 853 |
+
| grounding head | local | **yes** β `configs/base.yaml` Β§`grounding_training` |
|
| 854 |
+
| change head | local | **yes** β `configs/base.yaml` Β§`change` |
|
| 855 |
+
| optical_sar fusion head | local, seed sweep | **yes** |
|
| 856 |
+
| change_vqa head | **external GPU (Kaggle)** | **partly** β the promotion gate, evaluation and serving wiring are reproducible; there is no one-command retrain |
|
| 857 |
+
| vlm LoRA adapter | **external GPU** | **partly** β same |
|
| 858 |
+
|
| 859 |
+
(`release/repo/README.md` Β§What "reproduce" means.)
|
| 860 |
+
|
| 861 |
+
## 20. Local training (router, grounding, change, optical-SAR)
|
| 862 |
+
|
| 863 |
+
### 20.1 Router β CPU-only, and fast
|
| 864 |
+
|
| 865 |
+
```bash
|
| 866 |
+
python scripts/train_router.py --smoke # fast, stub encoder
|
| 867 |
+
python scripts/train_router.py --stub # fast, no model download
|
| 868 |
+
python scripts/train_router.py # full run, real MiniLM
|
| 869 |
+
```
|
| 870 |
+
|
| 871 |
+
(`scripts/train_router.py:1-10`)
|
| 872 |
+
|
| 873 |
+
The script's own measured note: *"Runs entirely on CPU. Measured cost with the real encoder: ~40 s to load
|
| 874 |
+
MiniLM on a cold cache, ~0.1 s to embed the corpus, ~1 s to train 60 epochs. There is no reason to spend
|
| 875 |
+
Kaggle GPU quota on this."* The encoder is frozen, so embeddings are cached and the adapter trains on
|
| 876 |
+
cached vectors (`configs/base.yaml:67-69`).
|
| 877 |
+
|
| 878 |
+
### 20.2 Grounding β two separable stages
|
| 879 |
+
|
| 880 |
+
```bash
|
| 881 |
+
# stage 1 only β measure the corpus before committing to training
|
| 882 |
+
python scripts/train_grounding.py --data-root <root> --extract-only
|
| 883 |
+
|
| 884 |
+
# both stages, real run
|
| 885 |
+
python scripts/train_grounding.py --data-root <root> --checkpoint <ckpt.pt>
|
| 886 |
+
|
| 887 |
+
# 2-batch forward+backward+checkpoint+reload, on CPU
|
| 888 |
+
python scripts/train_grounding.py --data-root <root> --debug
|
| 889 |
+
```
|
| 890 |
+
|
| 891 |
+
(`scripts/train_grounding.py:6-22`)
|
| 892 |
+
|
| 893 |
+
Splitting extraction from training means a hyperparameter change does not re-encode 15,699 images, and a
|
| 894 |
+
crashed training run does not lose the cache. The stated bar: *"the Phase 7 zero-shot baseline scored
|
| 895 |
+
0.0972 mean best IoU over 16,159 real eval records. This head must beat it by MIN_IMPROVEMENT_IOU to
|
| 896 |
+
justify existing. The comparison is printed whether or not it passes."*
|
| 897 |
+
|
| 898 |
+
### 20.3 Change β a thin CLI over `training/change/train.py`
|
| 899 |
+
|
| 900 |
+
```bash
|
| 901 |
+
# validate the data and the split; build nothing, train nothing
|
| 902 |
+
python scripts/train_change.py --data-root <LEVIR-CD root> --dry-run
|
| 903 |
+
|
| 904 |
+
# real run on CPU
|
| 905 |
+
python scripts/train_change.py --data-root <LEVIR-CD root> --device cpu
|
| 906 |
+
|
| 907 |
+
# a first real-data run that does not commit an hour
|
| 908 |
+
python scripts/train_change.py --data-root <root> --limit 256 --epochs 3
|
| 909 |
+
```
|
| 910 |
+
|
| 911 |
+
(`scripts/train_change.py:8-16`)
|
| 912 |
+
|
| 913 |
+
`--dry-run` is *"the leakage check without the cost: it loads the items, performs the scene-disjoint
|
| 914 |
+
split, runs `assert_image_disjoint`, prints the accounting and exits. It does NOT build the detector, so
|
| 915 |
+
it cannot trigger a pretrained-weights download, and it writes no files."*
|
| 916 |
+
|
| 917 |
+
### 20.4 Optical-SAR fusion β CPU-only, no result claimed
|
| 918 |
+
|
| 919 |
+
```bash
|
| 920 |
+
# validate the caches and the split; train nothing, write nothing
|
| 921 |
+
python scripts/train_fusion.py --train-cache train.npz --val-cache val.npz --dry-run
|
| 922 |
+
|
| 923 |
+
# a real run on CPU (the fusion head is CPU-only)
|
| 924 |
+
python scripts/train_fusion.py --train-cache train.npz --val-cache val.npz --arm A --epochs 20
|
| 925 |
+
```
|
| 926 |
+
|
| 927 |
+
(`scripts/train_fusion.py:8-15`)
|
| 928 |
+
|
| 929 |
+
> **No performance number is a result here.** The script states it: *"This loop has never seen the real
|
| 930 |
+
> paired BigEarthNet-S1+S2 corpus. Every run record it writes carries `result_status` and
|
| 931 |
+
> `pre_registered_metric_computed: false`; the pre-registered 11.5 metric is not computed."*
|
| 932 |
+
> (`scripts/train_fusion.py:17-20`.)
|
| 933 |
+
|
| 934 |
+
## 21. External GPU training (change-VQA, VLM LoRA)
|
| 935 |
+
|
| 936 |
+
Both external runs happen on **Kaggle** with **GPU T4 Γ2**. Neither is a one-command retrain from this
|
| 937 |
+
release.
|
| 938 |
+
|
| 939 |
+
### 21.1 Change-VQA (R-02)
|
| 940 |
+
|
| 941 |
+
**Guide:** `docs/R02_KAGGLE_TRAINING_GUIDE.md` (37 KB). **Runbook:** `RUNBOOK_CHANGE_VQA_KAGGLE.md`.
|
| 942 |
+
|
| 943 |
+
| Property | Value | Source |
|
| 944 |
+
|---|---|---|
|
| 945 |
+
| notebook | `notebooks/kaggle_change_vqa_train.ipynb` | `docs/R02_KAGGLE_TRAINING_GUIDE.md` Β§4 |
|
| 946 |
+
| accelerator | **GPU T4 Γ2** | Β§5 |
|
| 947 |
+
| internet | **On** (MiniLM downloads) | Β§5 |
|
| 948 |
+
| cells | **all 34, top to bottom** | Β§7 |
|
| 949 |
+
| hard stop | `HARD_STOP_SECONDS = 3 * 3600` (the plan's 3 h) | Β§7 |
|
| 950 |
+
| trainer defaults | `epochs=40, batch_size=128, seed=42, patience=6, time_limit=10800s` | Β§7 |
|
| 951 |
+
|
| 952 |
+
The path-discovery cells find the code root (four markers) and the CDVQA root (12 annotations + 4 image
|
| 953 |
+
dirs), and section 3c verifies the frozen STANet checkpoint by size and SHA256
|
| 954 |
+
(`RUNBOOK_CHANGE_VQA_KAGGLE.md` Β§5). **Test splits are not readable by the trainer:**
|
| 955 |
+
`training/change_vqa/train.py` loads only Train and Val, and *"there is no option in either file that
|
| 956 |
+
changes that"* (`scripts/train_change_vqa.py:11-15`).
|
| 957 |
+
|
| 958 |
+
### 21.2 VLM LoRA (Phase 6)
|
| 959 |
+
|
| 960 |
+
**Runbook:** `RUNBOOK_PHASE6_VLM_KAGGLE.md`.
|
| 961 |
+
|
| 962 |
+
| Property | Value | Source |
|
| 963 |
+
|---|---|---|
|
| 964 |
+
| notebook | `notebooks/kaggle_phase6_vlm_lora.ipynb` | `RUNBOOK_PHASE6_VLM_KAGGLE.md` Β§1 |
|
| 965 |
+
| accelerator | **GPU T4 Γ2** | Β§4 |
|
| 966 |
+
| internet | **off** (base model attached as input) | Β§4 |
|
| 967 |
+
| precision | **`fp16`, not the plan's `bf16`** β T4 is SM 7.5 | Β§4 |
|
| 968 |
+
| base model | `HuggingFaceTB/SmolVLM-500M-Instruct`, ~1.02 GB `model.safetensors` | Β§4 |
|
| 969 |
+
|
| 970 |
+
The trainer raises a `ConfigError` on `bf16` rather than silently falling back, and the deviation is
|
| 971 |
+
recorded in the manifest under `plan_deviations` (`RUNBOOK_PHASE6_VLM_KAGGLE.md` Β§4).
|
| 972 |
+
|
| 973 |
+
> **The VLM adapter's status is `ACCEPTANCE-REJECTED`.** Its metrics are *usable* (`exact_match 0.963`)
|
| 974 |
+
> but it was not promoted. **USABLE β ACCEPTED** (`DOCS_STYLE_GUIDE.md` Β§3). Do not describe the VLM path
|
| 975 |
+
> as accepted.
|
| 976 |
+
|
| 977 |
+
## 22. Calibration
|
| 978 |
+
|
| 979 |
+
```bash
|
| 980 |
+
python scripts/fit_calibration.py --dry-run # check inputs, exit
|
| 981 |
+
python scripts/fit_calibration.py # fit, write artifact
|
| 982 |
+
```
|
| 983 |
+
|
| 984 |
+
(`scripts/fit_calibration.py:19-21`)
|
| 985 |
+
|
| 986 |
+
It fits on **validation** data, and both the fitter entry point and the artifact writer **refuse a
|
| 987 |
+
held-out split** β *"Fitting on Test or Test2 would make the reported confidence a function of the answers
|
| 988 |
+
it is used to score β a leak, not a calibration."* (`scripts/fit_calibration.py:7-13`.)
|
| 989 |
+
|
| 990 |
+
> **Calibration made ECE worse** β `0.013755 β 0.014929` β and is **retained only because it is in the
|
| 991 |
+
> frozen config**. Never present it as an improvement (`DOCS_STYLE_GUIDE.md` Β§3).
|
| 992 |
+
|
| 993 |
+
---
|
| 994 |
+
|
| 995 |
+
# Part VIII β Coding conventions
|
| 996 |
+
|
| 997 |
+
## 23. What the code actually does
|
| 998 |
+
|
| 999 |
+
These conventions are visible across the files read. They are not aspirational style rules; each is
|
| 1000 |
+
observable in the code.
|
| 1001 |
+
|
| 1002 |
+
### 23.1 Every module has a substantial docstring stating *why*
|
| 1003 |
+
|
| 1004 |
+
The files read are densely commented at the module and function level, and the comments explain decisions
|
| 1005 |
+
and failure modes rather than restating the code. Examples: `core/registry.py:1-76` (a 76-line module
|
| 1006 |
+
docstring on the capability-key trap), `deploy/codespace/launch.sh:1-24` (why the launcher is defensive),
|
| 1007 |
+
`gateway/app.py:53-106` (the annotation-scope trap). **A change that removes the reasoning from a
|
| 1008 |
+
comment removes the reason a future reader will not re-introduce the bug.**
|
| 1009 |
+
|
| 1010 |
+
### 23.2 `from __future__ import annotations` is standard
|
| 1011 |
+
|
| 1012 |
+
It appears at the top of nearly every module (`core/config.py:11`, `core/registry.py:78`,
|
| 1013 |
+
`specialists/base.py:19`, `gateway/app.py:48`, `deploy/render/main.py:40`, `deploy/render/codespaces.py:16`,
|
| 1014 |
+
`deploy/codespace/warm_cache.py:28`, `scripts/*.py`). It is convenient β and it is the direct cause of
|
| 1015 |
+
Trap 6 (Β§30).
|
| 1016 |
+
|
| 1017 |
+
### 23.3 Findings are recorded as short codes, inline
|
| 1018 |
+
|
| 1019 |
+
The code refers to findings by code (`C-1`, `C-6`, `C-8`, `F5-2`, `P7-1`, `F-15`, `F-16`, `F-17`) and
|
| 1020 |
+
states the failure each guard prevents. This is a documentation convention enforced by comments, e.g.
|
| 1021 |
+
`core/config.py:121-142` (finding F5-2, the 17Γ overrun), `core/registry.py:227-235` (F-17, a config
|
| 1022 |
+
surface removed rather than left as a silent no-op).
|
| 1023 |
+
|
| 1024 |
+
### 23.4 Pydantic models set `extra="forbid"` and carry validators
|
| 1025 |
+
|
| 1026 |
+
Every schema model forbids extra fields and uses `@field_validator` / `@model_validator` for cross-field
|
| 1027 |
+
rules (Β§11). New models should follow the same pattern.
|
| 1028 |
+
|
| 1029 |
+
### 23.5 Loggers are module-level and named `satquery.<area>`
|
| 1030 |
+
|
| 1031 |
+
`logging.getLogger("satquery.orchestrator")` (`deploy/render/main.py:74`),
|
| 1032 |
+
`logging.getLogger("satquery.registry")` (`core/registry.py:95`), `logging.getLogger(__name__)`
|
| 1033 |
+
(`gateway/app.py:123`). Log calls use `%s`-style free text β there are no structured/JSON logs
|
| 1034 |
+
(`docs/architecture/10-observability-and-ops.md` Β§4.4).
|
| 1035 |
+
|
| 1036 |
+
### 23.6 Client-visible strings are scrubbed
|
| 1037 |
+
|
| 1038 |
+
Server-side diagnostics that name absolute paths must not reach a client. `core/errors.scrub_paths` reduces
|
| 1039 |
+
a path to a basename before it is published (`core/registry.py:297-312,560-598`;
|
| 1040 |
+
`core/errors.py:39`). Any new client-visible field that could carry a path or an exception string must go
|
| 1041 |
+
through the same scrub.
|
| 1042 |
+
|
| 1043 |
+
### 23.7 `noqa` comments state the reason
|
| 1044 |
+
|
| 1045 |
+
Where a lint suppression is used, the reason is written, e.g. `# noqa: E402 (see comment above)` in
|
| 1046 |
+
`gateway/app.py:85-106`, and `# noqa: BLE001 - report, never crash the warm step` in
|
| 1047 |
+
`deploy/codespace/warm_cache.py:74`.
|
| 1048 |
+
|
| 1049 |
+
### 23.8 Tests assert the *cause*, not the symptom
|
| 1050 |
+
|
| 1051 |
+
The regression test for the annotation-scope trap asserts that the route has no query parameters and that
|
| 1052 |
+
`gateway.app.Request is starlette.requests.Request` β *"rather than the symptom, because asserting the
|
| 1053 |
+
symptom would be brittle"* (`docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§13). Follow this when writing a
|
| 1054 |
+
regression test.
|
| 1055 |
+
|
| 1056 |
+
### 23.9 The 88-column soft limit
|
| 1057 |
+
|
| 1058 |
+
The files read wrap around 88 characters. There is no committed linter config in the files read, so this
|
| 1059 |
+
is a convention, not an enforced rule β **the exact formatter/linter configuration is `UNKNOWN β not
|
| 1060 |
+
established from the available evidence`.**
|
| 1061 |
+
|
| 1062 |
+
---
|
| 1063 |
+
|
| 1064 |
+
# Part IX β Known development traps
|
| 1065 |
+
|
| 1066 |
+
## 24. The traps, in one table
|
| 1067 |
+
|
| 1068 |
+
| # | Trap | One-line consequence |
|
| 1069 |
+
|---|---|---|
|
| 1070 |
+
| 1 | the monorepo `deploy/` is stale and untracked | it is **not** the deployed source |
|
| 1071 |
+
| 2 | the sandbox proxy is dead | outbound calls need `--noproxy '*'` / `ProxyHandler({})` |
|
| 1072 |
+
| 3 | pytest exists only in the venv | the system `pytest` resolves to the wrong interpreter |
|
| 1073 |
+
| 4 | Chrome drops synthetic CDP key events without OS focus | a browser-driven run silently answers the default query |
|
| 1074 |
+
| 5 | Cloudflare `_headers` rules concatenate | a later rule cannot "fix" an earlier one |
|
| 1075 |
+
| 6 | `from __future__ import annotations` + FastAPI | an unresolvable `Request` annotation becomes a **required query parameter** |
|
| 1076 |
+
| 7 | the full-suite run trips the bulk-delete guard | 4 spurious `test_safe_delete_shim` failures |
|
| 1077 |
+
| 8 | a stale serve process keeps answering | it reports the **previous revision's** capabilities |
|
| 1078 |
+
|
| 1079 |
+
## 25. Trap 1 β the stale untracked `deploy/`
|
| 1080 |
+
|
| 1081 |
+
**Symptom.** You edit `deploy/render/main.py`, deploy your change, and nothing changes in production β or
|
| 1082 |
+
you read `deploy/render/main.py` and cannot find the tunnel code the live service runs.
|
| 1083 |
+
|
| 1084 |
+
**Root cause.** `deploy/` inside the monorepo is **stale and untracked**. `git status` reports
|
| 1085 |
+
`?? deploy/` (verified in the working copy). It is **not** the deployed source
|
| 1086 |
+
(`release/repo/docs/DEPLOYMENT.md` Β§1; `release/CURRENT_RELEASE_STATE.md` Β§6).
|
| 1087 |
+
|
| 1088 |
+
**Evidence of divergence.** The monorepo `deploy/render/main.py` (532 lines) exposes `/api/health` with a
|
| 1089 |
+
`config` block that has **no** `tunnel` field and no `transport_mode` / `tunnel_timeout_s` /
|
| 1090 |
+
`wake_timeout_s` keys (`deploy/render/main.py:444-466`), whereas the **live** payload carries all of them
|
| 1091 |
+
(`release/repo/docs/DEPLOYMENT.md` Β§5). The monorepo copy also lacks `tunnel_agent.py` and `doctor.sh`,
|
| 1092 |
+
both of which `launch.sh` references (`deploy/codespace/launch.sh:83,89,153,160-168,179`).
|
| 1093 |
+
|
| 1094 |
+
**Fix.** Fetch the deployed file from the private repository and diff before editing. Treat the monorepo
|
| 1095 |
+
`deploy/` as documentation of intent, not as source.
|
| 1096 |
+
|
| 1097 |
+
## 26. Trap 2 β the dead sandbox proxy needs `--noproxy '*'`
|
| 1098 |
+
|
| 1099 |
+
**Symptom.** Outbound HTTP calls fail or hang in the authoring sandbox.
|
| 1100 |
+
|
| 1101 |
+
**Root cause.** The sandbox proxy is dead; requests are routed to it and never reach the target.
|
| 1102 |
+
|
| 1103 |
+
**Fix.** Disable proxies for the call (`release/repo/docs/REPRODUCIBILITY.md` Β§10.1):
|
| 1104 |
+
|
| 1105 |
+
```bash
|
| 1106 |
+
curl --noproxy '*' https://satquery-backend-m4yv.onrender.com/api/health
|
| 1107 |
+
```
|
| 1108 |
+
|
| 1109 |
+
```python
|
| 1110 |
+
opener = urllib.request.build_opener(urllib.request.ProxyHandler({}))
|
| 1111 |
+
```
|
| 1112 |
+
|
| 1113 |
+
`release/tools/hf_verify.py:37-38` does exactly this (`opener_no_proxy()`). In a normal environment the
|
| 1114 |
+
flag is harmless; in the sandbox it is mandatory.
|
| 1115 |
+
|
| 1116 |
+
## 27. Trap 3 β pytest only in the venv
|
| 1117 |
+
|
| 1118 |
+
**Symptom.** Invoking the system `pytest` fails or resolves to a different interpreter.
|
| 1119 |
+
|
| 1120 |
+
**Root cause.** pytest is installed only in `.venv`.
|
| 1121 |
+
|
| 1122 |
+
**Fix.** Always invoke the venv interpreter explicitly (`release/repo/docs/REPRODUCIBILITY.md` Β§10.2):
|
| 1123 |
+
|
| 1124 |
+
```bash
|
| 1125 |
+
.venv/Scripts/python.exe -m pytest tests/unit/test_frontend_live_wiring.py -q
|
| 1126 |
+
```
|
| 1127 |
+
|
| 1128 |
+
## 28. Trap 4 β Chrome drops synthetic CDP key events
|
| 1129 |
+
|
| 1130 |
+
**Symptom.** A browser-driven run silently answers the *default* query; the query box looks untouched;
|
| 1131 |
+
`mock_nodes` is non-zero.
|
| 1132 |
+
|
| 1133 |
+
**Root cause.** Chrome **drops synthesized key events when the browser window does not hold OS focus**.
|
| 1134 |
+
`press_key` / `fill_input` (real CDP key events) are focus-gated; `Input.insertText` (`type_text`) is not.
|
| 1135 |
+
|
| 1136 |
+
**Measured.** With Chrome backgrounded, `press_key("Z")` left `#qtext.value` unchanged, while
|
| 1137 |
+
`type_text("Q")` inserted fine (`release/repo/docs/REPRODUCIBILITY.md` Β§6.4).
|
| 1138 |
+
|
| 1139 |
+
**Fix.** Use `type_text` (not `fill_input`), and **assert the input state before dispatch** β `q_ok`
|
| 1140 |
+
(the query box really held the query), `obs_ok` (`#obsTail == 'ready'`), `t0_ok` (both frames present
|
| 1141 |
+
where required). This is the single most dangerous trap because it produces a **silent false pass** β the
|
| 1142 |
+
pipeline "works", it just answered a different question
|
| 1143 |
+
(`release/repo/docs/REPRODUCIBILITY.md` Β§10.5).
|
| 1144 |
+
|
| 1145 |
+
## 29. Trap 5 β Cloudflare `_headers` concatenate
|
| 1146 |
+
|
| 1147 |
+
**Symptom.** A specific cache-control rule does not take effect; the browser caches a file you expected it
|
| 1148 |
+
to revalidate.
|
| 1149 |
+
|
| 1150 |
+
**Root cause.** Cloudflare `_headers` rules **concatenate, they do not override.** Two matching rules are
|
| 1151 |
+
merged: a specific rule nested under a broad `/assets/img/*` rule yields
|
| 1152 |
+
`max-age=604800, β¦, max-age=0, must-revalidate` β and **Chromium takes the FIRST `max-age`**. The file's
|
| 1153 |
+
own "later rules override" comment is **false** (`release/repo/docs/DEPLOYMENT.md` Β§10).
|
| 1154 |
+
|
| 1155 |
+
**Fix.** Order rules so the *broadest* rule appears last, and never rely on a later rule overriding an
|
| 1156 |
+
earlier one. Related: Cloudflare **308-redirects `X.html` β `/X`**, so reference the extensionless path
|
| 1157 |
+
(`release/repo/docs/REPRODUCIBILITY.md` Β§10.4).
|
| 1158 |
+
|
| 1159 |
+
## 30. Trap 6 β the annotation-scope trap
|
| 1160 |
+
|
| 1161 |
+
**Symptom.** Every `POST` route returns `422` with FastAPI's own shape, without ever entering the handler:
|
| 1162 |
+
|
| 1163 |
+
```
|
| 1164 |
+
POST /v1/analyze β 422
|
| 1165 |
+
{"detail":[{"type":"missing","loc":["query","request"],"msg":"Field required"}]}
|
| 1166 |
+
```
|
| 1167 |
+
|
| 1168 |
+
(`docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§13.)
|
| 1169 |
+
|
| 1170 |
+
**Root cause.** The module uses `from __future__ import annotations`, so every annotation is a **string**
|
| 1171 |
+
at runtime. FastAPI resolves those strings via `get_typed_signature`, which calls
|
| 1172 |
+
`eval(annotation, func.__globals__)`. If `Request` is imported **inside** `create_app`, the route closures
|
| 1173 |
+
capture the name as a *local* of `create_app`; it never appears in `gateway.app.__globals__`, so
|
| 1174 |
+
resolution fails and FastAPI is left holding a bare `ForwardRef('Request')`. **FastAPI does not raise** β
|
| 1175 |
+
it silently falls back to treating the parameter as a **required query parameter named `request`**
|
| 1176 |
+
(`gateway/app.py:53-73`).
|
| 1177 |
+
|
| 1178 |
+
Three things were wrong at once: the request body was never read, the gateway's own validation never ran,
|
| 1179 |
+
and the error shape was FastAPI's `{"detail": ...}` rather than the contract's `{"error": {...}}`. **Every**
|
| 1180 |
+
POST route was affected, including `/v1/assets`.
|
| 1181 |
+
|
| 1182 |
+
**The asymmetry to internalise.** An unresolvable **parameter** annotation is *silently reinterpreted*
|
| 1183 |
+
(the handler never runs), while an unresolvable **return** annotation *raises*
|
| 1184 |
+
(`pydantic.errors.PydanticUndefinedAnnotation: name 'JSONResponse' is not defined`, which made
|
| 1185 |
+
`create_app()` unbuildable). Same root cause, opposite diagnosability
|
| 1186 |
+
(`gateway/app.py:88-106`).
|
| 1187 |
+
|
| 1188 |
+
**Fix.** Bind the annotation subjects at **module scope** (`gateway/app.py:79-106`):
|
| 1189 |
+
|
| 1190 |
+
```python
|
| 1191 |
+
if TYPE_CHECKING: # pragma: no cover
|
| 1192 |
+
from starlette.requests import Request
|
| 1193 |
+
from starlette.responses import Response
|
| 1194 |
+
|
| 1195 |
+
#: Runtime bindings used as annotation subjects in this module. Deliberately
|
| 1196 |
+
#: module-level so `eval()` can find them. See the comment above.
|
| 1197 |
+
from starlette.requests import Request # noqa: E402 (see comment above)
|
| 1198 |
+
from starlette.responses import Response # noqa: E402 (see comment above)
|
| 1199 |
+
from fastapi.responses import JSONResponse # noqa: E402 (see comment above)
|
| 1200 |
+
```
|
| 1201 |
+
|
| 1202 |
+
The regression test asserts the **cause**: the route has no query parameters, and
|
| 1203 |
+
`gateway.app.Request is starlette.requests.Request`.
|
| 1204 |
+
|
| 1205 |
+
> **The general rule.** Any name used as a FastAPI route annotation in a module with
|
| 1206 |
+
> `from __future__ import annotations` **must** be importable from that module's globals. Import it at
|
| 1207 |
+
> module scope, not inside the factory.
|
| 1208 |
+
|
| 1209 |
+
## 31. Trap 7 β the full-suite bulk-delete guard
|
| 1210 |
+
|
| 1211 |
+
**Symptom.** Running the whole `tests/unit` tree trips 4Γ `test_safe_delete_shim` failures.
|
| 1212 |
+
|
| 1213 |
+
**Root cause.** Windows **verbatim-path** defects in the sandbox's bulk-delete guard. The precise
|
| 1214 |
+
condition under which the shim intermittently triggers on Windows is `UNKNOWN β not established from the
|
| 1215 |
+
available evidence` (`release/repo/docs/REPRODUCIBILITY.md` Β§10.3).
|
| 1216 |
+
|
| 1217 |
+
**Fix / workaround.** Run the targeted suites (106 and 183 pass cleanly); treat the 4 shim failures as
|
| 1218 |
+
environmental, not regressions (Β§17).
|
| 1219 |
+
|
| 1220 |
+
## 32. Trap 8 β the stale serve process
|
| 1221 |
+
|
| 1222 |
+
**Symptom.** The Codespace answers `/v1/health` and `/v1/capabilities`, but reports the **previous
|
| 1223 |
+
revision's** capabilities.
|
| 1224 |
+
|
| 1225 |
+
**Root cause.** A serve process started before a code or environment change keeps serving from old code.
|
| 1226 |
+
The serve process reads its environment exactly once, at startup (`deploy/codespace/launch.sh:100-102`).
|
| 1227 |
+
|
| 1228 |
+
**Fix.** `launch.sh` already handles it: it records a **stamp** of the revision and the asset
|
| 1229 |
+
configuration and restarts the server when the stamp disagrees
|
| 1230 |
+
(`deploy/codespace/launch.sh:100-147`). The rule for a developer: **a restart, not a reload, is required
|
| 1231 |
+
after any env or revision change.**
|
| 1232 |
+
|
| 1233 |
+
> *"A stale serve process is worse than no process: it answers /v1/health and /v1/capabilities from OLD
|
| 1234 |
+
> code, so the deployment looks alive while reporting the previous revision's capabilities."*
|
| 1235 |
+
> (`deploy/codespace/launch.sh:111-113`)
|
| 1236 |
+
|
| 1237 |
+
### Related platform traps worth knowing
|
| 1238 |
+
|
| 1239 |
+
| Trap | Detail |
|
| 1240 |
+
|---|---|
|
| 1241 |
+
| a forwarded Codespace port returns `302` for a private repo | this is *why* the outbound tunnel exists (`release/repo/docs/DEPLOYMENT.md` Β§10) |
|
| 1242 |
+
| the tunnel agent must be started by the devcontainer `postStartCommand` | a restarted Codespace comes up with `agent_connected: false` otherwise |
|
| 1243 |
+
| never retry `POST /api/infer` at the gateway | a retry consumes inference twice |
|
| 1244 |
+
| `containerEnv` applies only at container **creation** | an env change needs a restart, which is why `launch.sh` re-exports on every start (`deploy/codespace/launch.sh:49-52`) |
|
| 1245 |
+
| the Codespace filesystem is **ephemeral** | uploaded assets and logs vanish with the Codespace (`deploy/codespace/launch.sh:44-47`) |
|
| 1246 |
+
|
| 1247 |
+
---
|
| 1248 |
+
|
| 1249 |
+
# Part X β Status and evidence
|
| 1250 |
+
|
| 1251 |
+
## 33. `NOT RUN` / `OPEN` / `BLOCKED` / `UNKNOWN` for development
|
| 1252 |
+
|
| 1253 |
+
| # | Item | Status |
|
| 1254 |
+
|---|---|---|
|
| 1255 |
+
| 1 | B-07 β tunnel gaps; patch prepared, **not deployed** | **`OPEN`** |
|
| 1256 |
+
| 2 | B-02 β `/api/health` `codespace_name` trailing `\n` | **`OPEN` (cosmetic)** |
|
| 1257 |
+
| 3 | A `LICENSE` file | **`OPEN`** β **no LICENSE file exists**; README says to add one before public release |
|
| 1258 |
+
| 4 | The monorepo README | **materially stale** β it calls the frontend "hermetic", describes a 4-endpoint `/v1/*` contract, omits the tunnel, and points at the stale `deploy/` (`release/CURRENT_RELEASE_STATE.md` Β§6) |
|
| 1259 |
+
| 5 | `hf/SETUP.md` and `hf/README.md` | **stale** β they assert the project ships no weights, which is now false (`release/CURRENT_RELEASE_STATE.md` Β§6) |
|
| 1260 |
+
| 6 | The exact formatter / linter configuration | **`UNKNOWN`** β no committed config in the files read (Β§23.9) |
|
| 1261 |
+
| 7 | The precise full-suite collected test count | **`UNKNOWN`** β per-suite counts are known, the total is not |
|
| 1262 |
+
| 8 | The exact Windows trigger for the delete-shim flake | **`UNKNOWN`** |
|
| 1263 |
+
| 9 | Whether `doctor.sh` / `tunnel_agent.py` exist in the deployed inference repo | **`UNKNOWN`** β the monorepo copy lacks them |
|
| 1264 |
+
| 10 | A system-level end-to-end benchmark | **`NOT RUN`** β none exists |
|
| 1265 |
+
| 11 | The router **test**-split number | **`NOT RUN`** β only validation (`n = 86`, ungated) exists |
|
| 1266 |
+
| 12 | Captioning benchmark | **`NOT RUN`** β implemented, not benchmarked |
|
| 1267 |
+
| 13 | A one-command retrain for the two external artifacts | **not implemented** β deliberate boundary (Β§21) |
|
| 1268 |
+
| 14 | `torch.compile`, CUDA, quantisation, thread capping | **not implemented** β CPU-first by design (Β§1) |
|
| 1269 |
+
| 15 | Enforcement of a one-model cache (`cache_max_models: 1`) | **UNVERIFIED** β the value is reported, not proven enforced (`docs/DEPLOYMENT_DECISION.md` Β§5) |
|
| 1270 |
+
|
| 1271 |
+
> **The stale-README trap is worth its own line.** `README.md` in the monorepo is *"materially stale"*: it
|
| 1272 |
+
> calls the frontend *"hermetic οΏ½οΏ½οΏ½ no backend calls"* (it calls `/api/*` on Render), puts Render/Codespace
|
| 1273 |
+
> as *"in progress"* (both deployed), describes a 4-endpoint `/v1/*` contract (the live contract is
|
| 1274 |
+
> `/api/*`), omits the tunnel, and points at the stale untracked `deploy/` as the deployment source
|
| 1275 |
+
> (`release/CURRENT_RELEASE_STATE.md` Β§6). The public release documentation is authoritative; the
|
| 1276 |
+
> monorepo README is not.
|
| 1277 |
+
|
| 1278 |
+
## 34. Where the evidence lives
|
| 1279 |
+
|
| 1280 |
+
| What | Where |
|
| 1281 |
+
|---|---|
|
| 1282 |
+
| the frozen config, invariants and hash | `configs/base.yaml`; `core/config.py:76-80,94-222` |
|
| 1283 |
+
| the binding schemas | `core/schemas.py` |
|
| 1284 |
+
| the specialist interface | `specialists/base.py` |
|
| 1285 |
+
| the capability table and lazy construction | `core/registry.py` |
|
| 1286 |
+
| the planner mapping | `core/planner.py:121-142` |
|
| 1287 |
+
| the dependency profiles and contract notes | `requirements.txt` |
|
| 1288 |
+
| the test markers | `pytest.ini` |
|
| 1289 |
+
| the launcher and warm-up | `deploy/codespace/launch.sh`, `deploy/codespace/warm_cache.py` |
|
| 1290 |
+
| the annotation-scope trap | `gateway/app.py:53-106`; `docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§13 |
|
| 1291 |
+
| the deploy manifest is inert | `docs/PHASE18_DEPLOYMENT_PACKAGING.md`; `scripts/validate_deploy_config.py` |
|
| 1292 |
+
| the change-VQA Kaggle guide | `docs/R02_KAGGLE_TRAINING_GUIDE.md`; `RUNBOOK_CHANGE_VQA_KAGGLE.md` |
|
| 1293 |
+
| the VLM LoRA Kaggle runbook | `RUNBOOK_PHASE6_VLM_KAGGLE.md` |
|
| 1294 |
+
| the test suites and their expected results | `release/repo/docs/REPRODUCIBILITY.md` Β§5 |
|
| 1295 |
+
| the environment traps | `release/repo/docs/REPRODUCIBILITY.md` Β§10 |
|
| 1296 |
+
| the platform traps | `release/repo/docs/DEPLOYMENT.md` Β§10 |
|
| 1297 |
+
| the live topology, env vars, cold start | `release/repo/docs/DEPLOYMENT.md` |
|
| 1298 |
+
| the operations manual | [../OPERATIONS.md](OPERATIONS.md) |
|
| 1299 |
+
| the factual inventory | `release/CURRENT_RELEASE_STATE.md` |
|
| 1300 |
+
|
| 1301 |
+
### Cross-references
|
| 1302 |
+
|
| 1303 |
+
| For | See |
|
| 1304 |
+
|---|---|
|
| 1305 |
+
| the frozen config and `Config.hash == 78f1e3700da15aa1` | [07 β Configuration and Freeze](architecture/07-configuration-freeze.md) |
|
| 1306 |
+
| the specialist contract in depth | [05 β Specialists](architecture/05-specialists.md) |
|
| 1307 |
+
| the router's five heads and the label space | [04 β Router](architecture/04-router.md) |
|
| 1308 |
+
| the evidence and confidence stages | [06 β Evidence and Confidence](architecture/06-evidence-and-confidence.md) |
|
| 1309 |
+
| the request lifecycle and the nine-state spine | [03 β Request Lifecycle](architecture/03-request-lifecycle.md) |
|
| 1310 |
+
| the four endpoints and error codes | [08 β The API Contract](architecture/08-api-contract.md) |
|
| 1311 |
+
| how to operate the live stack | [../OPERATIONS.md](OPERATIONS.md) |
|
| 1312 |
+
| per-artifact hyperparameters | [../TRAINING.md](TRAINING.md) |
|
| 1313 |
+
| what a third party can and cannot reproduce | [../REPRODUCIBILITY.md](REPRODUCIBILITY.md) |
|
| 1314 |
+
|
| 1315 |
+
---
|
| 1316 |
+
|
| 1317 |
+
> **Chapter summary.** Python 3.11+ and a CPU are sufficient; there is no CUDA requirement, and all
|
| 1318 |
+
> placement is `.to(device)`, never `.cuda()`. The single registry is `configs/base.yaml`, its hash
|
| 1319 |
+
> **`78f1e3700da15aa1`** is frozen, and **editing the config moves the hash and invalidates every artifact
|
| 1320 |
+
> keyed to it** β the one non-reversible action in the repository. `core/schemas.py` is the binding
|
| 1321 |
+
> contract and every specialist returns exactly `SpecialistResult`; the specialist interface is four
|
| 1322 |
+
> methods in `specialists/base.py`, and a new specialist is added by implementing it and adding a
|
| 1323 |
+
> `SpecialistSpec` row keyed on **capability**, not name. Tests run from the venv, targeted, because the
|
| 1324 |
+
> full suite trips the sandbox's bulk-delete guard. Four artifacts train locally; two β change-VQA and the
|
| 1325 |
+
> VLM LoRA β train on an external GPU and are not one-command reproducible. Eight development traps are
|
| 1326 |
+
> recorded, of which the most dangerous are the stale untracked `deploy/`, the annotation-scope trap that
|
| 1327 |
+
> turns a `Request` parameter into a required query parameter, and the Chrome focus trap that produces a
|
| 1328 |
+
> silent false pass. `NOT RUN`/`OPEN` items include B-07 and B-02 (both `OPEN`), the absent `LICENSE`, the
|
| 1329 |
+
> stale monorepo README and `hf/` docs, and several `UNKNOWN β not established from the available
|
| 1330 |
+
> evidence` gaps.
|