Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
release: add docs/TESTING.md
Browse files- docs/TESTING.md +1559 -0
docs/TESTING.md
ADDED
|
@@ -0,0 +1,1559 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# SatQuery AI — Testing, Verification and Evidence Integrity
|
| 2 |
+
|
| 3 |
+
**Status of this document.** This is the testing chapter of the public SatQuery AI release. It enumerates
|
| 4 |
+
the actual suites that exist, states what each one pins, and reports the full-suite result **including its
|
| 5 |
+
failures**. Every test file, count, command and code excerpt below was read from the repository. Where a
|
| 6 |
+
value could not be established it is written
|
| 7 |
+
`UNKNOWN — not established from the available evidence`.
|
| 8 |
+
|
| 9 |
+
**Status vocabulary** (see `DOCS_STYLE_GUIDE.md` §2): `IMPLEMENTED` · `VERIFIED` · `MEASURED` ·
|
| 10 |
+
`ATTEMPTED` · `NOT RUN` · `BLOCKED` · `DEFERRED` · `REJECTED` · `OPEN` · `RESOLVED` · `CLOSED`.
|
| 11 |
+
|
| 12 |
+
**The one rule this chapter follows without exception:** a green suite is not evidence about a property
|
| 13 |
+
nobody wrote an assertion for, and a failure is never reclassified as a pass. The project learned this the
|
| 14 |
+
hard way twice — a guard that pinned a *defect* as the specification
|
| 15 |
+
(`docs/DEPLOYMENT_ARCHITECTURE.md` §5.5) and a guard whose assertion was satisfied by an upstream fix so it
|
| 16 |
+
could not see the downstream site (`F-15b`). Both are recorded below rather than smoothed over.
|
| 17 |
+
|
| 18 |
+
---
|
| 19 |
+
|
| 20 |
+
## Table of contents
|
| 21 |
+
|
| 22 |
+
- [0. How to read this document](#0-how-to-read-this-document)
|
| 23 |
+
- [1. Testing philosophy: tests pin contracts, not implementation](#1-testing-philosophy-tests-pin-contracts-not-implementation)
|
| 24 |
+
- [2. The suite inventory](#2-the-suite-inventory)
|
| 25 |
+
- [3. The doc-guard tests: keeping documentation and code in sync](#3-the-doc-guard-tests-keeping-documentation-and-code-in-sync)
|
| 26 |
+
- [4. The evidence engine: purity, determinism and the reproducibility property](#4-the-evidence-engine-purity-determinism-and-the-reproducibility-property)
|
| 27 |
+
- [5. The frontend live-wiring suite](#5-the-frontend-live-wiring-suite)
|
| 28 |
+
- [6. The doc/frontend suite](#6-the-docfrontend-suite)
|
| 29 |
+
- [7. The full `tests/unit` result, and the 5–6 environmental failures](#7-the-full-testsunit-result-and-the-56-environmental-failures)
|
| 30 |
+
- [8. The live-validation harness and its integrity discipline](#8-the-live-validation-harness-and-its-integrity-discipline)
|
| 31 |
+
- [9. Verdict-independent evidence: recomputing from raw records](#9-verdict-independent-evidence-recomputing-from-raw-records)
|
| 32 |
+
- [10. How to run the tests](#10-how-to-run-the-tests)
|
| 33 |
+
- [11. What is NOT tested](#11-what-is-not-tested)
|
| 34 |
+
- [12. Open items and evidence index](#12-open-items-and-evidence-index)
|
| 35 |
+
|
| 36 |
+
---
|
| 37 |
+
|
| 38 |
+
## 0. How to read this document
|
| 39 |
+
|
| 40 |
+
### 0.1 The three evidence classes
|
| 41 |
+
|
| 42 |
+
The repository encodes an evidence class on every test via a pytest marker (`pytest.ini`):
|
| 43 |
+
|
| 44 |
+
```ini
|
| 45 |
+
markers =
|
| 46 |
+
unit: fast, no I/O, no server, no network. Evidence about a component in isolation.
|
| 47 |
+
integration: exercises two or more real components wired together in-process.
|
| 48 |
+
smoke: a minimal end-to-end path run against real local artifacts, not a mock.
|
| 49 |
+
real_inference: ran the real model on real inputs in this environment.
|
| 50 |
+
```
|
| 51 |
+
|
| 52 |
+
The comment above the markers states the purpose exactly, and it is the reason the markers exist rather
|
| 53 |
+
than being decorative:
|
| 54 |
+
|
| 55 |
+
> *"The markers exist to make a test's EVIDENCE CLASS explicit: a unit test never claims a deployment was
|
| 56 |
+
> exercised, and no test may be reported as 'real inference' unless it ran the real model on real data in
|
| 57 |
+
> this environment."*
|
| 58 |
+
|
| 59 |
+
And one deliberate non-marker, quoted in full because it is a subtle but important choice:
|
| 60 |
+
|
| 61 |
+
> *"`environment_blocked` is not a test marker: a blocked path is reported in the STEP 7 report rather than
|
| 62 |
+
> encoded as a permanently-skipped test, because **a skip can be mistaken for coverage**."*
|
| 63 |
+
|
| 64 |
+
That sentence is the spine of this chapter. A skip that looks like coverage is the failure mode the marker
|
| 65 |
+
design is built to avoid.
|
| 66 |
+
|
| 67 |
+
### 0.2 What "verified" means here
|
| 68 |
+
|
| 69 |
+
| Term | Means | Example |
|
| 70 |
+
|---|---|---|
|
| 71 |
+
| **Unit** | one component, no I/O, no server, no network | `tests/unit/test_evidence_engine.py` |
|
| 72 |
+
| **Integration** | two or more real components wired in-process | `tests/integration/test_step7_backend_chain.py` |
|
| 73 |
+
| **Smoke** | a minimal end-to-end path over real local artifacts | `tests/unit/phase6_smoke.py` (standalone) |
|
| 74 |
+
| **Real inference** | the real model on real inputs, in this environment | — see §11 |
|
| 75 |
+
| **Live validation** | the deployed stack driven by a browser | §8 |
|
| 76 |
+
|
| 77 |
+
The distinction matters most at the boundary between **integration** and **live**. `docs/API_CONTRACT.md`
|
| 78 |
+
§8 states the boundary plainly: *"no test in this repository dials a network address"*, and it points at
|
| 79 |
+
`docs/ITEM5_INTEGRATION_SUITE_SCOPE.md`, which *"records what the 26-test integration suite proves (the
|
| 80 |
+
app's boundary, in-process) and what only a live deployment can prove (reachability, cold start, memory
|
| 81 |
+
ceilings)."* A green `tests/integration` run is therefore **not** evidence about a deployment.
|
| 82 |
+
|
| 83 |
+
### 0.3 The two defects that shaped the discipline
|
| 84 |
+
|
| 85 |
+
Two incidents are referenced repeatedly below, because they are why this chapter is written the way it is:
|
| 86 |
+
|
| 87 |
+
1. **A guard that pinned the defect.** The test
|
| 88 |
+
`tests/unit/test_controller.py::test_trace_records_inputs_and_query` asserted
|
| 89 |
+
`envelope.trace.inputs == list(request.assets)` — which, after the upload handler rewrote handles into
|
| 90 |
+
paths, pinned a **filesystem-path disclosure as the specification**. The lesson, recorded in
|
| 91 |
+
`docs/DEPLOYMENT_ARCHITECTURE.md` §5.3: *"A guard is only as strong as the behaviour it pins, and this
|
| 92 |
+
one pinned the defect — a reminder that a green suite is not evidence about a property nobody wrote an
|
| 93 |
+
assertion for."*
|
| 94 |
+
2. **A guard that could not fail.** The `to_trace()` seam's scrub had **no covering test**: falsifying it
|
| 95 |
+
alone left the producer-side guard **GREEN**, because the producer had already scrubbed the value the
|
| 96 |
+
guard inspected. Recorded as F-15b in `docs/DEPLOYMENT_ARCHITECTURE.md` §5.5, and remedied by
|
| 97 |
+
`test_to_trace_scrubs_a_detail_that_arrives_from_any_constructor`, which constructs the entry
|
| 98 |
+
**directly** through the constructor.
|
| 99 |
+
|
| 100 |
+
Both are the same shape: *a guard whose assertion is satisfied by something other than the code under
|
| 101 |
+
test.*
|
| 102 |
+
|
| 103 |
+
---
|
| 104 |
+
|
| 105 |
+
## 1. Testing philosophy: tests pin contracts, not implementation
|
| 106 |
+
|
| 107 |
+
### 1.1 The philosophy, as the suites state it
|
| 108 |
+
|
| 109 |
+
The repository's test docstrings state a consistent philosophy, and the consistency is the point. Four
|
| 110 |
+
representative statements, each quoted from a module docstring:
|
| 111 |
+
|
| 112 |
+
**Doc-guards pin examples, not prose** (`tests/unit/test_api_contract_doc.py`):
|
| 113 |
+
|
| 114 |
+
> *"These tests validate the EXAMPLES, not the prose. They cannot detect a wrong sentence (e.g. a
|
| 115 |
+
> mis-stated latency budget). What they guarantee is that a frontend developer who copies an example
|
| 116 |
+
> verbatim gets a request the server accepts -- which is the failure mode that actually costs time."*
|
| 117 |
+
|
| 118 |
+
**Guards pin the route, not the validity** (`tests/unit/test_frontend_live_wiring.py`):
|
| 119 |
+
|
| 120 |
+
> *"It also pins the `chang` word-boundary defect: `\bchang\b` cannot match 'changed', so the page's own
|
| 121 |
+
> default question ('What changed here?') routed to `vqa` instead of the change path. That is invisible to
|
| 122 |
+
> any test that only checks 'is this a valid enum value' — both branches produce a valid value. The test
|
| 123 |
+
> therefore asserts the ROUTE, not just the validity."*
|
| 124 |
+
|
| 125 |
+
**Tests read shipped source when there is no importable symbol** (`tests/unit/test_frontend_live_wiring.py`):
|
| 126 |
+
|
| 127 |
+
> *"Why these are read from the source text: the frontend is plain ES5 with no build step and no module
|
| 128 |
+
> exports, so there is no importable symbol. Parsing the source is the only way to assert the shipped
|
| 129 |
+
> bytes. The parsing is anchored on named tokens rather than line numbers, so it does not silently pass if
|
| 130 |
+
> the file is restructured."*
|
| 131 |
+
|
| 132 |
+
**A runbook must not assert facts the repository contradicts** (`tests/unit/test_runbook_doc.py`):
|
| 133 |
+
|
| 134 |
+
> *"A runbook is worse than useless if a copied command fails or a quoted number was never measured: the
|
| 135 |
+
> operator concludes the *system* is broken rather than the *document*."*
|
| 136 |
+
|
| 137 |
+
### 1.2 The three anti-patterns the suites are written to avoid
|
| 138 |
+
|
| 139 |
+
| Anti-pattern | What it looks like | How the suites avoid it |
|
| 140 |
+
|---|---|---|
|
| 141 |
+
| **Pinning the defect** | an assertion that encodes today's buggy behaviour as correct | guards are **falsified** before being trusted: the fix is reverted and the guard is shown to fail (§1.3) |
|
| 142 |
+
| **Vacuous guards** | an assertion satisfied by an upstream layer, so the downstream site is untested | the seam is tested by constructing the object **directly** through its constructor (F-15b) |
|
| 143 |
+
| **Measurement by prose** | a text search that is tripped by documentation *about* a defect | guards walk the **parsed** examples, not the document text (`test_api_contract_doc.py::test_no_contract_example_fabricates_an_artifact_uri`) |
|
| 144 |
+
|
| 145 |
+
The third is worth quoting in full, because it is a rule about what a guard must fail on
|
| 146 |
+
(`tests/unit/test_api_contract_doc.py`):
|
| 147 |
+
|
| 148 |
+
> *"This walks the PARSED examples rather than the document text on purpose. The prose in section 2.4
|
| 149 |
+
> deliberately quotes the old `artifact://run/9f2c.../change_map.png` example when explaining why it was
|
| 150 |
+
> removed, so a text search would be tripped by documentation *about* the removal -- which is measuring
|
| 151 |
+
> the wrong thing. A guard must fail on the defect, not on a description of the defect."*
|
| 152 |
+
|
| 153 |
+
### 1.3 Falsification is part of the discipline
|
| 154 |
+
|
| 155 |
+
A guard is not trusted until it has been shown to **fail** against the defect it guards. The pattern is
|
| 156 |
+
recorded with measurements throughout the project. The clearest instance is the F-15 four-site
|
| 157 |
+
falsification (`docs/DEPLOYMENT_ARCHITECTURE.md` §5.5), where each site was reverted individually and the
|
| 158 |
+
guards re-run:
|
| 159 |
+
|
| 160 |
+
| Site reverted | `test_a_construction_failure_publishes_no_filesystem_path` | `test_to_trace_scrubs_a_detail_that_arrives_from_any_constructor` |
|
| 161 |
+
|---|---|---|
|
| 162 |
+
| `registry.py:597` (`_failure_entry`) | **FAILS** | passes *(correctly — it tests the seam, not the producer)* |
|
| 163 |
+
| `registry.py:309` (`to_trace`) | **passes — blind** | **FAILS** |
|
| 164 |
+
|
| 165 |
+
The "passes — blind" cell is the F-15b finding: the producer guard cannot see the seam, because the
|
| 166 |
+
producer has already scrubbed the value. The two guards are therefore **not duplicates**, and the table is
|
| 167 |
+
the evidence.
|
| 168 |
+
|
| 169 |
+
A second falsification table covers the two controller sites (`docs/DEPLOYMENT_ARCHITECTURE.md` §5.5):
|
| 170 |
+
|
| 171 |
+
| Site reverted | `test_a_failed_step_publishes_no_filesystem_path` — assertion that fires |
|
| 172 |
+
|---|---|
|
| 173 |
+
| `controller.py:801` (`_failure_evidence`) | line 590, `payload["message"]` |
|
| 174 |
+
| `controller.py:969` (`_warnings`) | line 582, `warnings[]` |
|
| 175 |
+
|
| 176 |
+
And a third, for the trace-path fix (`docs/DEPLOYMENT_ARCHITECTURE.md` §5.4):
|
| 177 |
+
|
| 178 |
+
| Step | Result |
|
| 179 |
+
|---|---|
|
| 180 |
+
| Revert only the `PARSE` site, run the module | **1 failed, 55 passed** at `tests/unit/test_controller.py:1153` |
|
| 181 |
+
| Restore byte-exact, verify hash | `0c00ae6f54300668122728dc706344044bc19d48328ceead30323c7147e04397` |
|
| 182 |
+
| Module re-run | **56 passed** |
|
| 183 |
+
| Is the guard vacuous? | **No** — the shape check only discriminates because the fixture writes real GeoTIFFs, and the name-equality check rejects the revert independently |
|
| 184 |
+
|
| 185 |
+
The "restore byte-exact, verify hash" row is the anti-tamper step: after reverting a site to falsify a
|
| 186 |
+
guard, the file is restored and its hash checked, so the falsification itself cannot leave the tree in a
|
| 187 |
+
modified state.
|
| 188 |
+
|
| 189 |
+
### 1.4 Determinism is asserted, not hoped for
|
| 190 |
+
|
| 191 |
+
Several suites pin **determinism** as a first-class property rather than a nice-to-have:
|
| 192 |
+
|
| 193 |
+
| Suite | Determinism property |
|
| 194 |
+
|---|---|
|
| 195 |
+
| `tests/unit/test_evidence_engine.py` | the same claims produce a byte-identical, identically-ordered collection even when uuids differ between processes (§4) |
|
| 196 |
+
| `tests/unit/test_router_threshold_sweep.py` | validation-only sweep contracts |
|
| 197 |
+
| `tests/unit/test_change_threshold_sweep.py` | validation-only sweep contracts |
|
| 198 |
+
| `tests/unit/test_report_generator.py` | *"pure, deterministic, fixtures only"* |
|
| 199 |
+
| `tests/unit/test_vlm_evaluate.py` | the predeclared acceptance rule |
|
| 200 |
+
| `tests/unit/test_fusion_seed_variance.py` | the predeclared seed-variance pass |
|
| 201 |
+
|
| 202 |
+
The determinism claim that matters most is the evidence engine's, because it is what makes the live
|
| 203 |
+
validation's verdict-recomputation property hold (§9).
|
| 204 |
+
|
| 205 |
+
---
|
| 206 |
+
|
| 207 |
+
## 2. The suite inventory
|
| 208 |
+
|
| 209 |
+
### 2.1 The layout
|
| 210 |
+
|
| 211 |
+
`pytest.ini` sets the roots:
|
| 212 |
+
|
| 213 |
+
```ini
|
| 214 |
+
[pytest]
|
| 215 |
+
testpaths = tests
|
| 216 |
+
pythonpath = .
|
| 217 |
+
addopts = -q --tb=short
|
| 218 |
+
filterwarnings =
|
| 219 |
+
ignore::DeprecationWarning
|
| 220 |
+
ignore::PendingDeprecationWarning
|
| 221 |
+
ignore::rasterio.errors.NotGeoreferencedWarning
|
| 222 |
+
```
|
| 223 |
+
|
| 224 |
+
`tests/` contains eight directories and two top-level modules:
|
| 225 |
+
|
| 226 |
+
| Location | Contents | Evidence class |
|
| 227 |
+
|---|---|---|
|
| 228 |
+
| `tests/unit/` | 98 Python files — 94 collected test modules + `__init__.py`, `conftest.py`, `_qa_probe_tokens.py`, `phase6_smoke.py` | `unit` (mostly) |
|
| 229 |
+
| `tests/e2e/` | `test_demo.py`, `test_gate4_e2e.py`, `_demo_probe.py` | end-to-end |
|
| 230 |
+
| `tests/geospatial/` | `test_transform.py` | `unit` |
|
| 231 |
+
| `tests/integration/` | `test_step7_backend_chain.py`, `test_step7_r02_serving.py` | `integration` |
|
| 232 |
+
| `tests/leakage/` | `test_leakage.py` | `unit` |
|
| 233 |
+
| `tests/model/` | `test_vlm_contract.py`, `test_vqa_quality_gate.py` | `unit` / `integration` |
|
| 234 |
+
| `tests/routing/` | `test_router.py` | `unit` |
|
| 235 |
+
| `tests/` | `test_config.py`, `test_schemas.py` | `unit` |
|
| 236 |
+
|
| 237 |
+
`tests/unit` totals **49,673 lines** across its 98 Python files. Two files there are **deliberately not
|
| 238 |
+
collected**:
|
| 239 |
+
|
| 240 |
+
| File | Why it is not collected |
|
| 241 |
+
|---|---|
|
| 242 |
+
| `tests/unit/_qa_probe_tokens.py` | *"QA scratch probe (NOT collected by pytest -- underscore prefix)."* |
|
| 243 |
+
| `tests/unit/phase6_smoke.py` | *"Phase 6 real CPU smoke test -- standalone, NOT collected by pytest."* |
|
| 244 |
+
| `tests/unit/conftest.py` | *"Shared fixtures for the Phase 6 VLM unit tests."* — a fixture module, not a test module |
|
| 245 |
+
| `tests/unit/__init__.py` | package marker |
|
| 246 |
+
|
| 247 |
+
### 2.2 The full file inventory, with what each covers
|
| 248 |
+
|
| 249 |
+
The descriptions below are the **first line of each module's own docstring** — the module's statement of
|
| 250 |
+
what it pins, not a paraphrase.
|
| 251 |
+
|
| 252 |
+
#### 2.2.1 Gateway, deployment and configuration
|
| 253 |
+
|
| 254 |
+
| File | What it covers (module docstring) |
|
| 255 |
+
|---|---|
|
| 256 |
+
| `test_gateway_app.py` | *"The gateway ASGI layer: its contract, its allowlist, and real HTTP execution."* |
|
| 257 |
+
| `test_gateway_assets.py` | *"The ephemeral asset store: opaque handles, three obligations, real lifetimes."* |
|
| 258 |
+
| `test_gateway_policy.py` | *"The gateway's decisions must be correct, cheap, and never reach the Space on a rejection."* |
|
| 259 |
+
| `test_gateway_responsibilities.py` | *"STEP 7 §C — the eighteen gateway responsibilities, as UNIT tests."* |
|
| 260 |
+
| `test_space_app.py` | *"The Space entrypoint must be cheap, honest, and hash-preserving."* |
|
| 261 |
+
| `test_deployment_adapter.py` | *"The deployment capability adapter: registry-authoritative, contract-shaped."* |
|
| 262 |
+
| `test_deploy_config.py` | *"Tests for the Phase 18 deploy packaging manifest and its validator."* |
|
| 263 |
+
| `test_deploy_requests.py` | *"Phase 18 §79 DEPLOYMENT: cold-start and sequential-request behaviour."* |
|
| 264 |
+
| `test_config.py` | *"Phase 1 tests — configuration registry and frozen contract guards."* |
|
| 265 |
+
| `test_code_revision.py` | *"Tests for `core.code_revision` — provenance defect section 6b."* |
|
| 266 |
+
| `test_packaging_script.py` | *"Regression tests for `scripts/package_kaggle_code.py`."* |
|
| 267 |
+
|
| 268 |
+
#### 2.2.2 Controller, planner, registry and schemas
|
| 269 |
+
|
| 270 |
+
| File | What it covers |
|
| 271 |
+
|---|---|
|
| 272 |
+
| `test_controller.py` | *"Tests for `core.controller` — dispatch, partial failure and assembly (Phase 15)."* |
|
| 273 |
+
| `test_controller_seam.py` | *"The controller's asset-resolution seam — regression tests for a real defect."* |
|
| 274 |
+
| `test_planner.py` | *"Tests for `core.planner` — the deterministic policy planner (Phase 15)."* |
|
| 275 |
+
| `test_registry.py` | *"Tests for `core.registry` — the specialist capability registry (Phase 15)."* |
|
| 276 |
+
| `test_schemas.py` | *"Phase 1 tests — canonical schema contract."* |
|
| 277 |
+
| `test_errors.py` | *"Tests for `core.errors` — the taxonomy, and the F-15 path scrubber."* |
|
| 278 |
+
| `test_evidence_engine.py` | *"Tests for the evidence engine — the aggregator (Phase 13)."* (§4) |
|
| 279 |
+
| `test_report_generator.py` | *"Unit tests for `reports.generator` — pure, deterministic, fixtures only."* |
|
| 280 |
+
|
| 281 |
+
#### 2.2.3 Router
|
| 282 |
+
|
| 283 |
+
| File | What it covers |
|
| 284 |
+
|---|---|
|
| 285 |
+
| `test_router_threshold_sweep.py` | *"Phase 4/13 — contracts for the validation-only ROUTER threshold sweep."* |
|
| 286 |
+
| `tests/routing/test_router.py` | *"Phase 4 tests — the intent router."* |
|
| 287 |
+
|
| 288 |
+
#### 2.2.4 Change detection and change-VQA (the largest subsystem by file count)
|
| 289 |
+
|
| 290 |
+
| File | What it covers |
|
| 291 |
+
|---|---|
|
| 292 |
+
| `test_change.py` | *"Phase 9 tests — change detection model, dataset, and post-processing."* |
|
| 293 |
+
| `test_change_levir_real_layout.py` | *"Phase 9 — the LEVIR-CD loader against the REAL dataset layout."* |
|
| 294 |
+
| `test_change_phase9_decontamination.py` | *"Phase 9 -- pins that stop the change-detection module from being re-contaminated."* |
|
| 295 |
+
| `test_change_specialist.py` | *"Phase 9 — the change specialist serving contract."* |
|
| 296 |
+
| `test_change_threshold_sweep.py` | *"Phase 9 — contracts for the validation-only threshold sweep."* |
|
| 297 |
+
| `test_change_train_script_contract.py` | *"Phase 9 — contracts for `scripts/train_change.py`."* |
|
| 298 |
+
| `test_change_vqa_amp.py` | *"R-02 — the mixed-precision training contract."* |
|
| 299 |
+
| `test_change_vqa_dataset.py` | *"R-02 test areas E-J — the CDVQA dataset layer."* |
|
| 300 |
+
| `test_change_vqa_eval_cli.py` | *"R-02 — the evaluation CLI's empty-split diagnosis."* |
|
| 301 |
+
| `test_change_vqa_features.py` | *"R-02 test areas K-L — frozen change features and the feature caches."* |
|
| 302 |
+
| `test_change_vqa_head.py` | *"R-02 test areas M-Q — batching, the reasoning head, and the metrics."* |
|
| 303 |
+
| `test_change_vqa_integration.py` | *"R-02 test area R — the change-VQA specialist and its planner/registry wiring."* |
|
| 304 |
+
| `test_change_vqa_kaggle_notebook.py` | *"R-02 — the Kaggle notebook's discovery contract."* |
|
| 305 |
+
| `test_change_vqa_prepare_script.py` | *"R-02 — `prepare_record.json` must be the provenance of the whole directory."* |
|
| 306 |
+
| `test_change_vqa_promotion.py` | *"STEP 1 — the promoted R-02 change-VQA head at its canonical serving path."* |
|
| 307 |
+
| `test_change_vqa_smoke.py` | *"R-02 — the local smoke training run, and the artifact it leaves behind."* |
|
| 308 |
+
| `test_change_vqa_vocab.py` | *"R-02 test areas A-D — the CDVQA answer space and question ontology."* |
|
| 309 |
+
| `test_cdvqa_adapter.py` | *"Tests for the CDVQA reader, join and example construction."* |
|
| 310 |
+
| `test_cdvqa_second_overlap.py` | *"Tests for scripts/check_cdvqa_second_overlap.py."* |
|
| 311 |
+
|
| 312 |
+
#### 2.2.5 Optical-SAR
|
| 313 |
+
|
| 314 |
+
| File | What it covers |
|
| 315 |
+
|---|---|
|
| 316 |
+
| `test_optical_sar_croma.py` | *"CROMA wrapper — the interface verified against the real model."* |
|
| 317 |
+
| `test_optical_sar_croma_geometry.py` | *"Phase 11 geometry guards — the four requirements that were *measured but…"* |
|
| 318 |
+
| `test_optical_sar_fusion_head.py` | *"Optical-SAR fusion head — the frozen concatenation and channel dropout."* |
|
| 319 |
+
| `test_optical_sar_prompts.py` | *"Optical-SAR prompts — the narration boundary."* |
|
| 320 |
+
| `test_optical_sar_radiometry.py` | *"CROMA encoder-input radiometry — the DEV-2 ruling, as executable guards."* |
|
| 321 |
+
| `test_optical_sar_sensor_adapter.py` | *"Optical-SAR sensor adapter — the two hard rules, enforced."* |
|
| 322 |
+
| `test_optical_sar_specialist.py` | *"Phase 11/12 — the optical-SAR specialist serving contract."* |
|
| 323 |
+
| `test_app_serving_optical_sar.py` | *"Regression guards for the optical-SAR serving wiring (Phase 14 / Pass 16)."* |
|
| 324 |
+
| `test_eval_fusion_115.py` | *"Tests for the pre-registered 11.5 metric tool."* |
|
| 325 |
+
| `test_fusion_extraction.py` | *"Tests for the Phase 12 frozen-feature extraction pipeline (fixtures only)."* |
|
| 326 |
+
| `test_fusion_seed_variance.py` | *"Tests for the predeclared seed-variance pass (Gate F D-03 / D-05)."* |
|
| 327 |
+
| `test_fusion_training.py` | *"Tests for the Phase 12 fusion-head training loop (fixtures only)."* |
|
| 328 |
+
| `test_reben_adapter.py` | *"Tests for the reBEN v2 -> PairedSample adapter (Phase 12)."* |
|
| 329 |
+
| `test_analyze_reben_labels.py` | *"Tests for the reBEN label analyser."* |
|
| 330 |
+
|
| 331 |
+
#### 2.2.6 Grounding
|
| 332 |
+
|
| 333 |
+
| File | What it covers |
|
| 334 |
+
|---|---|
|
| 335 |
+
| `test_grounding_dataset.py` | *"Tests for the Phase 8 grounding dataset."* |
|
| 336 |
+
| `test_grounding_degenerate_fallback.py` | *"Regression tests for the degenerate fallback-box guard (Phase 8)."* |
|
| 337 |
+
| `test_grounding_feature_extraction.py` | *"Tests for the frozen-feature cache."* |
|
| 338 |
+
| `test_grounding_head_wiring.py` | *"Wiring the trained RemoteCLIP grounding head into production (work order §5)."* |
|
| 339 |
+
| `test_grounding_metrics.py` | *"Tests for the grounding metrics."* |
|
| 340 |
+
| `test_grounding_nms.py` | *"Unit tests for grounding NMS (specialists/grounding/head.py::nms)."* |
|
| 341 |
+
| `test_vrsbench_loader.py` | *"Tests for the VRSBench referring-expression loader."* |
|
| 342 |
+
| `test_vrsbench_lazy_images.py` | *"Tests for the lazy-image mode of the VRSBench loader."* |
|
| 343 |
+
|
| 344 |
+
#### 2.2.7 VLM training and evaluation
|
| 345 |
+
|
| 346 |
+
| File | What it covers |
|
| 347 |
+
|---|---|
|
| 348 |
+
| `test_vlm_artifact.py` | *"Tests for the Phase 6 adapter artifact (`training.vlm.artifact`)."* |
|
| 349 |
+
| `test_vlm_collate.py` | *"Tests for Phase 6 batch assembly (`training.vlm.collate`)."* |
|
| 350 |
+
| `test_vlm_config.py` | *"Tests for the Phase 6 training recipe (`training.vlm.config`)."* |
|
| 351 |
+
| `test_vlm_dataset.py` | *"Tests for the Phase 6 BigEarthNet -> SmolVLM instruction corpus."* |
|
| 352 |
+
| `test_vlm_evaluate.py` | *"Tests for the predeclared acceptance rule (`training.vlm.evaluate`)."* |
|
| 353 |
+
| `test_vlm_evaluate_cache.py` | *"Tests for the resumable per-question answer cache in `training.vlm.evaluate`."* |
|
| 354 |
+
| `test_vlm_evaluate_decision_split.py` | *"Tests for `decide_acceptance(..., decision_split=...)`."* |
|
| 355 |
+
| `test_vlm_formatting.py` | *"Tests for Phase 6 prompt construction and loss masking (`training.vlm.formatting`)."* |
|
| 356 |
+
| `test_vlm_lora.py` | *"Tests for Phase 6 LoRA injection (`training.vlm.lora`)."* |
|
| 357 |
+
| `tests/model/test_vlm_contract.py` | *"Phase 5 tests — SmolVLM contract, prompts, and the VQA specialist."* |
|
| 358 |
+
| `tests/model/test_vqa_quality_gate.py` | *"Integration tests: the F5-5 quality gate blocks the VLM before it runs."* |
|
| 359 |
+
|
| 360 |
+
#### 2.2.8 Datasets, metrics and evaluation
|
| 361 |
+
|
| 362 |
+
| File | What it covers |
|
| 363 |
+
|---|---|
|
| 364 |
+
| `test_bigearthnet_blocks.py` | *"Tests for the T2 spatial-block scene key (Phase 12)."* |
|
| 365 |
+
| `test_bigearthnet_pairing.py` | *"Tests for the BigEarthNet optical<->SAR pairing step (Phase 12)."* |
|
| 366 |
+
| `test_bigearthnet_prep.py` | *"Tests for the BigEarthNet reader and instruction-pair construction."* |
|
| 367 |
+
| `test_prepare_script.py` | *"Tests for the BigEarthNet preparation CLI."* |
|
| 368 |
+
| `test_caption_metrics.py` | *"Caption metrics (ruling R-16): BLEU, ROUGE-L, CIDEr, BERTScore."* |
|
| 369 |
+
| `test_vqa_metrics.py` | *"Tests for the Tier-1 VQA / Change-VQA text metrics (`evaluation.metrics.vqa`)."* |
|
| 370 |
+
| `test_eval_normalize.py` | *"Tests for `evaluation.normalize` — the plan section 63 metric normaliser."* |
|
| 371 |
+
| `test_evaluation_runner.py` | *"STEP 5 — the evaluation runner."* |
|
| 372 |
+
| `test_calibration.py` | *"STEP 2 — the calibration fitter, the artifact, and the consumer wiring."* |
|
| 373 |
+
| `test_quality_gate.py` | *"Tests for the deterministic input-quality gate (finding F5-5)."* |
|
| 374 |
+
| `test_benchmark_adapters.py` | *"STEP 5 — benchmark adapters: contract, registry, and the non-invention guard."* |
|
| 375 |
+
| `test_benchmark_adapters_real.py` | *"The four real-corpus benchmark adapters: correctness against synthetic fixtures."* |
|
| 376 |
+
| `test_public_test_corpus.py` | *"STEP 5 — the immutable public-test corpus (plan section 37, Test Set Firewall)."* |
|
| 377 |
+
| `test_manifest_freeze.py` | *"Tests for the frozen dataset-manifest record (`evaluation/manifest_freeze.json`)."* |
|
| 378 |
+
| `test_prompt_freeze.py` | *"Tests for the frozen prompt-set record (`evaluation/prompt_freeze.json`)."* |
|
| 379 |
+
| `test_run_manifest_population.py` | *"Tests for STEP 4 — runtime population of the run manifest."* |
|
| 380 |
+
| `test_provenance_fixes.py` | *"Tests for the R-02 provenance fixes: sections 6a, 6b and 6c."* |
|
| 381 |
+
| `test_diagnose_feature_cache.py` | *"Tests for `scripts/diagnose_feature_cache.py`'s exit-code contract."* |
|
| 382 |
+
|
| 383 |
+
#### 2.2.9 Serving composition and the app boundary
|
| 384 |
+
|
| 385 |
+
| File | What it covers |
|
| 386 |
+
|---|---|
|
| 387 |
+
| `test_app_serving.py` | *"Regression tests for `app.serving` — the public serving composition root."* |
|
| 388 |
+
| `test_phase6_closure.py` | *"Tests for the Phase 6 closure record."* |
|
| 389 |
+
|
| 390 |
+
#### 2.2.10 Contract-conformance, frontend and documentation guards
|
| 391 |
+
|
| 392 |
+
| File | What it covers |
|
| 393 |
+
|---|---|
|
| 394 |
+
| `test_api_contract_doc.py` | *"The API contract must not drift from the schemas it documents."* |
|
| 395 |
+
| `test_frontend_guide_doc.py` | *"The contract documents must stay consistent with each other and the schemas."* |
|
| 396 |
+
| `test_frontend_live_wiring.py` | *"Regression tests for the frontend <-> orchestrator live wiring."* (§5) |
|
| 397 |
+
| `test_runbook_doc.py` | *"The deployment runbook must not assert facts the repository contradicts."* |
|
| 398 |
+
| `test_deploy_config.py` | *"Tests for the Phase 18 deploy packaging manifest and its validator."* |
|
| 399 |
+
| `test_step7_api_contract_conformance.py` | *"STEP 7 section D -- the API contract, tested against the schema that implements it."* |
|
| 400 |
+
| `test_safe_delete_shim.py` | *"The Windows safe-delete shim: verbatim prefixes, and the shell's return code."* (§7) |
|
| 401 |
+
|
| 402 |
+
#### 2.2.11 The other directories
|
| 403 |
+
|
| 404 |
+
| File | What it covers |
|
| 405 |
+
|---|---|
|
| 406 |
+
| `tests/geospatial/test_transform.py` | *"Phase 2 tests — geospatial transforms and the raster input contract."* |
|
| 407 |
+
| `tests/leakage/test_leakage.py` | *"Phase 3 tests — Gate 1: dataset manifests and scene-level leakage isolation."* |
|
| 408 |
+
| `tests/integration/test_step7_backend_chain.py` | *"STEP 7 sections H, I, M -- the real inference service, driven over HTTP."* |
|
| 409 |
+
| `tests/integration/test_step7_r02_serving.py` | *"STEP 7 section I -- the R-02 head, verified through the real serving path."* |
|
| 410 |
+
| `tests/e2e/test_demo.py` | *"End-to-end tests for the Phase-19 demonstration driver (`demo.run_demo`)."* |
|
| 411 |
+
| `tests/e2e/test_gate4_e2e.py` | *"Gate-4 end-to-end test — the full control tier in its real local state."* |
|
| 412 |
+
| `tests/e2e/_demo_probe.py` | *"Subprocess probe: run the demo in a FRESH interpreter and report which…"* (helper) |
|
| 413 |
+
|
| 414 |
+
### 2.3 What the inventory says about coverage shape
|
| 415 |
+
|
| 416 |
+
Three observations a reader should draw from the inventory:
|
| 417 |
+
|
| 418 |
+
1. **The largest single cluster is change detection and change-VQA** (19 files), which matches the
|
| 419 |
+
project's stated engineering priority ordering (`docs/MASTER_ARCHITECTURE_PLAN.md` §2.1: optical-SAR
|
| 420 |
+
first, change second).
|
| 421 |
+
2. **Every capability has a *serving contract* test, separate from its *model* tests.** For example
|
| 422 |
+
`test_change.py` covers the model/dataset/post-processing while `test_change_specialist.py` covers
|
| 423 |
+
*"the change specialist serving contract"*; the same split exists for optical-SAR
|
| 424 |
+
(`test_optical_sar_fusion_head.py` vs `test_optical_sar_specialist.py`) and grounding. The split is
|
| 425 |
+
deliberate: a model that works in a notebook and a specialist that serves the contract are different
|
| 426 |
+
claims.
|
| 427 |
+
3. **Documentation is tested as code.** Five of the 94 modules are doc-guards (§3), and they are counted in
|
| 428 |
+
the 183-passed doc/frontend suite (§6) rather than being an afterthought.
|
| 429 |
+
|
| 430 |
+
### 2.4 The integration suite's documented scope
|
| 431 |
+
|
| 432 |
+
`docs/ITEM5_INTEGRATION_SUITE_SCOPE.md` is the record of what the integration suite proves. The contract
|
| 433 |
+
summarises it (`docs/API_CONTRACT.md` §8):
|
| 434 |
+
|
| 435 |
+
> *"`docs/ITEM5_INTEGRATION_SUITE_SCOPE.md` records what the 26-test integration suite proves (the app's
|
| 436 |
+
> *boundary*, in-process) and what only a live deployment can prove (reachability, cold start, memory
|
| 437 |
+
> ceilings). Read it before treating a green `tests/integration` run as evidence about a deployment —
|
| 438 |
+
> **no test in this repository dials a network address**, including this contract's own `/v1/*`
|
| 439 |
+
> examples."*
|
| 440 |
+
|
| 441 |
+
The exact number of tests **collected** in a single `tests/integration` invocation is
|
| 442 |
+
`UNKNOWN — not established from the available evidence`. The documented figure is 26; the evidence this
|
| 443 |
+
chapter can establish is the two module files and their stated scope.
|
| 444 |
+
|
| 445 |
+
---
|
| 446 |
+
|
| 447 |
+
## 3. The doc-guard tests: keeping documentation and code in sync
|
| 448 |
+
|
| 449 |
+
### 3.1 Why documentation is tested
|
| 450 |
+
|
| 451 |
+
Five modules treat documentation as a **tested artefact**. The rationale is stated in each module's
|
| 452 |
+
docstring, and the common thread is that a documentation defect costs a real debugging session:
|
| 453 |
+
|
| 454 |
+
| Guard | The failure it prevents |
|
| 455 |
+
|---|---|
|
| 456 |
+
| `test_api_contract_doc.py` | a frontend builds against an example that does not validate, and the cause looks like a frontend bug |
|
| 457 |
+
| `test_frontend_guide_doc.py` | a value a frontend must send or read is documented wrong, producing a `422` and a wasted session |
|
| 458 |
+
| `test_runbook_doc.py` | an operator copies a command that fails, or quotes a number that was never measured, and concludes the *system* is broken |
|
| 459 |
+
| `test_deploy_config.py` | the deploy manifest silently moves the frozen `Config.hash` |
|
| 460 |
+
| `test_step7_api_contract_conformance.py` | the contract and the schema that implements it drift apart |
|
| 461 |
+
|
| 462 |
+
`test_api_contract_doc.py` records what the guards actually catch, with a list of four real defects the
|
| 463 |
+
act of writing the contract produced:
|
| 464 |
+
|
| 465 |
+
> *"Writing this contract produced exactly that failure four times over -- `Box` fields were nested instead
|
| 466 |
+
> of flat, `Evidence` used `summary`/`value`/`source` instead of
|
| 467 |
+
> `coordinates`/`score`/`source_specialist`, `CoordinateSystem` values were shorthand rather than the real
|
| 468 |
+
> `normalized_0_1`/`pixel`/`geo`, and `ExecutionTrace.inputs`/`outputs` were dicts rather than lists. Each
|
| 469 |
+
> was caught by validating the examples against the real models, which is what these tests do
|
| 470 |
+
> permanently."*
|
| 471 |
+
|
| 472 |
+
### 3.2 `test_api_contract_doc.py`
|
| 473 |
+
|
| 474 |
+
**Scope:** validate every fenced ```json block in `docs/API_CONTRACT.md` against the real Pydantic models.
|
| 475 |
+
|
| 476 |
+
| Mechanism | Detail |
|
| 477 |
+
|---|---|
|
| 478 |
+
| Fixture `contract_text` | asserts the document exists, then reads it |
|
| 479 |
+
| Fixture `json_blocks` | regex-extracts every ```json block and parses it; a block that is not valid JSON raises `AssertionError` naming the block index |
|
| 480 |
+
| `_find(blocks, predicate)` | locates an example by shape rather than by position |
|
| 481 |
+
|
| 482 |
+
| Test | What it pins |
|
| 483 |
+
|---|---|
|
| 484 |
+
| `test_every_json_block_in_the_contract_parses` | *"A malformed example is unusable; this fails loudly rather than silently."* |
|
| 485 |
+
| `test_the_result_envelope_example_validates` | the envelope example validates against `ResultEnvelope`; `schema_version == SCHEMA_VERSION`; `confidence.method in ("uncalibrated", "temperature_scaling")` |
|
| 486 |
+
| `test_no_contract_example_fabricates_an_artifact_uri` | no parsed example contains an `artifact://` URI or a filesystem path in `change_map` / `artifact_ref` (walks the **parsed** examples, §1.2) |
|
| 487 |
+
| (further tests) | documented task/coordinate-system values are the real ones; every error code documented; the frozen config hash is recorded; the calibration caveat is recorded |
|
| 488 |
+
|
| 489 |
+
The `confidence.method` assertion is the honest-signal guard: it pins that the contract's example cannot
|
| 490 |
+
claim a calibration method the engine does not have. §4 shows the engine side of the same rule.
|
| 491 |
+
|
| 492 |
+
### 3.3 `test_frontend_guide_doc.py`
|
| 493 |
+
|
| 494 |
+
**Scope:** `docs/API_CONTRACT.md` and `docs/FRONTEND_INTEGRATION.md` must agree with each other **and** with
|
| 495 |
+
`core/schemas.py`.
|
| 496 |
+
|
| 497 |
+
It imports the real models — `Box, CoordinateSystem, Evidence, EvidenceType, Region, Task` — and compares
|
| 498 |
+
the prose against them. The docstring lists the five defects that motivated it:
|
| 499 |
+
|
| 500 |
+
> *"Writing them produced, in order: `CoordinateSystem` documented as `normalized`/`geographic` (real:
|
| 501 |
+
> `normalized_0_1`/`geo`), `Box` geometry documented as nested (real: flat), `Evidence` fields named
|
| 502 |
+
> `summary`/`value`/`source` (real: `coordinates`/`score`/`source_specialist`),
|
| 503 |
+
> `ExecutionTrace.inputs`/`outputs` as objects (real: lists), and `Task.unsupported` missing from both
|
| 504 |
+
> tables."*
|
| 505 |
+
|
| 506 |
+
The representative test asserts a **bidirectional** requirement:
|
| 507 |
+
|
| 508 |
+
```python
|
| 509 |
+
def test_every_task_is_documented_in_both_documents(api_text, frontend_text):
|
| 510 |
+
for task in Task:
|
| 511 |
+
assert f"`{task.value}`" in api_text, f"Task.{task.value} missing from API contract"
|
| 512 |
+
assert f"`{task.value}`" in frontend_text, (
|
| 513 |
+
f"Task.{task.value} missing from the frontend guide; a value the client "
|
| 514 |
+
f"may receive but cannot find documented is a crash waiting to happen"
|
| 515 |
+
)
|
| 516 |
+
```
|
| 517 |
+
|
| 518 |
+
The message on the second assertion is the rationale: *"a value the client may receive but cannot find
|
| 519 |
+
documented is a crash waiting to happen."*
|
| 520 |
+
|
| 521 |
+
The suite also pins: no login screen, no secrets in the browser, the confidence-property trap, and
|
| 522 |
+
serialized analyses (`docs/EVALUATION.md` §7.2).
|
| 523 |
+
|
| 524 |
+
### 3.4 `test_runbook_doc.py`
|
| 525 |
+
|
| 526 |
+
**Scope:** `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` must not assert facts the repository contradicts.
|
| 527 |
+
|
| 528 |
+
The module names **two failure modes**:
|
| 529 |
+
|
| 530 |
+
> 1. *"**Invented measurements.** The runbook quotes command output. If it quotes a byte size, a
|
| 531 |
+
> capability list, a function name or an exit code that the repository does not produce, the quote is
|
| 532 |
+
> fabricated. Where a fact is cheap to verify locally, this module verifies it."*
|
| 533 |
+
> 2. *"**Invented API surface.** The runbook calls into the codebase (`build_serving_registry`,
|
| 534 |
+
> `available()`, `CHANGE_CHECKPOINT`, `CHANGE_VQA_HEAD`). A runbook that names a function the code does
|
| 535 |
+
> not export is a runbook whose commands raise `AttributeError` on the operator's first copy-paste."*
|
| 536 |
+
|
| 537 |
+
It reads three documents — the runbook, `docs/DEPLOYMENT_ARCHITECTURE.md` and `docs/API_CONTRACT.md` — so
|
| 538 |
+
that a claim contradicted *across* documents is caught, not only a claim contradicted by code.
|
| 539 |
+
|
| 540 |
+
**The most important class of assertion in this module** is stated in its docstring:
|
| 541 |
+
|
| 542 |
+
> *"The absence of a deployment cannot be tested here, so these tests assert the runbook's *local* claims
|
| 543 |
+
> and, separately, assert that the document itself labels its unimplemented sections as unimplemented. The
|
| 544 |
+
> second class matters more than it looks: a runbook that describes a live deployment without having
|
| 545 |
+
> performed one is the exact overclaim the work order forbids."*
|
| 546 |
+
|
| 547 |
+
So the guard pins **honesty about status**, not merely factual accuracy: an "undecided" item must be
|
| 548 |
+
declared undecided, and a "designed" item must not be described as "verified" (`docs/EVALUATION.md` §7.2).
|
| 549 |
+
|
| 550 |
+
### 3.5 `test_deploy_config.py`
|
| 551 |
+
|
| 552 |
+
**Scope:** the Phase 18 deploy manifest and its validator, and the **config-hash regression guard**.
|
| 553 |
+
|
| 554 |
+
The docstring states the two properties it proves:
|
| 555 |
+
|
| 556 |
+
> *"These tests prove the manifest is INERT — it is never merged into the config registry and cannot move
|
| 557 |
+
> `Config.hash` — and that the validator rejects each failure mode. Negative cases use in-memory dicts via
|
| 558 |
+
> `validate_documents`, so the real config files are never mutated."*
|
| 559 |
+
|
| 560 |
+
Key elements:
|
| 561 |
+
|
| 562 |
+
| Element | Detail |
|
| 563 |
+
|---|---|
|
| 564 |
+
| `REGISTRY_MARKER` | `test_deploy_manifest_exists_and_declares_the_non_registry_marker` asserts `doc[REGISTRY_MARKER] is False` |
|
| 565 |
+
| `test_validator_reports_no_errors_on_the_real_files` | `assert validate() == []` |
|
| 566 |
+
| `test_validator_cli_exits_zero_as_a_standalone_script` | drives the real entry point **in a child interpreter** (the same one running the suite) *"to prove the module's own `sys.path` bootstrap makes `core` importable WITHOUT pytest's `pythonpath = .` shortcut"* |
|
| 567 |
+
| `GPU_DURATION_KEYS` | `gpu_duration_vqa`, `gpu_duration_grounding`, `gpu_duration_change`, `gpu_duration_optical_sar` |
|
| 568 |
+
| `FROZEN_CONFIG_HASH` | imported from `tests.test_config` rather than re-declared — *"We import it rather than re-declare it, so the two cannot drift."* |
|
| 569 |
+
|
| 570 |
+
The `FROZEN_CONFIG_HASH` import is the anti-drift move applied to the guard itself: the deploy-config test
|
| 571 |
+
does not carry its own copy of `78f1e3700da15aa1`, it imports the authority.
|
| 572 |
+
|
| 573 |
+
The suite also pins the C-8 constraints: `torch_compile` false, CPU mode required, lazy load, single model
|
| 574 |
+
cache (`docs/EVALUATION.md` §7.2).
|
| 575 |
+
|
| 576 |
+
### 3.6 `test_step7_api_contract_conformance.py`
|
| 577 |
+
|
| 578 |
+
*"STEP 7 section D -- the API contract, tested against the schema that implements it."* This is the
|
| 579 |
+
schema-level counterpart to `test_api_contract_doc.py`: where the doc-guard validates the *examples* in the
|
| 580 |
+
markdown, this module tests the *contract* against the implementing schema directly.
|
| 581 |
+
|
| 582 |
+
### 3.7 What the doc-guards cannot do
|
| 583 |
+
|
| 584 |
+
The modules say so themselves, and the honesty is worth reproducing because it bounds what a green run
|
| 585 |
+
means:
|
| 586 |
+
|
| 587 |
+
> *"Prose can still be wrong in ways these tests cannot see -- a mis-stated latency, a wrong reason for a
|
| 588 |
+
> decision. What is asserted is that every value a frontend must send or read is the real value, which is
|
| 589 |
+
> the class of error that produces a 422 and a wasted debugging session."*
|
| 590 |
+
> — `tests/unit/test_frontend_guide_doc.py`
|
| 591 |
+
|
| 592 |
+
> *"These tests validate the EXAMPLES, not the prose. They cannot detect a wrong sentence (e.g. a
|
| 593 |
+
> mis-stated latency budget)."*
|
| 594 |
+
> — `tests/unit/test_api_contract_doc.py`
|
| 595 |
+
|
| 596 |
+
So a doc-guard guarantees **example-level and value-level** conformance. It does not guarantee that a
|
| 597 |
+
sentence is true. That distinction is the reason `test_runbook_doc.py` exists separately — it targets a
|
| 598 |
+
different class of claim (numbers, names, and status honesty) that the example-validators cannot reach.
|
| 599 |
+
|
| 600 |
+
---
|
| 601 |
+
|
| 602 |
+
## 4. The evidence engine: purity, determinism and the reproducibility property
|
| 603 |
+
|
| 604 |
+
### 4.1 Why this suite exists at all
|
| 605 |
+
|
| 606 |
+
The evidence engine is the component that turns a `SpecialistResult` into a *proof*: an ordered,
|
| 607 |
+
identified, digestible `EvidenceCollection`. Everything the release claims about "evidence" downstream —
|
| 608 |
+
the trace, the audit trail, the ability to recompute a verdict from a recorded run — rests on one
|
| 609 |
+
property of this component: **the same inputs must produce the same output, byte for byte.**
|
| 610 |
+
|
| 611 |
+
If that property fails, then two runs of the same request can produce two different evidence sets, and
|
| 612 |
+
no verdict can be recomputed from a record. The evidence engine is therefore the load-bearing component
|
| 613 |
+
behind §9's "verdict-independent evidence" claim, and its suite is the guard on that load-bearing
|
| 614 |
+
property.
|
| 615 |
+
|
| 616 |
+
The suite is `tests/unit/test_evidence_engine.py` — **73 test functions** in 877 lines
|
| 617 |
+
(`tests/unit/test_evidence_engine.py`).
|
| 618 |
+
|
| 619 |
+
### 4.2 The contract, in the module's own words
|
| 620 |
+
|
| 621 |
+
The module docstring states exactly two failure modes it is written against, and a third honesty rule:
|
| 622 |
+
|
| 623 |
+
> *"The engine's whole value is that it is *reproducible* and *lossless*. So the tests are mostly about
|
| 624 |
+
> the two ways that can silently break:*
|
| 625 |
+
>
|
| 626 |
+
> *determinism — the same claims must produce the same ordered, identified collection even when
|
| 627 |
+
> specialists finish in a different order, or when the uuids differ between processes.*
|
| 628 |
+
> *no loss — no specialist's evidence may vanish without being counted, and agreement between
|
| 629 |
+
> specialists must be recorded rather than collapsed with one name thrown away.*
|
| 630 |
+
>
|
| 631 |
+
> *Plus the honesty rule on confidence: an uncalibrated engine must pass the raw score through and SAY
|
| 632 |
+
> it is uncalibrated, never manufacture a fitted mapping."*
|
| 633 |
+
> — `tests/unit/test_evidence_engine.py`
|
| 634 |
+
|
| 635 |
+
Three properties, then: **determinism**, **no loss**, and **confidence honesty**. The suite is
|
| 636 |
+
organised around them.
|
| 637 |
+
|
| 638 |
+
### 4.3 The determinism family
|
| 639 |
+
|
| 640 |
+
The core purity claim is a single test whose docstring is the whole argument:
|
| 641 |
+
|
| 642 |
+
```python
|
| 643 |
+
def test_aggregate_is_deterministic_across_repeat_calls() -> None:
|
| 644 |
+
"""Same input -> byte-identical output. The core purity claim."""
|
| 645 |
+
```
|
| 646 |
+
|
| 647 |
+
It asserts three things across two calls with the same input — identical `ids()`, identical
|
| 648 |
+
`evidence_digest()`, and identical `model_dump()` lists. "Byte-identical output" is not a figure of
|
| 649 |
+
speech: the digest is the serialised form, so a difference anywhere in the collection changes the digest.
|
| 650 |
+
|
| 651 |
+
The determinism family, read from the file:
|
| 652 |
+
|
| 653 |
+
| Test | What it pins |
|
| 654 |
+
|---|---|
|
| 655 |
+
| `test_aggregate_is_deterministic_across_repeat_calls` | the core purity claim (line 132) |
|
| 656 |
+
| `test_order_is_independent_of_input_order` | forward vs. backward input order give the same digest |
|
| 657 |
+
| `test_equal_scores_order_deterministically_by_coordinates` | *"Ties must not fall back to insertion order — that is input-order dependence wearing a disguise"* |
|
| 658 |
+
| `test_higher_score_sorts_first_within_a_type` | the ordering rule inside an evidence type |
|
| 659 |
+
| `test_unscored_items_sort_after_scored_items` | a total order even when some items carry no score |
|
| 660 |
+
| `test_types_are_grouped_together` | type grouping is stable |
|
| 661 |
+
| `test_ids_are_sequential_and_zero_padded` | the identifier format |
|
| 662 |
+
| `test_ids_restart_from_one_for_each_aggregation` | ids are a function of the collection, not global state |
|
| 663 |
+
| `test_source_results_are_not_mutated` | purity — computing evidence does not write back |
|
| 664 |
+
| `test_result_objects_are_not_mutated` | purity, at the result level |
|
| 665 |
+
|
| 666 |
+
The two mutation tests are the *purity* half of "purity and determinism": the engine is a function, not
|
| 667 |
+
a transformer of its inputs. A caller's `SpecialistResult` is the same object after `aggregate` as
|
| 668 |
+
before.
|
| 669 |
+
|
| 670 |
+
### 4.4 The no-loss family
|
| 671 |
+
|
| 672 |
+
Losslessness is the harder property to test, because the ways evidence can vanish are subtle. The
|
| 673 |
+
deduplication tests are where it is pinned:
|
| 674 |
+
|
| 675 |
+
| Test | What it pins |
|
| 676 |
+
|---|---|
|
| 677 |
+
| `test_payload_only_differences_are_deduplicated_as_one_claim` | two items differing only in payload are one claim |
|
| 678 |
+
| `test_identical_items_from_one_specialist_collapse` | intra-specialist dedup |
|
| 679 |
+
| `test_deduplication_ignores_random_uuid` | dedup is on content, not on the process-local uuid |
|
| 680 |
+
| `test_deduplication_ignores_payload_differences_and_merges_them` | payloads merge rather than one being dropped |
|
| 681 |
+
| `test_different_coordinates_are_not_deduplicated` | dedup does not over-collapse |
|
| 682 |
+
| `test_different_coordinate_systems_are_not_deduplicated` | a normalised box and a pixel box are different claims |
|
| 683 |
+
| `test_same_claim_from_two_specialists_is_kept_and_marked_corroborated` | **agreement is recorded, not collapsed** — the second specialist's name is kept as corroboration |
|
| 684 |
+
| `test_unrelated_items_carry_no_corroboration_marker` | the marker is meaningful, not always-on |
|
| 685 |
+
| `test_multiple_results_aggregate_without_loss` | the headline no-loss assertion |
|
| 686 |
+
| `test_aggregation_does_not_double_count_a_specialists_own_evidence` | lossless ≠ duplicating |
|
| 687 |
+
|
| 688 |
+
The corroboration tests are the direct expression of the docstring's *"agreement between specialists
|
| 689 |
+
must be recorded rather than collapsed with one name thrown away."* This is a design decision encoded
|
| 690 |
+
as a test: when two specialists agree, the collection records **both** names and marks the claim
|
| 691 |
+
corroborated, rather than keeping one and discarding the other.
|
| 692 |
+
|
| 693 |
+
The cap tests pin that truncation is *counted*, not silent:
|
| 694 |
+
|
| 695 |
+
| Test | What it pins |
|
| 696 |
+
|---|---|
|
| 697 |
+
| `test_zero_max_items_is_rejected` | an invalid cap is refused |
|
| 698 |
+
| `test_cap_is_respected_and_the_drop_is_counted` | items past the cap are dropped **and counted** |
|
| 699 |
+
| `test_no_truncation_flag_when_under_the_cap` | the truncation flag is truthful |
|
| 700 |
+
| `test_default_cap_matches_the_config_value` | `DEFAULT_MAX_ITEMS` tracks the config |
|
| 701 |
+
| `test_sources_are_reported_before_the_cap_is_applied` | *which* specialists contributed is recorded before truncation |
|
| 702 |
+
|
| 703 |
+
That last test is a specific honesty property: if the cap drops a specialist's only item, the collection
|
| 704 |
+
still records that the specialist contributed. Losslessness at the *source* level survives truncation at
|
| 705 |
+
the *item* level.
|
| 706 |
+
|
| 707 |
+
The empty/degenerate cases are pinned too: `test_empty_input_yields_an_empty_collection`,
|
| 708 |
+
`test_a_result_with_no_evidence_contributes_nothing`, `test_accepts_a_single_result_without_wrapping`,
|
| 709 |
+
and the two argument-error tests `test_both_results_and_evidence_is_rejected` /
|
| 710 |
+
`test_neither_results_nor_evidence_is_rejected`.
|
| 711 |
+
|
| 712 |
+
### 4.5 The confidence honesty rule
|
| 713 |
+
|
| 714 |
+
This is where the project's style guide (`DOCS_STYLE_GUIDE.md` §3) meets a test. The rule the engine
|
| 715 |
+
must obey: an uncalibrated engine **passes the raw score through and says it is uncalibrated**; it never
|
| 716 |
+
manufactures a fitted mapping. The tests that pin this:
|
| 717 |
+
|
| 718 |
+
| Test | What it pins |
|
| 719 |
+
|---|---|
|
| 720 |
+
| `test_no_calibration_passes_raw_through_and_says_so` | the default is identity + an explicit "uncalibrated" label |
|
| 721 |
+
| `test_uncalibrated_is_never_labelled_as_fitted` | the label cannot be silently upgraded |
|
| 722 |
+
| `test_identity_temperature_is_reported_as_uncalibrated` | `T = 1.0` is *not* calibration |
|
| 723 |
+
| `test_engine_without_calibration_does_not_claim_calibration` | the engine-level claim |
|
| 724 |
+
| `test_raw_outside_unit_range_is_clamped_not_rejected` | a robustness decision, pinned |
|
| 725 |
+
| `test_non_finite_raw_is_refused` | NaN/Inf are refused, not clamped |
|
| 726 |
+
|
| 727 |
+
This matters because the project's calibration genuinely made the metric **worse** (ECE
|
| 728 |
+
`0.013755 → 0.014929`, `DOCS_STYLE_GUIDE.md` §3) and is retained only because it is in the frozen
|
| 729 |
+
config. The engine must therefore never present calibration as an improvement, and these tests are what
|
| 730 |
+
enforce that at the code level. The `METHOD_TEMPERATURE` / `METHOD_UNCALIBRATED` constants are the two
|
| 731 |
+
labels, and `test_uncalibrated_is_never_labelled_as_fitted` is the guard that the first is never used
|
| 732 |
+
for the second.
|
| 733 |
+
|
| 734 |
+
The fitted path is tested separately and is deterministic:
|
| 735 |
+
`test_temperature_scaling_is_applied_when_fitted`,
|
| 736 |
+
`test_temperature_above_one_softens_and_below_one_sharpens`,
|
| 737 |
+
`test_temperature_scaling_is_monotonic`, `test_temperature_scaling_is_deterministic`,
|
| 738 |
+
`test_endpoint_scores_do_not_produce_nan`, `test_invalid_temperatures_are_rejected`, and
|
| 739 |
+
`test_specialist_degradation_survives_calibration` / `test_specialist_components_are_preserved_through_calibration`.
|
| 740 |
+
|
| 741 |
+
### 4.6 Calibration artifact I/O
|
| 742 |
+
|
| 743 |
+
The engine reads a calibration artifact from disk, and the tests pin the failure modes explicitly rather
|
| 744 |
+
than letting them degrade silently:
|
| 745 |
+
|
| 746 |
+
| Test | What it pins |
|
| 747 |
+
|---|---|
|
| 748 |
+
| `test_artifact_round_trips_through_json` | write→read is lossless |
|
| 749 |
+
| `test_nested_artifact_form_is_accepted` | both artifact shapes are accepted |
|
| 750 |
+
| `test_missing_artifact_degrades_instead_of_failing` | a missing artifact is a *degradation* by default |
|
| 751 |
+
| `test_missing_artifact_raises_when_required` | …but raises when the caller says it is required |
|
| 752 |
+
| `test_malformed_artifact_is_not_silently_swallowed` | a corrupt artifact is an error, not a shrug |
|
| 753 |
+
| `test_artifact_without_a_temperature_is_rejected` | the required field is required |
|
| 754 |
+
| `test_load_calibration_respects_the_master_switch` | the config switch is honoured |
|
| 755 |
+
| `test_load_calibration_reads_the_configured_file` | the configured path is the one read |
|
| 756 |
+
| `test_load_calibration_without_a_file_configured_is_none` | no config ⇒ `None`, not a guess |
|
| 757 |
+
| `test_engine_from_config_uses_the_configured_cap` / `..._falls_back_to_the_default_cap` | the cap wiring |
|
| 758 |
+
|
| 759 |
+
The pairing of `degrades_instead_of_failing` with `raises_when_required` is the pattern this repository
|
| 760 |
+
uses throughout: a soft default for a convenience path, and a hard error for the path where silence
|
| 761 |
+
would be a lie.
|
| 762 |
+
|
| 763 |
+
### 4.7 The scratch-root workaround, and why it is documented in the test file
|
| 764 |
+
|
| 765 |
+
The suite does **not** use pytest's built-in `tmp_path`. The module says why, and the reason is a
|
| 766 |
+
sandbox property, not a preference:
|
| 767 |
+
|
| 768 |
+
> *"Scratch root, repo-local. The built-in `tmp_path` fixture cannot finalize under this machine's
|
| 769 |
+
> sandbox, so artifact tests use this instead. It is the same directory `--basetemp=.pytest_tmp` points
|
| 770 |
+
> pytest at, so nothing is written outside the workspace."*
|
| 771 |
+
> — `tests/unit/test_evidence_engine.py`
|
| 772 |
+
|
| 773 |
+
```python
|
| 774 |
+
SCRATCH_ROOT = Path(__file__).resolve().parents[2] / ".pytest_tmp"
|
| 775 |
+
```
|
| 776 |
+
|
| 777 |
+
And the module docstring carries the operational requirement:
|
| 778 |
+
|
| 779 |
+
> *"The `--basetemp=.pytest_tmp` flag is mandatory (see the project handoff): the default pytest temp
|
| 780 |
+
> root triggers a sandbox denial on this machine."*
|
| 781 |
+
|
| 782 |
+
The `scratch` fixture is deliberately **not** cleaned up on teardown, and the docstring says why: *"the
|
| 783 |
+
sandbox's safe-delete guard rejects the recursive delete, and a test that writes into the repo's own
|
| 784 |
+
ignored scratch directory is harmless. Each test gets a unique path so no state leaks between them."*
|
| 785 |
+
This is the same guard that causes §7's four `test_safe_delete_shim` failures — documented here as a
|
| 786 |
+
design constraint rather than hidden as a quirk.
|
| 787 |
+
|
| 788 |
+
### 4.8 The public surface
|
| 789 |
+
|
| 790 |
+
`test_module_public_surface_is_declared` pins the exported names, and
|
| 791 |
+
`test_serialises_through_the_schema_serialiser` pins that the collection serialises through the *schema*
|
| 792 |
+
serialiser (not a bespoke one), so the wire form and the evidence form cannot drift. The shorthand and
|
| 793 |
+
helpers are pinned by `test_aggregate_evidence_shorthand_matches_the_engine`,
|
| 794 |
+
`test_collection_helpers_filter_correctly`, `test_collection_summary_reports_the_observable_facts`,
|
| 795 |
+
`test_digest_changes_when_a_claim_changes`, `test_digest_is_stable_across_regenerated_uuids`,
|
| 796 |
+
`test_collection_is_iterable_and_sized`.
|
| 797 |
+
|
| 798 |
+
`test_digest_is_stable_across_regenerated_uuids` is the one that makes §9 possible: because the digest
|
| 799 |
+
ignores process-local uuids, a recorded run's evidence digest can be recomputed in a *different process*
|
| 800 |
+
and still match. Without it, "recompute the verdict from the record" would be a claim you could not test.
|
| 801 |
+
|
| 802 |
+
### 4.9 What the evidence-engine suite does NOT establish
|
| 803 |
+
|
| 804 |
+
- It does **not** establish that the specialists produce correct evidence — only that the *aggregator*
|
| 805 |
+
is deterministic and lossless given whatever they produce.
|
| 806 |
+
- It does **not** establish end-to-end reproducibility of a live run; that is §8's job.
|
| 807 |
+
- It does **not** establish that the calibration is *good* — the tests pin that calibration is
|
| 808 |
+
*honestly labelled*, not that it improves anything (it does not; ECE got worse).
|
| 809 |
+
|
| 810 |
+
---
|
| 811 |
+
|
| 812 |
+
## 5. The frontend live-wiring suite
|
| 813 |
+
|
| 814 |
+
### 5.1 What it proves, and why it reads source text
|
| 815 |
+
|
| 816 |
+
`tests/unit/test_frontend_live_wiring.py` is the regression net for the frontend↔orchestrator boundary.
|
| 817 |
+
Its docstring states two things it proves, **both of which were broken before the change**:
|
| 818 |
+
|
| 819 |
+
> *"A. **CORS.** The deployment allowed exactly one origin (`https://satquery.pages.dev`), so the
|
| 820 |
+
> frontend could not be driven from a local development server at all: every request from
|
| 821 |
+
> `http://localhost:8080` was refused with `Disallowed CORS origin`, and a developer had to deploy to
|
| 822 |
+
> Cloudflare to test a one-line JavaScript change. The fix adds explicit localhost origins while keeping
|
| 823 |
+
> production listed and keeping `*` rejected.*
|
| 824 |
+
>
|
| 825 |
+
> *B. **The task vocabulary.** The page's router and the server's `Task` enum must agree. They are two
|
| 826 |
+
> independently written vocabularies in two languages, and nothing previously checked that a value the
|
| 827 |
+
> frontend would send is a value the server accepts. This module reads BOTH and compares them."*
|
| 828 |
+
> — `tests/unit/test_frontend_live_wiring.py`
|
| 829 |
+
|
| 830 |
+
And it pins a third, subtler defect — one invisible to a validity check:
|
| 831 |
+
|
| 832 |
+
> *"It also pins the `chang` word-boundary defect: `\bchang\b` cannot match 'changed', so the page's own
|
| 833 |
+
> default question ('What changed here?') routed to `vqa` instead of the change path. That is invisible
|
| 834 |
+
> to any test that only checks 'is this a valid enum value' — both branches produce a valid value. The
|
| 835 |
+
> test therefore asserts the ROUTE, not just the validity."*
|
| 836 |
+
|
| 837 |
+
That last sentence is the philosophical centre of the whole suite: **assert the route, not the
|
| 838 |
+
validity.** A test that asks "is `vqa` a valid task?" passes for both the correct and the defective
|
| 839 |
+
router, because both emit valid tasks. Only a test that asks "does *this question* route to *this
|
| 840 |
+
task*?" can see the defect.
|
| 841 |
+
|
| 842 |
+
**Why source-text parsing.** The module reads the frontend source rather than importing it, and says so:
|
| 843 |
+
|
| 844 |
+
> *"the frontend is plain ES5 with no build step and no module exports, so there is no importable
|
| 845 |
+
> symbol. Parsing the source is the only way to assert the shipped bytes. The parsing is anchored on
|
| 846 |
+
> named tokens rather than line numbers, so it does not silently pass if the file is restructured."*
|
| 847 |
+
|
| 848 |
+
The anchors are module-level path constants — `REPO_ROOT`, `DEPLOY_RENDER`, `FRONTEND_JS`, `MISSION_JS`,
|
| 849 |
+
`LIVE_JS`, `CORE_JS`, `MISSION_HTML` — and the production origin is a single constant,
|
| 850 |
+
`PRODUCTION_ORIGIN = "https://satquery.pages.dev"`. The CORS half imports the orchestrator with a clean
|
| 851 |
+
environment via `_render_module()`, because `_allowed_origins` reads `os.environ` at **call** time while
|
| 852 |
+
`_DEV_ORIGINS` is built at **import** time — a distinction the module's helper comment records
|
| 853 |
+
explicitly.
|
| 854 |
+
|
| 855 |
+
### 5.2 The measured result
|
| 856 |
+
|
| 857 |
+
**106 passed** (`release/CURRENT_RELEASE_STATE.md:121`; `README.md` §The test suites;
|
| 858 |
+
`docs/EVALUATION.md` §7.1). The command is:
|
| 859 |
+
|
| 860 |
+
```bash
|
| 861 |
+
.venv/Scripts/python.exe -m pytest tests/unit/test_frontend_live_wiring.py -q # 106 passed
|
| 862 |
+
```
|
| 863 |
+
|
| 864 |
+
This is a **re-run-this-session** figure, and it is the number this chapter quotes.
|
| 865 |
+
|
| 866 |
+
### 5.3 The fifteen test classes
|
| 867 |
+
|
| 868 |
+
The classes were read from the file. Each name is a claim:
|
| 869 |
+
|
| 870 |
+
| # | Class | What it pins |
|
| 871 |
+
|---|---|---|
|
| 872 |
+
| 1 | `TestTheCorsAllowlistKeepsProductionAndAddsDevelopment` | production origin always allowed, not duplicated, dev origins added, `*` rejected even when hidden in a list |
|
| 873 |
+
| 2 | `TestTheCorsDecisionMatchesTheAllowlist` | the response-leg decision matches the allowlist (echo, no-headers, lookalike-host rejection) |
|
| 874 |
+
| 3 | `TestTheFrontendSpeaksTheServersTaskVocabulary` | every mapped value is a real server task; the mapping covers every task the router can emit; specialist names are human labels, not tasks |
|
| 875 |
+
| 4 | `TestTheDefaultQuestionReachesTheChangePath` | the page's own default question routes to the change path; the change stem matches its inflections; the temporal slot is required for a change question |
|
| 876 |
+
| 5 | `TestALocationQuestionIsNotAChangeQuestion` | a "where" question reads as grounding, with one or two assets |
|
| 877 |
+
| 6 | `TestTheArchitecturePolicyDoesNotReadALocationAsAChange` | the architecture sample routes to grounding; a place word in a "where" question is not a change |
|
| 878 |
+
| 7 | `TestTheTaskRespectsThePairRequirement` | paired tasks substitute correctly when given one asset; single-asset tasks are never substituted |
|
| 879 |
+
| 8 | `TestTheLiveClientUsesTheRealIngestionPath` | the client targets the orchestrator's routes; uploads send raw bytes with a derived content type; the infer body carries asset ids under `assets`; the client never sends a filesystem path |
|
| 880 |
+
| 9 | `TestThePageLoadsTheLiveClient` | `mission.html` includes live JS before mission JS; the declared API base is an absolute origin, not the relative fallback, not the frontend origin; no fixture image is wired into the analysis path |
|
| 881 |
+
| 10 | `TestTheLivePathKeepsTheHonestFailureContract` | a failure is attributed to the step that failed; the preview is kept for the no-file case |
|
| 882 |
+
| 11 | `TestTheCaptionStopsCallingARealUploadIllustrative` | a successful run rewrites the leading word and the alt text; the failure path does not claim an analysis happened; a new file clears the previous run's claim |
|
| 883 |
+
| 12 | `TestTheFrontendSendsOnlyTheAssetsTheTaskRequires` | a single-asset task with a pair uploaded uses only `t1`; the old wiring would have sent both; change/optical-SAR send the pair when present |
|
| 884 |
+
| 13 | `TestDescriptiveQueriesRouteToCaption` | descriptive queries reach caption; a description is not mistaken for change; caption is single-asset |
|
| 885 |
+
| 14 | `TestTheOpticalSarInputValidation` | a real optical+GeoTIFF SAR pair is ok; two plain photos warn; a missing SAR image is an error; the warning names the optical+radar expectation |
|
| 886 |
+
| 15 | `TestTheServerErrorTranslation` | invalid request gets asset guidance; 422 gets malformed guidance; recoverable `false` gets restart guidance; the server's message and detail are preserved; a null error is handled; a recoverable transport code is not called unrecoverable |
|
| 887 |
+
|
| 888 |
+
Two of these deserve a note because they encode *counter-intuitive* decisions:
|
| 889 |
+
|
| 890 |
+
**`TestTheFrontendSendsOnlyTheAssetsTheTaskRequires` (class 12).** The class contains
|
| 891 |
+
`test_the_old_wiring_would_have_sent_both_files_for_vqa` and
|
| 892 |
+
`test_the_fix_sends_one_file_where_the_old_sent_two`. This is a *differential* test: it asserts not just
|
| 893 |
+
that the new code is right, but that the *old* code was wrong in the specific way claimed. That is the
|
| 894 |
+
falsification discipline (§1.3) applied inside a suite — the fix is only meaningful if the defect it
|
| 895 |
+
fixes is real, and the test proves the defect by asserting the old behaviour differs.
|
| 896 |
+
|
| 897 |
+
**`TestTheCaptionStopsCallingARealUploadIllustrative` (class 11).** This is a *copy-honesty* test. The
|
| 898 |
+
page used to caption a real analysis "illustrative"; the suite asserts the leading word and the alt text
|
| 899 |
+
are rewritten on success, and — crucially — that the **failure path does not claim an analysis
|
| 900 |
+
happened** (`test_the_failure_path_does_not_claim_an_analysis_happened`). A UI that says "analysis
|
| 901 |
+
complete" on a failure is a claim defect, and it is tested as one.
|
| 902 |
+
|
| 903 |
+
### 5.4 The historical count: 94 → 100 → 106
|
| 904 |
+
|
| 905 |
+
A reviewer may see a different number. `docs/FINAL_DELIVERY_REPORT.md §7` records **94 passed** for this
|
| 906 |
+
file, while the current figure is **106**. The difference is the suite growing after the delivery report
|
| 907 |
+
was written: the regression tests added for the router fix took this file from **100 → 106** tests
|
| 908 |
+
(`DELIVERY_REPORT_2026-09-25.md`, regression-tests note). This chapter quotes **106** and cites both
|
| 909 |
+
figures so neither looks like a contradiction (`docs/REPRODUCIBILITY.md` §5.2).
|
| 910 |
+
|
| 911 |
+
### 5.5 What the frontend live-wiring suite does NOT establish
|
| 912 |
+
|
| 913 |
+
- It does **not** run the browser. It asserts the *shipped bytes* of the frontend against the
|
| 914 |
+
*orchestrator module* in-process. Whether the page actually loads and runs is §8's job (live
|
| 915 |
+
validation).
|
| 916 |
+
- It does **not** prove the router's *quality* — only that specific questions route to the specific
|
| 917 |
+
tasks asserted. The router's residual misroutes (`"What is the new runway?"` → `change`) are real and
|
| 918 |
+
recorded (`docs/LIMITATIONS.md`; `LIVE_VALIDATION_POSTFIX.md`), not hidden by a green suite.
|
| 919 |
+
- It does **not** test the inference service. It tests the boundary the frontend speaks to, which is the
|
| 920 |
+
orchestrator.
|
| 921 |
+
|
| 922 |
+
---
|
| 923 |
+
|
| 924 |
+
## 6. The doc/frontend suite
|
| 925 |
+
|
| 926 |
+
### 6.1 The five files, and the measured result
|
| 927 |
+
|
| 928 |
+
The "doc/frontend suite" is a named grouping of **five files** that together report **183 passed**
|
| 929 |
+
(`docs/FINAL_DELIVERY_REPORT.md §7`; `docs/EVALUATION.md` §7.2; `docs/REPRODUCIBILITY.md` §5.3). The
|
| 930 |
+
five are:
|
| 931 |
+
|
| 932 |
+
| File | What it pins |
|
| 933 |
+
|---|---|
|
| 934 |
+
| `tests/unit/test_frontend_guide_doc.py` | both documents agree on tasks, coordinate systems, evidence types, geometry shapes; no login screen; no secrets in the browser; the confidence-property trap; serialized analyses |
|
| 935 |
+
| `tests/unit/test_frontend_live_wiring.py` | the frontend↔orchestrator boundary — see §5 |
|
| 936 |
+
| `tests/unit/test_api_contract_doc.py` | every JSON block parses and validates; no contract example fabricates an artifact URI; documented task/coordinate-system values are the real ones; every error code documented; the frozen config hash is recorded; the calibration caveat is recorded |
|
| 937 |
+
| `tests/unit/test_runbook_doc.py` | the runbook names only public serving entrypoints; quoted artifact sizes match the real files; the capability list matches the registry; env vars are the architecture ones; undecided items declared undecided; **no claim that a deployment happened**; verified distinguished from designed |
|
| 938 |
+
| `tests/unit/test_deploy_config.py` | the deploy manifest declares the non-registry marker; the validator reports no errors; the **config-hash regression guard**; the C-8 constraints (`torch_compile` false, cpu mode required, lazy load, single model cache) |
|
| 939 |
+
|
| 940 |
+
These are **documentation conformance tests**: they fail if a doc drifts from the code it describes.
|
| 941 |
+
They are the mechanism that keeps this release's documentation honest — the reason a doc claim can be
|
| 942 |
+
trusted is that a test would go red if the claim stopped matching the code.
|
| 943 |
+
|
| 944 |
+
### 6.2 Why four of the five are doc-guards, and one is not
|
| 945 |
+
|
| 946 |
+
Four of the five (`test_frontend_guide_doc`, `test_api_contract_doc`, `test_runbook_doc`,
|
| 947 |
+
`test_deploy_config`) are doc-guards, analysed individually in §3. The fifth
|
| 948 |
+
(`test_frontend_live_wiring`) is not a doc-guard — it is a code-behaviour suite that happens to be
|
| 949 |
+
grouped here because it reads frontend source text. The grouping is by *audience* (frontend +
|
| 950 |
+
documentation) rather than by mechanism, and that is why the count 183 covers both.
|
| 951 |
+
|
| 952 |
+
### 6.3 The doc-guard guarantee, restated
|
| 953 |
+
|
| 954 |
+
The doc-guards guarantee **example-level and value-level** conformance, not prose truth (§3.7). The two
|
| 955 |
+
module docstrings state the limit themselves:
|
| 956 |
+
|
| 957 |
+
> *"These tests validate the EXAMPLES, not the prose. They cannot detect a wrong sentence (e.g. a
|
| 958 |
+
> mis-stated latency budget)."* — `tests/unit/test_api_contract_doc.py`
|
| 959 |
+
|
| 960 |
+
> *"Prose can still be wrong in ways these tests cannot see -- a mis-stated latency, a wrong reason for a
|
| 961 |
+
> decision."* — `tests/unit/test_frontend_guide_doc.py`
|
| 962 |
+
|
| 963 |
+
`test_runbook_doc.py` reaches a different class of claim — numbers, names, and **status honesty** — and
|
| 964 |
+
its most important assertion is that the document itself *labels its unimplemented sections as
|
| 965 |
+
unimplemented* (§3.4). That is the guard against the exact overclaim the style guide forbids.
|
| 966 |
+
|
| 967 |
+
### 6.4 The config-hash regression guard
|
| 968 |
+
|
| 969 |
+
`test_deploy_config.py` carries the anti-drift move that makes it more than a validator: it does **not**
|
| 970 |
+
re-declare the frozen config hash. It imports it:
|
| 971 |
+
|
| 972 |
+
> *"We import it rather than re-declare it, so the two cannot drift."*
|
| 973 |
+
> — `tests/unit/test_deploy_config.py`
|
| 974 |
+
|
| 975 |
+
The authority is `tests.test_config`, and the value is the frozen hash `78f1e3700da15aa1`
|
| 976 |
+
(`DOCS_STYLE_GUIDE.md` §3). A second copy of the hash anywhere in the tree would be a latent
|
| 977 |
+
contradiction; the import makes drift structurally impossible.
|
| 978 |
+
|
| 979 |
+
The same file proves the manifest is **INERT** — *"it is never merged into the config registry and
|
| 980 |
+
cannot move `Config.hash`"* — and that negative cases use in-memory dicts via `validate_documents` *"so
|
| 981 |
+
the real config files are never mutated."* It also drives the validator CLI **in a child interpreter**
|
| 982 |
+
to prove the module's own `sys.path` bootstrap makes `core` importable *"WITHOUT pytest's `pythonpath =
|
| 983 |
+
.` shortcut"* (`test_validator_cli_exits_zero_as_a_standalone_script`).
|
| 984 |
+
|
| 985 |
+
### 6.5 What the doc/frontend suite does NOT establish
|
| 986 |
+
|
| 987 |
+
- It does **not** establish that any deployment exists. `test_runbook_doc.py` explicitly asserts the
|
| 988 |
+
opposite direction: that the document does **not** claim a deployment happened (§3.4).
|
| 989 |
+
- It does **not** validate prose (§6.3).
|
| 990 |
+
- It does **not** check that the *code* is correct — only that the docs match the code. If the code and
|
| 991 |
+
the docs are wrong together, a doc-guard is green.
|
| 992 |
+
|
| 993 |
+
---
|
| 994 |
+
|
| 995 |
+
## 7. The full `tests/unit` result, and the 5–6 environmental failures
|
| 996 |
+
|
| 997 |
+
This is the section where the release's honesty discipline is most visible. Running the entire unit tree
|
| 998 |
+
produces failures. They are reported, classified, and *not* reclassified as passes.
|
| 999 |
+
|
| 1000 |
+
### 7.1 The result
|
| 1001 |
+
|
| 1002 |
+
Running `python -m pytest tests/unit` in the authoring sandbox produces **5–6 failures**. Every one is
|
| 1003 |
+
**environmental or ordering**-related, not a regression in shipped code. The classification, exactly as
|
| 1004 |
+
`docs/FINAL_DELIVERY_REPORT.md §7` records it:
|
| 1005 |
+
|
| 1006 |
+
| # | Failure | Attribution | Regression? |
|
| 1007 |
+
|---|---|---|---|
|
| 1008 |
+
| 1–4 | `test_safe_delete_shim` — **4 failures** | the sandbox's bulk-**delete guard** (Windows verbatim-path behaviour) | **No** |
|
| 1009 |
+
| 5 | one **ordering flake** in the router route test | test **ordering** (passes in isolation) | **No** |
|
| 1010 |
+
| 6 | one **stale adapter test** (`optical_sar` absent when CROMA unshipped) | stale test — **CROMA is now shipped** | **No** |
|
| 1011 |
+
|
| 1012 |
+
### 7.2 The evidence that they are not regressions
|
| 1013 |
+
|
| 1014 |
+
The evidence is the **re-run**. Re-running the affected files **together** gives **137 passed**
|
| 1015 |
+
(`docs/FINAL_DELIVERY_REPORT.md §7`; `docs/EVALUATION.md` §7.3–7.4). The logic is stated plainly:
|
| 1016 |
+
|
| 1017 |
+
> *"if the failures were real regressions in the code under test, re-running those files together would
|
| 1018 |
+
> still fail. They do not — which is what distinguishes an environmental/ordering failure from a
|
| 1019 |
+
> regression."*
|
| 1020 |
+
|
| 1021 |
+
This is the falsification discipline (§1.3) applied to a *test result* rather than to a *guard*: the
|
| 1022 |
+
claim "these are not regressions" is itself tested, by a re-run designed to fail if it were false.
|
| 1023 |
+
|
| 1024 |
+
### 7.3 The four `test_safe_delete_shim` failures, in detail
|
| 1025 |
+
|
| 1026 |
+
`tests/unit/test_safe_delete_shim.py` (338 lines, `pytestmark = pytest.mark.unit`) guards two defects in
|
| 1027 |
+
the WorkBuddy Windows safe-delete shim (`cli/vendor/shim/sitecustomize.py`) that produced **phantom test
|
| 1028 |
+
failures** in this repository. The module's docstring records both with measurements.
|
| 1029 |
+
|
| 1030 |
+
**Defect 1 — a verbatim temp path was not recognised as a temp path.**
|
| 1031 |
+
|
| 1032 |
+
> *"`_path_for_compare` returned `normcase(realpath(abspath(path)))`, which preserves the Windows verbatim
|
| 1033 |
+
> prefix `\\?\`. `os.path.relpath` cannot relate `\\?\c:\...` to `c:\...`, so `_is_under_root` returned
|
| 1034 |
+
> False and the OS-temp exemption in `_should_bypass_safe_delete` did not apply."*
|
| 1035 |
+
|
| 1036 |
+
Measured, before the fix:
|
| 1037 |
+
|
| 1038 |
+
```
|
| 1039 |
+
plain temp subdir -> bypass True
|
| 1040 |
+
verbatim temp subdir -> bypass False <-- the bug
|
| 1041 |
+
non-temp subdir -> bypass False
|
| 1042 |
+
```
|
| 1043 |
+
|
| 1044 |
+
That is why routine pytest `garbage-*` collection — which walks
|
| 1045 |
+
`\\?\C:\Users\...\Temp\pytest-of-anish\garbage-*` — reached the bulk guard, tripped `confirmRequired` at
|
| 1046 |
+
**69 entries against a threshold of 50**, and **latched a rejection that then blocked every delete in
|
| 1047 |
+
the conversation**.
|
| 1048 |
+
|
| 1049 |
+
**Defect 2 — a successful delete was reported as a failure.**
|
| 1050 |
+
|
| 1051 |
+
> *"`_platform_trash` raised whenever `SHFileOperationW` returned non-zero."*
|
| 1052 |
+
|
| 1053 |
+
Measured on the authoring host with `FOF_ALLOWUNDO|FOF_NOCONFIRMATION|FOF_NOERRORUI|FOF_SILENT`, calling
|
| 1054 |
+
shell32 directly (the shim not involved):
|
| 1055 |
+
|
| 1056 |
+
```
|
| 1057 |
+
series non-temp rc temp rc
|
| 1058 |
+
4 regions x 3 interleaved rounds 2, 2, 2, ... 0, 0, 0
|
| 1059 |
+
6 processes x 20 deletes 2 (120/120) 0 (120/120)
|
| 1060 |
+
```
|
| 1061 |
+
|
| 1062 |
+
The target is removed in **every** case and the user's Recycle Bin is populated and active (1,300+
|
| 1063 |
+
`$I`/`$R` entries), so the return code is **not** a reliable "could not delete" signal. The module is
|
| 1064 |
+
explicit that the code is **intermittent** and that the precise Windows-internal trigger is not known:
|
| 1065 |
+
|
| 1066 |
+
> *"across pytest invocations the same non-temp delete sometimes returned 0, and in one run a *temp*
|
| 1067 |
+
> delete returned non-zero. The precise Windows-internal trigger is **NOT established** and is not
|
| 1068 |
+
> claimed here."*
|
| 1069 |
+
|
| 1070 |
+
This is a place where this chapter writes **`UNKNOWN — not established from the available evidence`**,
|
| 1071 |
+
following the module's own honesty.
|
| 1072 |
+
|
| 1073 |
+
**What the guard asserts** (the `WHAT IS ASSERTED` block of the docstring):
|
| 1074 |
+
|
| 1075 |
+
- A verbatim-prefixed path inside the OS temp root compares equal to its plain form and is exempted.
|
| 1076 |
+
- A verbatim-prefixed path *outside* the temp root is still guarded — *"the fix normalises the prefix, it
|
| 1077 |
+
does not widen the exemption."*
|
| 1078 |
+
- Deleting through a verbatim temp path succeeds end to end.
|
| 1079 |
+
- A non-zero shell code on a delete whose target is genuinely gone does not raise (stubbed shell, stubbed
|
| 1080 |
+
`lexists` — deterministic, deletes nothing).
|
| 1081 |
+
- A non-zero shell code on a delete whose target **survives** still raises — *"Fail-closed is preserved;
|
| 1082 |
+
only the false positive is gone."*
|
| 1083 |
+
- Against the real shell, a non-temp delete does not raise (behavioural confirmation, not the guard).
|
| 1084 |
+
|
| 1085 |
+
**Why the guard is stub-based.** The module states the reason, and it is a direct consequence of defect
|
| 1086 |
+
2's intermittency:
|
| 1087 |
+
|
| 1088 |
+
> *"the guard for this defect stubs the shell: a test that depends on the real return code would
|
| 1089 |
+
> sometimes pass against the broken shim. The stub tests fail on the original shim on every run."*
|
| 1090 |
+
|
| 1091 |
+
That is the difference between a test that *sometimes* catches a defect and a test that *always* does.
|
| 1092 |
+
|
| 1093 |
+
**The rule for extending the file** is stated in the docstring and is worth reproducing, because it is
|
| 1094 |
+
the exact failure this file exists to prevent:
|
| 1095 |
+
|
| 1096 |
+
> *"**Do not perform a real delete outside the OS temp root.** … An earlier draft of the defect-2 test
|
| 1097 |
+
> did exactly that, tripped `confirmRequired` at the threshold and latched a rejection that blocked every
|
| 1098 |
+
> subsequent delete in the session — the very failure this file exists to prevent."*
|
| 1099 |
+
|
| 1100 |
+
The module also **skips cleanly** when the shim is absent or disabled, via
|
| 1101 |
+
`pytest.importorskip("sitecustomize", ...)` plus `skipif` markers (`needs_shim_helpers`, `windows_only`).
|
| 1102 |
+
So on a normal user environment these tests skip; in the authoring sandbox they reach the guard and 4 of
|
| 1103 |
+
the 8 fail. The constants involved are `VERBATIM = "\\\\?\\"` and a `NON_TEMP` probe path that is *"never
|
| 1104 |
+
created"* — the predicates tested are pure.
|
| 1105 |
+
|
| 1106 |
+
### 7.4 The ordering flake
|
| 1107 |
+
|
| 1108 |
+
One router route test fails only when the full suite runs in a particular order, and **passes in
|
| 1109 |
+
isolation**. That is the definition of an ordering flake: the failure is a function of collection order,
|
| 1110 |
+
not of the code. `docs/REPRODUCIBILITY.md` §5.4 records it in the same row-set as the shim failures.
|
| 1111 |
+
|
| 1112 |
+
### 7.5 The stale adapter test
|
| 1113 |
+
|
| 1114 |
+
One test asserts `optical_sar` is **absent when CROMA is unshipped**. CROMA is now shipped
|
| 1115 |
+
(`CURRENT_RELEASE_STATE.md` §1 capability contract: `optical_sar` `available: true`), so the assertion's
|
| 1116 |
+
premise has changed. This is a **stale test**, not a code defect: the test encodes an assumption about
|
| 1117 |
+
the environment that the environment has outgrown.
|
| 1118 |
+
|
| 1119 |
+
### 7.6 A different environment state: the 2,237-passed run
|
| 1120 |
+
|
| 1121 |
+
`docs/PHASE12_115_METRIC_COMPUTED.md` §8 records, for **2026-09-22**, a full unit suite result of
|
| 1122 |
+
**2,237 passed, 0 failed (622.68 s)**. That is a real, measured result in a *different* environment state
|
| 1123 |
+
(the safe-delete shim's guard state and the CROMA shipping status differ). It is recorded rather than
|
| 1124 |
+
suppressed, *"because the two results are both true and the difference is exactly the
|
| 1125 |
+
environmental/ordering story"* (`docs/EVALUATION.md` §7.3).
|
| 1126 |
+
|
| 1127 |
+
The same document records a correction that is itself an honesty lesson: an earlier revision claimed
|
| 1128 |
+
`2,179 passed` computed as `2,171 + 8`, *"presented as a check but the total was never measured — it was
|
| 1129 |
+
inferred from a stale baseline."* The measured figure was **2,237**. This is the style guide's rule
|
| 1130 |
+
(`DOCS_STYLE_GUIDE.md` §1.3) in the wild: an inferred number presented as a measurement was corrected to
|
| 1131 |
+
the measured one.
|
| 1132 |
+
|
| 1133 |
+
### 7.7 The UNKNOWN: the collected count
|
| 1134 |
+
|
| 1135 |
+
The exact **collected** test count for the current-session full-suite run is
|
| 1136 |
+
**`UNKNOWN — not established from the available evidence`**. The record gives the failure classification
|
| 1137 |
+
and the 137-passed re-run, but not a collected total. Per-suite counts are known (106, 183, 137, 73, 51);
|
| 1138 |
+
the single collected total is not.
|
| 1139 |
+
|
| 1140 |
+
### 7.8 Why this is reported rather than hidden
|
| 1141 |
+
|
| 1142 |
+
The style guide's first rule is `DO NOT FABRICATE`, and its status rule is *"a mixed result is never 'all
|
| 1143 |
+
work perfectly'"* (`DOCS_STYLE_GUIDE.md` §1.7). A full-suite run with 5–6 failures is a mixed result.
|
| 1144 |
+
Presenting only the 106 and 183 would be exactly the "all work perfectly" overclaim the guide forbids.
|
| 1145 |
+
The correct presentation is the three-row table (§7.1) plus the re-run evidence (§7.2) plus the explicit
|
| 1146 |
+
UNKNOWN (§7.7) — which is what this section is.
|
| 1147 |
+
|
| 1148 |
+
---
|
| 1149 |
+
|
| 1150 |
+
## 8. The live-validation harness and its integrity discipline
|
| 1151 |
+
|
| 1152 |
+
### 8.1 What live validation is, and why it is separate from the suites
|
| 1153 |
+
|
| 1154 |
+
Every suite in §2–§7 runs in-process. `docs/API_CONTRACT.md` §8 states the boundary plainly: *"no test in
|
| 1155 |
+
this repository dials a network address."* A green `tests/integration` run is therefore **not** evidence
|
| 1156 |
+
about a deployment. **Live validation** is the separate activity that drives the *deployed* stack — the
|
| 1157 |
+
real Cloudflare Pages frontend, the real Render orchestrator, the real tunnel, the real Codespace — with
|
| 1158 |
+
a browser, and records what happened.
|
| 1159 |
+
|
| 1160 |
+
The record lives at `.workbuddy-ai/scratch/live_validation/` — `LIVE_VALIDATION_POSTFIX.md`, the raw
|
| 1161 |
+
harness outputs (`run_output.txt`, `run_final2.txt`, `run_final3.txt`), the recomputed verdicts
|
| 1162 |
+
(`results_final.json`, `results_pass3.json`), and eight per-case screenshots (`A1_vqa.png` … `B2_newairport.png`).
|
| 1163 |
+
|
| 1164 |
+
### 8.2 The three passes, summarised
|
| 1165 |
+
|
| 1166 |
+
| Pass | HEAD | Harness | Result |
|
| 1167 |
+
|---|---|---|---|
|
| 1168 |
+
| 1 | `ff46eba42b18` + `d413d3672311` | v1 (`fill_input`) | 8/8 |
|
| 1169 |
+
| 2 | `2d7ae53b482d` | v2 asserting | 8/8 |
|
| 1170 |
+
| 3 | `2d7ae53b482d` | v2 asserting (pre-discriminator-fix) | 8/8 (recomputed) |
|
| 1171 |
+
|
| 1172 |
+
**24 live runs, 24 correct dispatches, 0 mock nodes, no run id repeated across passes.** The trace fill
|
| 1173 |
+
was **94.4444 %** on every case, and every case carries a real `run_*` id, the Hugging Face link in the
|
| 1174 |
+
DOM, and the `capabilities → assets → infer` call sequence all addressed to
|
| 1175 |
+
`satquery-backend-m4yv.onrender.com` (two `assets` calls for the pair tasks).
|
| 1176 |
+
|
| 1177 |
+
The eight cases are Phase A (six regression cases: `vqa`, `caption`, `grounding`, `change`,
|
| 1178 |
+
`change_vqa`, `optical_sar`) and Phase B (the two defect cases: `"Where are the built-up areas in this
|
| 1179 |
+
image?"` and `"Where is the new airport?"`, both expected to reach `grounding` with a **single** asset —
|
| 1180 |
+
the exact condition under which the old router collapsed to `vqa` and answered "River").
|
| 1181 |
+
|
| 1182 |
+
### 8.3 The CDP synthetic-key-event trap
|
| 1183 |
+
|
| 1184 |
+
This is the integrity lesson that shaped the harness, and it is the most transferable finding in this
|
| 1185 |
+
chapter.
|
| 1186 |
+
|
| 1187 |
+
**The shape of the false pass.** Pass 1's harness drove the query box with `fill_input()`, which types
|
| 1188 |
+
with **real CDP key events**. A re-run attempt failed on case 1 with `run_id=0002`, `mock_nodes=9`,
|
| 1189 |
+
`answer="No answer yet"`, and only a `capabilities` call — the **mock path**. The root cause, measured
|
| 1190 |
+
directly:
|
| 1191 |
+
|
| 1192 |
+
> *"`fill_input` types with **real CDP key events**, and Chrome **drops synthesized key events when the
|
| 1193 |
+
> browser window does not hold OS focus**; it has **no assertion**, so it clicked Run with the page's
|
| 1194 |
+
> default query still in the box. Measured directly: with Chrome backgrounded, `press_key("Z")` left
|
| 1195 |
+
> `#qtext.value` unchanged, while `type_text("Q")` (CDP `Input.insertText`, not focus-gated) inserted
|
| 1196 |
+
> fine."*
|
| 1197 |
+
> — `LIVE_VALIDATION_POSTFIX.md`
|
| 1198 |
+
|
| 1199 |
+
The measurement is recorded in `diag_focus.harness`, which prints the value of `#qtext` after
|
| 1200 |
+
`press_key("Z")` and after `type_text("Q")` — showing the first is a no-op and the second works.
|
| 1201 |
+
|
| 1202 |
+
**The generalisation.** A harness that types into a form and then reads a *result* cannot tell the
|
| 1203 |
+
difference between:
|
| 1204 |
+
|
| 1205 |
+
- "the query was entered, and the run used it", and
|
| 1206 |
+
- "the query was never entered, and the run used the page's default".
|
| 1207 |
+
|
| 1208 |
+
Both produce a result. The result of the second is a *plausible* answer to a *different* question, and a
|
| 1209 |
+
harness without an assertion records it as the verdict for the first. That is a **silent false pass**,
|
| 1210 |
+
and it is the reason the fix is not "use a different typing helper" but "**assert the form state before
|
| 1211 |
+
dispatch**".
|
| 1212 |
+
|
| 1213 |
+
### 8.4 The fix: deterministic query entry plus pre-dispatch assertions
|
| 1214 |
+
|
| 1215 |
+
The harness was rebuilt as `run_all_postfix2.harness`, whose header states the change:
|
| 1216 |
+
|
| 1217 |
+
> *"v2 sets the query deterministically (js value + input event, falling back to `Input.insertText`) and
|
| 1218 |
+
> ASSERTS the form state before clicking, recording `q_ok` / `obs_ok` / `t0_ok` and a computed verdict
|
| 1219 |
+
> per case."*
|
| 1220 |
+
|
| 1221 |
+
The three assertions, and the failure each prevents:
|
| 1222 |
+
|
| 1223 |
+
| Assertion | Checks | Failure it prevents |
|
| 1224 |
+
|---|---|---|
|
| 1225 |
+
| `q_ok` | `#qtext.value` holds the intended query | the silent-drop false pass |
|
| 1226 |
+
| `obs_ok` | `#obsTail == 'ready'` **and** one file on `#fileInput` | an upload that did not land |
|
| 1227 |
+
| `t0_ok` | both frames present, for pair tasks | a pair task run on one asset |
|
| 1228 |
+
| `no_mock_nodes` | `mock_nodes == 0` | the preview path being recorded as live |
|
| 1229 |
+
|
| 1230 |
+
The rule, stated in the session handoff (`HANDOFF_NEXT_AGENT.md` §5.2) and reproduced in
|
| 1231 |
+
`docs/architecture/09-frontend.md` §10.3:
|
| 1232 |
+
|
| 1233 |
+
> *"**Do NOT use `fill_input()` or `press_key()` to enter the query.** They type with real CDP key events,
|
| 1234 |
+
> which Chrome **silently drops when the browser window does not hold OS focus** … Use `js()` to set
|
| 1235 |
+
> `#qtext.value` (plus `input`/`change` events) and/or `type_text()` (CDP `Input.insertText`, not
|
| 1236 |
+
> focus-gated). **Always assert the form state before clicking Run** … or a no-op will be recorded as a
|
| 1237 |
+
> pass."*
|
| 1238 |
+
|
| 1239 |
+
The `set_query()` function in the harness implements exactly this: it focuses the box, clears it, tries
|
| 1240 |
+
`type_text(q)` (the non-focus-gated CDP path), reads the value back, and **falls back** to a direct
|
| 1241 |
+
`js` set with `input`/`change` events if the read-back does not match. It returns the value actually read
|
| 1242 |
+
back, and `q_ok = (qgot == c["q"])`. The assertion is on the *read-back*, not on the *attempt*.
|
| 1243 |
+
|
| 1244 |
+
`obs_ok`'s second half is a real DOM fact: `#obsTail` reads `none` in the markup (`mission.html:67`) and
|
| 1245 |
+
the live driver flips it to `ready` on upload, so `obs_ok` proves the upload landed rather than merely
|
| 1246 |
+
that a click happened.
|
| 1247 |
+
|
| 1248 |
+
### 8.5 The three discriminators that proved the earlier 8/8 was clean
|
| 1249 |
+
|
| 1250 |
+
The response to a silent-false-pass risk is not to re-run and hope; it is to check whether the *earlier*
|
| 1251 |
+
result was contaminated, using evidence the harness recorded. The report does exactly that:
|
| 1252 |
+
|
| 1253 |
+
> *"I then checked whether the earlier 8/8 run was infected by the same silent failure. It was not:*
|
| 1254 |
+
>
|
| 1255 |
+
> * *its recorded intents are **query-specific** — A1 reads `taskvqa…temporalnone`, whereas the default
|
| 1256 |
+
> query "What changed here?" would read `taskchange…temporalrequired` (exactly what the failed run
|
| 1257 |
+
> showed);*
|
| 1258 |
+
> * *its answers **embed the query text** — e.g. `[grounding] Located 6 candidate region(s) for 'Where
|
| 1259 |
+
> are the built-up areas in this image?'`;*
|
| 1260 |
+
> * *A6 required two files (`optical 4/12 + SAR 2/2` channels), which only the uploaded pair supplies.*
|
| 1261 |
+
>
|
| 1262 |
+
> *So the 8/8 result is a valid measurement."*
|
| 1263 |
+
> — `DELIVERY_REPORT_2026-09-25.md` §3
|
| 1264 |
+
|
| 1265 |
+
The three discriminators generalise into a reusable checklist:
|
| 1266 |
+
|
| 1267 |
+
| Discriminator | What it proves |
|
| 1268 |
+
|---|---|
|
| 1269 |
+
| the recorded **intent** is query-specific | the query reached the router |
|
| 1270 |
+
| the **answer** embeds the query text | the server received the intended query |
|
| 1271 |
+
| a case **requires an artefact** only the setup supplies | the setup really happened |
|
| 1272 |
+
|
| 1273 |
+
This is why the harness records the full per-case payload — `intent`, `answer`, `files_t1`, `files_t0`,
|
| 1274 |
+
`obs_tail`, `mock_nodes`, `api_calls` — rather than just a pass/fail bit. The bit can be wrong; the raw
|
| 1275 |
+
record can be re-examined.
|
| 1276 |
+
|
| 1277 |
+
### 8.6 The two discriminator bugs (and why they were false *failures*, not false passes)
|
| 1278 |
+
|
| 1279 |
+
Two harness bugs were found and fixed, both causing **false failures** — the safe direction:
|
| 1280 |
+
|
| 1281 |
+
1. **The answer `[task]` tag exists only for region tasks.** `vqa`/`caption` answers are bare
|
| 1282 |
+
("Grassland", prose), so a tag-based verdict heuristic reported them as failures. The fix computes the
|
| 1283 |
+
dispatched task as `answer_tag` when present, else the intent's reading.
|
| 1284 |
+
2. **The intent panel is a concatenated string.** `task([a-z_]+)` must be matched **non-greedily** up to
|
| 1285 |
+
`modality`; a greedy match swallows the whole string. The fix is
|
| 1286 |
+
`re.search(r"task([a-z_]+?)modality", intent)`.
|
| 1287 |
+
|
| 1288 |
+
The harness header states the asymmetry: *"Two further harness bugs were found and fixed, both causing
|
| 1289 |
+
**false failures**."* A false failure costs a re-run; a false pass costs the integrity of the result. The
|
| 1290 |
+
harness was built so that its bugs err toward false failure, and §9 makes even the false-failure case
|
| 1291 |
+
recoverable without a re-run.
|
| 1292 |
+
|
| 1293 |
+
### 8.7 Liveness traps discovered during the runs
|
| 1294 |
+
|
| 1295 |
+
Two operational traps are recorded because they look exactly like stalls:
|
| 1296 |
+
|
| 1297 |
+
- **`browser-use` block-buffers stdout even when redirected to a file**, so the output file stays at
|
| 1298 |
+
**0 bytes until the process exits** — indistinguishable from a stall. The harness now forces
|
| 1299 |
+
line-buffering (`sys.stdout.reconfigure(line_buffering=True)`) and prints `CASE_START <id>` per case,
|
| 1300 |
+
so the file grows case by case.
|
| 1301 |
+
- **`grep` block-buffers when piped**, so piping the harness through `grep` swallows all output if the
|
| 1302 |
+
pipeline is killed. Redirect to a file instead.
|
| 1303 |
+
|
| 1304 |
+
The harness also records `ALL_DONE` as a final marker, so a truncated run is detectable from the output
|
| 1305 |
+
file alone.
|
| 1306 |
+
|
| 1307 |
+
### 8.8 What the live validation does NOT establish
|
| 1308 |
+
|
| 1309 |
+
- It is **not** the independent audit. `LIVE_VALIDATION_POSTFIX.md` opens by saying so: *"This is a
|
| 1310 |
+
*re-validation*, not the independent audit: the prior independent audit is
|
| 1311 |
+
`../2026-09-25-00-37-41/live_validation/INDEPENDENT_LIVE_VALIDATION.md`."*
|
| 1312 |
+
- It does **not** establish model *quality*. The two verdicts are kept separate: **Deployment/integration:
|
| 1313 |
+
PASS** and **Model quality: MIXED** (caption and grounding meaningful; change/change_vqa plausible; VQA
|
| 1314 |
+
weak-but-related — A1 answers "Grassland"; optical-SAR returns a bare class index `class_18`, not a
|
| 1315 |
+
human label).
|
| 1316 |
+
- It does **not** establish load behaviour, latency ceilings, or cold-start times. It is eight cases, run
|
| 1317 |
+
three times.
|
| 1318 |
+
- It does **not** clear B-07: transient tunnel-agent gaps remain **OPEN**, and the patch is **prepared,
|
| 1319 |
+
NOT deployed** (`CURRENT_RELEASE_STATE.md` §6).
|
| 1320 |
+
|
| 1321 |
+
---
|
| 1322 |
+
|
| 1323 |
+
## 9. Verdict-independent evidence: recomputing from raw records
|
| 1324 |
+
|
| 1325 |
+
### 9.1 The property
|
| 1326 |
+
|
| 1327 |
+
The live-validation harness has a `verdict` field. The project's design decision is that **the verdict
|
| 1328 |
+
must not be the primary record.** The recorded evidence — `q_ok`, `obs_ok`, `t0_ok`, `run_id`,
|
| 1329 |
+
`mock_nodes`, `intent`, `answer`, `api_calls` — is recorded *independently of* the verdict computation,
|
| 1330 |
+
so a harness bug in the verdict logic can never silently turn a real failure into a pass, and can be
|
| 1331 |
+
corrected without a re-run.
|
| 1332 |
+
|
| 1333 |
+
### 9.2 The recomputation script
|
| 1334 |
+
|
| 1335 |
+
`.workbuddy-ai/scratch/recompute_verdicts.py` implements the property. Its docstring states the case that
|
| 1336 |
+
motivated it:
|
| 1337 |
+
|
| 1338 |
+
> *"The v2 harness recorded, per case: `q_ok`, `obs_ok`, `t0_ok`, `run_id`, `mock_nodes`, `intent`,
|
| 1339 |
+
> `answer`, `api_calls` -- all independent of the verdict computation. Its verdict field was wrong for
|
| 1340 |
+
> vqa/caption because the tag heuristic assumed every answer starts with `"[task]"`; vqa/caption answers
|
| 1341 |
+
> are bare. This recomputes the verdict from the intent's `task<name>` slot (primary) plus the answer tag
|
| 1342 |
+
> (cross-check only when present)."*
|
| 1343 |
+
|
| 1344 |
+
The script is read-only over the recorded run and writes `results_final.json`. The core of it:
|
| 1345 |
+
|
| 1346 |
+
```python
|
| 1347 |
+
mi = re.search(r"task([a-z_]+?)modality", intent)
|
| 1348 |
+
intent_task = mi.group(1) if mi else ""
|
| 1349 |
+
tag = answer[1:answer.find("]")] if (answer.startswith("[") and "]" in answer) else ""
|
| 1350 |
+
expect = d["expect"]
|
| 1351 |
+
|
| 1352 |
+
# The DISPATCHED task is what the server actually ran. The answer carries a "[task]"
|
| 1353 |
+
# prefix for the region tasks (grounding/change/change_vqa/optical_sar); vqa and
|
| 1354 |
+
# caption answers are bare, so fall back to the intent's reading there.
|
| 1355 |
+
# NOTE: A5 is a legitimate case where the two differ -- the router READS "change" but
|
| 1356 |
+
# dispatches "change_vqa" (the documented quantifier upgrade). That is not a failure.
|
| 1357 |
+
dispatched = tag if tag else intent_task
|
| 1358 |
+
|
| 1359 |
+
checks = {
|
| 1360 |
+
"query_in_box": bool(d.get("q_ok")),
|
| 1361 |
+
"asset_attached": bool(d.get("obs_ok")),
|
| 1362 |
+
"pair_attached": bool(d.get("t0_ok")),
|
| 1363 |
+
"live_run": str(d.get("run_id", "")).startswith("run_"),
|
| 1364 |
+
"no_mock_nodes": d.get("mock_nodes") == 0,
|
| 1365 |
+
"task_dispatched": dispatched == expect,
|
| 1366 |
+
}
|
| 1367 |
+
verdict = "PASS" if all(checks.values()) else "FAIL"
|
| 1368 |
+
```
|
| 1369 |
+
|
| 1370 |
+
Every input to the verdict is a *recorded observation*; the verdict is a pure function of them. The
|
| 1371 |
+
script also records `read_vs_dispatched_differs` — so the one legitimate read≠dispatch case (A5, the
|
| 1372 |
+
documented quantifier upgrade) is **flagged, not failed**.
|
| 1373 |
+
|
| 1374 |
+
### 9.3 Pass 3 is the proof of the property
|
| 1375 |
+
|
| 1376 |
+
Pass 3 was launched with the harness build that still carried the two discriminator bugs (§8.6), so its
|
| 1377 |
+
raw output says **`SUMMARY 0/8`**. The verdicts were recomputed from the recorded evidence by
|
| 1378 |
+
`recompute_verdicts.py` and are **8/8 PASS** (`results_pass3.json`). The record states the significance:
|
| 1379 |
+
|
| 1380 |
+
> *"That is the intended safety property — **the recorded evidence (run id, `mock_nodes`, intent, answer,
|
| 1381 |
+
> assertions) is independent of the verdict computation**, so a harness bug can never silently turn a
|
| 1382 |
+
> real failure into a pass, and never costs a re-run to correct."*
|
| 1383 |
+
> — `LIVE_VALIDATION_POSTFIX.md`
|
| 1384 |
+
|
| 1385 |
+
This is the strongest form of the claim: the property was not merely designed, it was **exercised**. A
|
| 1386 |
+
run with a broken verdict computation was corrected *post hoc* from its own raw record, without touching
|
| 1387 |
+
the deployment or re-running the browser.
|
| 1388 |
+
|
| 1389 |
+
### 9.4 The generalisation
|
| 1390 |
+
|
| 1391 |
+
The pattern is: **record observations, compute verdicts separately, keep the observations.** It is the
|
| 1392 |
+
same shape as the evidence engine's digest (§4.8): the *record* is content-addressed and process-independent,
|
| 1393 |
+
and the *interpretation* is a function over it. Both make the same guarantee — that a conclusion can be
|
| 1394 |
+
re-derived from the evidence rather than trusted as an assertion.
|
| 1395 |
+
|
| 1396 |
+
---
|
| 1397 |
+
|
| 1398 |
+
## 10. How to run the tests
|
| 1399 |
+
|
| 1400 |
+
### 10.1 The interpreter
|
| 1401 |
+
|
| 1402 |
+
Use the repository virtual environment, not a bare `python`. On Windows:
|
| 1403 |
+
|
| 1404 |
+
```bash
|
| 1405 |
+
.venv/Scripts/python.exe -m pytest <target> -q
|
| 1406 |
+
```
|
| 1407 |
+
|
| 1408 |
+
`docs/REPRODUCIBILITY.md` §10.2 records this as the invocation that reproduces the published results. On
|
| 1409 |
+
Windows, pytest must be given `-p no:cacheprovider` because the sandbox refuses `.pytest_cache` writes
|
| 1410 |
+
(`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §0).
|
| 1411 |
+
|
| 1412 |
+
### 10.2 Run targeted files, not the whole tree
|
| 1413 |
+
|
| 1414 |
+
**A full `tests/unit` run trips the sandbox's bulk-delete guard** (the 4× `test_safe_delete_shim`
|
| 1415 |
+
failures of §7.3). The workaround is to run targeted suites:
|
| 1416 |
+
|
| 1417 |
+
```bash
|
| 1418 |
+
.venv/Scripts/python.exe -m pytest tests/unit/test_frontend_live_wiring.py -q # 106 passed
|
| 1419 |
+
.venv/Scripts/python.exe -m pytest tests/unit/test_evidence_engine.py --basetemp=.pytest_tmp -q
|
| 1420 |
+
.venv/Scripts/python.exe -m pytest tests/unit/test_deploy_config.py -p no:cacheprovider -q
|
| 1421 |
+
.venv/Scripts/python.exe -m pytest tests/unit/test_api_contract_doc.py tests/unit/test_frontend_guide_doc.py -p no:cacheprovider -q
|
| 1422 |
+
```
|
| 1423 |
+
|
| 1424 |
+
The runbook notes that **multi-suite invocations in one command have been refused by the environment
|
| 1425 |
+
before; single suites are reliable** (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §2.6). The `--basetemp`
|
| 1426 |
+
flag is **mandatory** for the evidence-engine suite (§4.7): the default pytest temp root triggers a
|
| 1427 |
+
sandbox denial on this machine, and the module's `scratch` fixture writes to a repo-local `.pytest_tmp`
|
| 1428 |
+
instead.
|
| 1429 |
+
|
| 1430 |
+
### 10.3 `pytest.ini`
|
| 1431 |
+
|
| 1432 |
+
The configuration is four lines plus the marker registrations:
|
| 1433 |
+
|
| 1434 |
+
```ini
|
| 1435 |
+
[pytest]
|
| 1436 |
+
testpaths = tests
|
| 1437 |
+
pythonpath = .
|
| 1438 |
+
addopts = -q --tb=short
|
| 1439 |
+
```
|
| 1440 |
+
|
| 1441 |
+
(`pytest.ini:1-4`). `pythonpath = .` is what makes `from core.config import ...` resolve from the repo
|
| 1442 |
+
root without an installed package — and `test_deploy_config.py` deliberately proves the module *also*
|
| 1443 |
+
bootstraps `sys.path` itself, *"WITHOUT pytest's `pythonpath = .` shortcut"* (§6.4). The markers are
|
| 1444 |
+
registered at `pytest.ini:23-27` (§0.1), and `filterwarnings` silences three warning classes including
|
| 1445 |
+
`rasterio.errors.NotGeoreferencedWarning` (`pytest.ini:5-8`).
|
| 1446 |
+
|
| 1447 |
+
### 10.4 The output-buffering trap
|
| 1448 |
+
|
| 1449 |
+
**Do not pipe pytest through `grep`** in the authoring sandbox — output is block-buffered and a killed
|
| 1450 |
+
pipeline swallows it. Redirect to a file instead (`docs/REPRODUCIBILITY.md` §5.6). The same trap applies
|
| 1451 |
+
to the live-validation harness (§8.7).
|
| 1452 |
+
|
| 1453 |
+
### 10.5 The expected results
|
| 1454 |
+
|
| 1455 |
+
| Suite | Command | Expected | Status |
|
| 1456 |
+
|---|---|---|---|
|
| 1457 |
+
| Frontend live-wiring | `pytest tests/unit/test_frontend_live_wiring.py -q` | **106 passed** | `VERIFIED` |
|
| 1458 |
+
| Doc/frontend suite | `pytest` on the 5 doc/frontend files | **183 passed** | `VERIFIED` |
|
| 1459 |
+
| Evidence engine | `pytest tests/unit/test_evidence_engine.py --basetemp=.pytest_tmp` | **73 passed** | `VERIFIED` |
|
| 1460 |
+
| Gateway policy | `pytest tests/unit/test_gateway_policy.py` | **51 passed** | `VERIFIED` |
|
| 1461 |
+
| Full unit suite | `pytest tests/unit` | **5–6 environmental/ordering failures**, rest pass; a re-run passes **137** | `MEASURED` |
|
| 1462 |
+
|
| 1463 |
+
(`release/repo/README.md` §The test suites; `docs/PHASE19_FINAL_HARDENING.md` §6;
|
| 1464 |
+
`docs/REPRODUCIBILITY.md` §5.1.) The precise full-suite **collected** count is
|
| 1465 |
+
`UNKNOWN — not established from the available evidence` (§7.7).
|
| 1466 |
+
|
| 1467 |
+
---
|
| 1468 |
+
|
| 1469 |
+
## 11. What is NOT tested
|
| 1470 |
+
|
| 1471 |
+
A green suite is not evidence about a property nobody wrote an assertion for (`DOCS_STYLE_GUIDE.md` §1.6;
|
| 1472 |
+
§0.3). This section lists the properties the release does **not** have test evidence for, so that no
|
| 1473 |
+
reader mistakes coverage for completeness.
|
| 1474 |
+
|
| 1475 |
+
| Property | State | Why / what exists instead |
|
| 1476 |
+
|---|---|---|
|
| 1477 |
+
| **Continuous integration** | **does not exist** | there is **no `.github/` directory** and no CI workflow in the repository. Every test result in this chapter was produced by a **manual** run in the authoring sandbox. Nothing runs the suites automatically on push. |
|
| 1478 |
+
| **End-to-end system benchmark** | **NOT RUN (none exists)** | `CURRENT_RELEASE_STATE.md` §4: *"System-level end-to-end benchmark — **NOT RUN (none exists)**."* No system-level accuracy is claimed anywhere in the release. |
|
| 1479 |
+
| **Load / soak / concurrency** | **NOT RUN** | no test exercises sustained load, concurrent users, or a long-running process. The per-IP rate limiter is a **fairness** control, not a security or capacity control (`docs/DEPLOYMENT_ARCHITECTURE.md` §5.2). |
|
| 1480 |
+
| **Adversarial / fuzz testing** | **NOT RUN** | no fuzzing or adversarial-input suite. Input validation is tested by example (§2.2.9), not by search. |
|
| 1481 |
+
| **Cross-dataset generalization** | **NOT RUN** | the measured metrics are per-dataset (LEVIR-CD-256, VRSBench, held-out test sets); no test measures transfer to an unseen dataset. |
|
| 1482 |
+
| **Browser / device matrix** | **NOT RUN** | live validation ran in **headed Chromium** only. No Firefox/Safari/WebKit, no mobile, no accessibility audit. |
|
| 1483 |
+
| **Deployment verification in CI** | **NOT RUN** | the runbook guard asserts the document does **not** claim a deployment happened (§3.4). Deployment reachability is established only by the manual live validation (§8). |
|
| 1484 |
+
| **Router test split** | **NOT RUN** | `CURRENT_RELEASE_STATE.md` §4: router accuracy `0.965116` is **validation, ungated, n = 86**; *"the test split was NOT RUN"* (`DOCS_STYLE_GUIDE.md` §3). |
|
| 1485 |
+
| **Calibration as an improvement** | **not established (and measured the other way)** | ECE went `0.013755 → 0.014929` — **worse**. The engine suite pins that calibration is *honestly labelled* (§4.5), not that it helps. |
|
| 1486 |
+
| **Optical-SAR / change-VQA acceptance** | **OPEN** | the rulings are OPEN (`CURRENT_RELEASE_STATE.md` §4; `DOCS_STYLE_GUIDE.md` §3). A suite passing is not a model acceptance. |
|
| 1487 |
+
| **VLM adapter acceptance** | **REJECTED** | metrics `USABLE_VERIFIED` (exact_match 0.963) but status **ACCEPTANCE-REJECTED**. USABLE ≠ ACCEPTED (`DOCS_STYLE_GUIDE.md` §3). |
|
| 1488 |
+
| **Tunnel reliability (B-07)** | **OPEN** | transient tunnel-agent gaps; patch **prepared, NOT deployed**. No test can cover a defect that is not fixed in production. |
|
| 1489 |
+
| **The inference service's internals** | **UNKNOWN** | the tunnel-agent source is not in the monorepo; `docs/architecture/02-deployment-topology.md` §3.3/§9.4 documents the topology as a **measured** fact, but the agent's internals are `UNKNOWN — not established from the available evidence`. |
|
| 1490 |
+
| **Credential write permission (HF)** | **UNKNOWN** | `CURRENT_RELEASE_STATE.md` §7: HF token identity `thundercode`, role `fineGrained`; *"write permission not yet proven."* |
|
| 1491 |
+
|
| 1492 |
+
Two of these deserve emphasis:
|
| 1493 |
+
|
| 1494 |
+
- **No CI is the largest gap.** Every count in this chapter is a manual measurement. A reader should not
|
| 1495 |
+
read "106 passed" as "the suite passes on every commit" — it means "it passed when it was run, in the
|
| 1496 |
+
recorded environment." There is no automation that would catch a regression on push.
|
| 1497 |
+
- **No end-to-end benchmark is the second.** The individual task metrics are artifact-backed and
|
| 1498 |
+
measured, but there is no measurement of the *system* — router + specialist + envelope + frontend — on
|
| 1499 |
+
a held-out end-to-end corpus. The live validation (§8) proves the *pipeline runs and dispatches
|
| 1500 |
+
correctly*; it does not measure system accuracy.
|
| 1501 |
+
|
| 1502 |
+
---
|
| 1503 |
+
|
| 1504 |
+
## 12. Open items and evidence index
|
| 1505 |
+
|
| 1506 |
+
### 12.1 Open / blocked / not-run items for this chapter's topic
|
| 1507 |
+
|
| 1508 |
+
Following the style guide's requirement that every doc end with an explicit list
|
| 1509 |
+
(`DOCS_STYLE_GUIDE.md` §4):
|
| 1510 |
+
|
| 1511 |
+
| Item | State | Note |
|
| 1512 |
+
|---|---|---|
|
| 1513 |
+
| B-07 — transient tunnel-agent gaps | **OPEN** | patch **prepared, NOT deployed**; root shape measured (in `auto` mode a tunnel timeout falls through to the forward path, burning `wake_timeout_s=120` on a 302 ≈ 249 s) |
|
| 1514 |
+
| B-02 — `codespace_name` trailing `\n` | **OPEN (cosmetic)** | confirmed still live; the wake path strips it |
|
| 1515 |
+
| No CI | **does not exist** | no `.github/` directory; all results are manual |
|
| 1516 |
+
| System-level end-to-end benchmark | **NOT RUN (none exists)** | no system-level accuracy claimed |
|
| 1517 |
+
| Router test split | **NOT RUN** | val-only, n = 86 |
|
| 1518 |
+
| Full-suite collected test count | **UNKNOWN** | per-suite counts known; the single collected total is not |
|
| 1519 |
+
| `test_safe_delete_shim` failures | **environmental** | sandbox delete-guard; not regressions; precise Windows trigger **UNKNOWN** |
|
| 1520 |
+
| Ordering flake | **environmental** | passes in isolation |
|
| 1521 |
+
| Stale adapter test | **stale** | asserts CROMA unshipped; CROMA is shipped |
|
| 1522 |
+
| Live validation | **MEASURED** | 3 passes × 8 cases, 8/8 each, 24 runs, 0 mock nodes, trace fill 94.4444 % |
|
| 1523 |
+
| Optical-SAR / change-VQA rulings | **OPEN** | 0.931 / macro-F1 0.434161; 0.697626 / 0.378373 |
|
| 1524 |
+
| VLM adapter acceptance | **REJECTED** | metrics usable; acceptance rejected |
|
| 1525 |
+
| Calibration | **not an improvement** | ECE worse (0.013755 → 0.014929); retained only because it is in the frozen config |
|
| 1526 |
+
|
| 1527 |
+
### 12.2 Where the evidence lives
|
| 1528 |
+
|
| 1529 |
+
| What | Where |
|
| 1530 |
+
|---|---|
|
| 1531 |
+
| Test configuration and markers | `pytest.ini` |
|
| 1532 |
+
| Frontend live-wiring suite | `tests/unit/test_frontend_live_wiring.py` — **106 passed** |
|
| 1533 |
+
| Evidence-engine suite | `tests/unit/test_evidence_engine.py` — **73 passed** |
|
| 1534 |
+
| Safe-delete shim guard | `tests/unit/test_safe_delete_shim.py` — 338 lines; skips cleanly off-sandbox |
|
| 1535 |
+
| Doc-guards | `tests/unit/test_api_contract_doc.py`, `test_frontend_guide_doc.py`, `test_runbook_doc.py`, `test_deploy_config.py`, `test_step7_api_contract_conformance.py` |
|
| 1536 |
+
| Integration suite scope | `tests/integration/test_step7_backend_chain.py`, `test_step7_r02_serving.py`; `docs/ITEM5_INTEGRATION_SUITE_SCOPE.md` |
|
| 1537 |
+
| The suite results, reported in full | `docs/EVALUATION.md` §7.1–7.5; `docs/REPRODUCIBILITY.md` §5; `release/repo/README.md` §The test suites |
|
| 1538 |
+
| The full-suite failure classification | `docs/FINAL_DELIVERY_REPORT.md` §7; `docs/PHASE12_115_METRIC_COMPUTED.md` §8 (the 2,237-passed run) |
|
| 1539 |
+
| Live-validation record | `.workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md` |
|
| 1540 |
+
| Live-validation raw output | `.workbuddy-ai/scratch/live_validation/run_output.txt`, `run_final2.txt`, `run_final3.txt` |
|
| 1541 |
+
| Live-validation screenshots | `.workbuddy-ai/scratch/live_validation/A1_vqa.png` … `B2_newairport.png` (8) |
|
| 1542 |
+
| Live-validation harness (v1) | `.workbuddy-ai/scratch/run_all_postfix.harness` |
|
| 1543 |
+
| Live-validation harness (v2, asserting) | `.workbuddy-ai/scratch/run_all_postfix2.harness` |
|
| 1544 |
+
| CDP focus diagnostic | `.workbuddy-ai/scratch/diag_focus.harness` |
|
| 1545 |
+
| Verdict recomputation | `.workbuddy-ai/scratch/recompute_verdicts.py` |
|
| 1546 |
+
| Prior independent audit | `../2026-09-25-00-37-41/live_validation/INDEPENDENT_LIVE_VALIDATION.md` |
|
| 1547 |
+
| Release-state inventory | `release/CURRENT_RELEASE_STATE.md` |
|
| 1548 |
+
|
| 1549 |
+
### 12.3 Cross-references
|
| 1550 |
+
|
| 1551 |
+
- The trust model, secret custody, CORS allowlist, size limits and rate-limiting-as-fairness that the
|
| 1552 |
+
frontend suite asserts against: [`SECURITY.md`](SECURITY.md).
|
| 1553 |
+
- The API contract the doc-guards validate: [`API_CONTRACT.md`](API_CONTRACT.md).
|
| 1554 |
+
- The deployment topology the live validation drives: [`DEPLOYMENT.md`](DEPLOYMENT.md),
|
| 1555 |
+
[`architecture/02-deployment-topology.md`](architecture/02-deployment-topology.md).
|
| 1556 |
+
- The measured metrics this chapter does **not** re-derive: [`EVALUATION.md`](EVALUATION.md).
|
| 1557 |
+
- The reproduction commands and their expected outputs: [`REPRODUCIBILITY.md`](REPRODUCIBILITY.md) §5.
|
| 1558 |
+
- The router defect the live validation re-validates: [`architecture/04-router.md`](architecture/04-router.md),
|
| 1559 |
+
[`architecture/09-frontend.md`](architecture/09-frontend.md) §10.
|