thundercode commited on
Commit
0bace1a
·
verified ·
1 Parent(s): 80bcc68

release: add README_FULL.md

Browse files
Files changed (1) hide show
  1. README_FULL.md +1437 -0
README_FULL.md ADDED
@@ -0,0 +1,1437 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # SatQuery AI — Satellite Imagery Question Answering
2
+
3
+ **Satellite imagery question answering, remote-sensing change detection, visual grounding and
4
+ optical-SAR fusion in one open, CPU-runnable system.**
5
+
6
+ Ask a natural-language question about a satellite or aerial image — or a pair of images — and SatQuery
7
+ AI routes it to the right specialist model, collects evidence, and returns a single confidence-scored
8
+ result envelope. It is a **remote-sensing vision-language system** built as a *router plus specialists*
9
+ pipeline: satellite image captioning, visual question answering, **text-guided visual grounding**,
10
+ **bi-temporal satellite image change detection**, change-VQA, and **Sentinel-1 / Sentinel-2 optical-SAR
11
+ fusion**. It runs on **CPU**, is served from a static frontend, and is live at
12
+ **https://satquery.pages.dev**.
13
+
14
+ [![Release](https://img.shields.io/badge/release-1.0.0-blue)](RELEASE_MANIFEST.md)
15
+ [![Live demo](https://img.shields.io/badge/demo-satquery.pages.dev-success)](https://satquery.pages.dev)
16
+ [![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97-models-yellow)](https://huggingface.co/thundercode/SatQuery)
17
+ [![Docs](https://img.shields.io/badge/docs-37%20documents-informational)](docs/architecture/README.md)
18
+ [![License](https://img.shields.io/badge/license-not%20selected-red)](#license)
19
+
20
+ <p align="center">
21
+ <img src="screenshots/analyze-grounding.png" alt="The SatQuery AI Analyze console answering a text-guided visual grounding question about satellite imagery against live inference" width="820">
22
+ </p>
23
+
24
+ > **Status: research prototype, pre-1.0.** The architecture is frozen. This repository is the
25
+ > documented public release of the system, its trained artifacts, and its measured results —
26
+ > **including the negative ones.**
27
+
28
+ This README is the front door to a long-form documentation set. It is written to the same standard as
29
+ the rest of the release: **every number, path, run identifier and status is taken from a file that was
30
+ actually read**, and where something was never run, that is stated rather than implied.
31
+
32
+ ---
33
+
34
+ ## Topics
35
+
36
+ This system spans several searchable areas. If you are looking for a **remote-sensing
37
+ vision-language model**, a **satellite image change-detection** baseline, a **text-guided visual
38
+ grounding** implementation for Earth observation, **Sentinel-1/Sentinel-2 optical-SAR fusion**, or a
39
+ **router-and-specialists** design that runs on CPU, this repository covers that ground.
40
+
41
+ | Area | What is here |
42
+ |---|---|
43
+ | **Satellite imagery question answering** | six task specialists behind one typed `ResultEnvelope` |
44
+ | **Remote-sensing VQA & captioning** | SmolVLM-500M over frozen backbones, with a trained LoRA adapter |
45
+ | **Visual grounding (text → box)** | RemoteCLIP ViT-B/32 + a trained grounding head; VRSBench protocol |
46
+ | **Change detection (bi-temporal)** | STANet-style Siamese detector on LEVIR-CD-256 |
47
+ | **Change-VQA** | natural-language questions over a temporal pair |
48
+ | **Optical-SAR fusion** | CROMA-base, 19-class BigEarthNet CLC head |
49
+ | **Multimodal / vision-language** | evidence aggregation, calibrated confidence, execution traces |
50
+ | **Deployment** | static frontend + gateway + outbound-tunnel inference on CPU |
51
+
52
+ Search terms this project is described by: `satellite imagery question answering` ·
53
+ `remote sensing vision language model` · `satellite image change detection` · `visual grounding remote
54
+ sensing` · `optical SAR fusion` · `Sentinel-1 Sentinel-2 fusion` · `satellite image captioning` ·
55
+ `LEVIR-CD` · `VRSBench` · `BigEarthNet` · `router and specialists architecture` · `vision-language model
56
+ on CPU`.
57
+
58
+ ---
59
+
60
+ ## Table of contents
61
+
62
+ - [Topics](#topics)
63
+ - [Motivation](#motivation)
64
+ - [What the system supports](#what-the-system-supports)
65
+ - [Supported inputs](#supported-inputs)
66
+ - [Architecture](#architecture)
67
+ - [Repository map](#repository-map)
68
+ - [The four endpoints](#the-four-endpoints)
69
+ - [Documentation map](#documentation-map)
70
+ - [Routing and the execution trace](#routing-and-the-execution-trace)
71
+ - [The eight execution events](#the-eight-execution-events)
72
+ - [Real inference vs. the preview path](#real-inference-vs-the-preview-path)
73
+ - [Models](#models)
74
+ - [Measured results](#measured-results)
75
+ - [Grounding: three decode variants, two protocols](#grounding-three-decode-variants-two-protocols)
76
+ - [Fusion: measured but the ruling is open](#fusion-measured-but-the-ruling-is-open)
77
+ - [Change-VQA: two test sets, and they disagree](#change-vqa-two-test-sets-and-they-disagree)
78
+ - [Phase 6 / VLM: deployment success ≠ model acceptance](#phase-6--vlm-deployment-success--model-acceptance)
79
+ - [Calibration: it got worse, and we say so](#calibration-it-got-worse-and-we-say-so)
80
+ - [Live validation](#live-validation)
81
+ - [Representative real run IDs](#representative-real-run-ids)
82
+ - [Screenshots](#screenshots)
83
+ - [Installation](#installation)
84
+ - [Local development](#local-development)
85
+ - [Deployment](#deployment)
86
+ - [Deployment caveats](#deployment-caveats)
87
+ - [Reproducibility](#reproducibility)
88
+ - [Known limitations](#known-limitations)
89
+ - [Links](#links)
90
+ - [Citation](#citation)
91
+ - [License](#license)
92
+
93
+ ---
94
+
95
+ ## Motivation
96
+
97
+ Remote-sensing analysis is fragmented. Detecting change between two acquisitions, localising an
98
+ object, captioning a scene, answering a question about it, and fusing optical with SAR each live in a
99
+ different model, a different preprocessing convention, and a different output schema. Assembling them
100
+ into one answer means re-solving the same problems — tiling, band handling, coordinate systems,
101
+ confidence — every time.
102
+
103
+ SatQuery AI explores a single hypothesis: **a small deterministic router plus a shared evidence
104
+ contract can make a heterogeneous specialist ensemble behave like one system**, without a large
105
+ language model in the control path. The router *understands* the query. A deterministic policy
106
+ *decides* which specialists run. The specialists *compute*. The evidence engine *proves* the answer.
107
+
108
+ Two design rules follow from that, and they are non-negotiable in the codebase:
109
+
110
+ - **No LLM-generated coordinates. No LLM-generated confidence.** Coordinates come from detection and
111
+ segmentation heads; confidence comes from a calibrated scoring path.
112
+ - **Every result carries an observable execution trace** — never chain-of-thought.
113
+
114
+ ### Why not one end-to-end model
115
+
116
+ The design is a *router plus specialists*, not a single model that "does satellite QA". Each layer has
117
+ exactly one verb, assigned in the architecture freeze:
118
+
119
+ > router **understands**; policy engine **decides**; specialists **compute**; VLM **explains**;
120
+ > evidence engine **proves**.
121
+
122
+ That division is not decoration — it resolves real ambiguities about *where* a decision belongs. The
123
+ `Intent` type makes the first row explicit in code:
124
+
125
+ ```python
126
+ class Intent(BaseModel):
127
+ """Output of the learned router. Advisory only — the controller decides."""
128
+ ```
129
+
130
+ The four constraints that force the modular design:
131
+
132
+ | Constraint | Consequence |
133
+ |---|---|
134
+ | The system must run on **CPU** | End-to-end VLM inference at usable quality needs a GPU; small per-task modules do not. |
135
+ | Different tasks have **incompatible outputs** | `change` returns a spatial change map; `caption` returns prose; `optical_sar` returns a class distribution. One head cannot emit all three. |
136
+ | Tasks have **different data and metrics** | Each specialist is trained and evaluated on its own split with its own protocol. |
137
+ | **Truthfulness** | Per-task metrics are auditable. A single end-to-end number would hide which component failed. |
138
+
139
+ The cost of this design is that there is **no system-level accuracy number** — because there is no
140
+ single model to measure. That absence is stated rather than papered over; see
141
+ [Known limitations](#known-limitations) item 1.
142
+
143
+ ### What the system is deliberately not
144
+
145
+ | Absent | Why |
146
+ |---|---|
147
+ | Database, authentication, users, job queue | the gateway is stateless by design; inference is synchronous |
148
+ | GPU requirement | device is selected via `SATQUERY_DEVICE`; all placement is `.to(device)`, never `.cuda()` |
149
+ | Gradio GUI | the frontend is a separate static tier; `app/space_app.py` serves JSON only |
150
+ | Chain-of-thought in traces | traces carry observable facts only — states, timings, counts, config hash, model refs |
151
+ | Backbone redistribution | backbones are fetched from the Hugging Face Hub, pinned by revision |
152
+ | System-level end-to-end benchmark | **NOT RUN — none exists** |
153
+ | Router test-split evaluation | **NOT RUN** |
154
+
155
+ ## What the system supports
156
+
157
+ Six specialist tasks. All six are reported `available: true` by the live capability contract
158
+ (`GET /api/capabilities`, probed 2026-09-25, `schema_version 1.0`).
159
+
160
+ | Task | What it answers | Assets | `requires_pair` | `max_assets` |
161
+ |---|---|---|---|---|
162
+ | `vqa` | A free-form question about a single scene | 1 | false | 1 |
163
+ | `caption` | A description of a single scene | 1 | false | 1 |
164
+ | `grounding` | *Where* is a described object or region — returns boxes | 1 | false | 1 |
165
+ | `change` | *What changed* between two co-registered acquisitions — returns change regions | 2 | true | 2 |
166
+ | `change_vqa` | A yes/no or short question about a detected change | 2 | true | 2 |
167
+ | `optical_sar` | Joint scene classification from an optical + SAR pair | 2 | true | 2 |
168
+
169
+ The live contract also carries per-task notes — `vqa`/`caption` fetch SmolVLM weights from the Hub on
170
+ first use, `grounding` fetches the RemoteCLIP encoder on first use, and `optical_sar` declares
171
+ `modalities: ["optical", "sar"]`.
172
+
173
+ ### The ontology is wider than the capability list
174
+
175
+ There are **two** related vocabularies, and they are not the same six:
176
+
177
+ - `core/schemas.py::Task` carries **seven** values: `vqa`, `caption`, `grounding`, `change`,
178
+ `optical_sar`, `change_vqa`, `unsupported`.
179
+ - The router's label space (`router/label_space.py::TASK_CLASSES`) is **six** classes:
180
+ `vqa, caption, grounding, change, optical_sar, unsupported`.
181
+ - `GET /api/capabilities` lists **six** tasks — the same six as the router's *minus* `unsupported`,
182
+ *plus* `change_vqa`.
183
+
184
+ This asymmetry is intentional. `unsupported` is a **routing outcome** ("this is not a satellite-imagery
185
+ question"), not a servable capability. `change_vqa` is reached through the change family rather than
186
+ being a separate router class, and it is servable. The three sets are reconciled in one place — the
187
+ frontend's `ROUTER_TASK_TO_SERVER` map (`frontend/assets/js/mission.js`) — because `AnalysisRequest` is
188
+ `extra="forbid"` and any string outside the `Task` enum is a 422.
189
+
190
+ ## Supported inputs
191
+
192
+ Confirmed by the implementation, not assumed:
193
+
194
+ | Modality | Task(s) | Format |
195
+ |---|---|---|
196
+ | Optical, single image | `vqa`, `caption`, `grounding` | JPEG, PNG, TIFF |
197
+ | Temporal optical pair | `change`, `change_vqa` | Two images of **identical dimensions** |
198
+ | Optical + SAR pair | `optical_sar` | GeoTIFF/TIFF preferred |
199
+
200
+ **Modality is inferred server-side from band count**, not from the file extension: `{1, 2}` bands ⇒
201
+ SAR, `{3, 4, 8, 11, 12, 13}` bands ⇒ optical. The browser cannot read band count, so the console warns
202
+ when a submitted pair looks like two ordinary photographs rather than an optical/SAR pair.
203
+
204
+ **Per-file upload limit: 4,194,304 bytes (4 MiB).** Larger files are refused with HTTP 413 — imagery
205
+ must be downscaled first.
206
+
207
+ ### The size cap is one number shared by two layers
208
+
209
+ The cap is not a literal in two places; both the gateway and the inference service read
210
+ `SATQUERY_MAX_FILE_BYTES`, and the default is `4 * 1024 * 1024` in both. The inference service's
211
+ `_asset_max_file_bytes()` (`app/space_app.py`) is deliberately strict about it:
212
+
213
+ - a **missing** variable ⇒ the 4 MiB default;
214
+ - a **non-integer** value ⇒ `ValueError` (not silently defaulted);
215
+ - a **non-positive** value ⇒ `ValueError` (a cap of `0` refuses every upload, which is a configuration
216
+ error rather than a limit).
217
+
218
+ The reason is recorded in the source: a silent default would let a deployment whose operator typed a
219
+ malformed cap keep accepting uploads against a limit nobody chose, while the gateway refused to boot
220
+ for the *same* value — the two layers disagreeing about what "too large" means, which is exactly the
221
+ failure the shared variable exists to prevent.
222
+
223
+ The accepted content types mirror the gateway's allowlist (defence in depth — the Space validates
224
+ independently rather than trusting the gateway to be its only caller):
225
+
226
+ ```
227
+ image/tiff, image/geotiff, image/png, image/jpeg, application/octet-stream
228
+ ```
229
+
230
+ ### Uploads are handle-based, and the Space owns the bytes
231
+
232
+ `POST /v1/assets` accepts one file and returns an **opaque ephemeral handle**. The store lives on the
233
+ inference host, not the gateway, because the inference host is where `inspect_raster` reads the bytes
234
+ and where `cache_max_models: 1` serialises their consumption — a gateway-side store would put the bytes
235
+ on a different machine from the reader. The upload endpoint is **off by default** and enabled only when
236
+ `SATQUERY_ASSET_ENABLED` and `SATQUERY_ASSET_DIR` are both set; otherwise it answers a named
237
+ `model_unavailable` envelope explaining the switch. Handle capacity defaults to **32** and the handle
238
+ TTL to **900.0 s**, both overridable by environment (and read from the environment rather than
239
+ `configs/base.yaml` on purpose — adding a key there would move the frozen config hash).
240
+
241
+ ### Input geometry and normalisation
242
+
243
+ Downstream of the format check, input handling is governed by the frozen registry
244
+ (`configs/base.yaml`):
245
+
246
+ | Key | Value | Meaning |
247
+ |---|---|---|
248
+ | `image.max_pixels` | 25,000,000 | hard ceiling on decoded pixels |
249
+ | `image.tile_size` | 512 | tile edge |
250
+ | `image.tile_overlap` | 128 | tile stride overlap |
251
+ | `image.max_tiles` | 64 | hard ceiling on tiles *examined* |
252
+ | `image.top_k_tiles` | 4 | tiles actually sent through a specialist |
253
+ | `optical.normalization` | percentile | 2nd–98th percentile stretch |
254
+ | `optical.lower_percentile` / `upper_percentile` | 2 / 98 | stretch bounds |
255
+ | `optical.canonical_channels` | 12 | CROMA expects exactly 12 optical channels |
256
+ | `sar.representation` | db | SAR is converted to decibels |
257
+ | `sar.clip_min_db` / `clip_max_db` | −30 / 5 | dB clip window |
258
+ | `sar.canonical_channels` | 2 | CROMA expects exactly 2 SAR channels (VV, VH) |
259
+
260
+ The tile policy follows the plan's section 9.1: whole-image thumbnail first, then top-K tiles. `max_tiles`
261
+ is the hard ceiling on tiles examined; `top_k_tiles` is how many are actually dispatched — and the loader
262
+ **rejects** a config where `top_k_tiles > max_tiles`.
263
+
264
+ ## Architecture
265
+
266
+ This is the **actually deployed** topology. An older direct-client-to-inference design is superseded.
267
+
268
+ ```mermaid
269
+ flowchart TD
270
+ B["Browser<br/>(static console)"] -->|HTTPS| CF["Cloudflare Pages<br/>satquery.pages.dev"]
271
+ CF -->|"HTTPS JSON · /api/*"| R["Render<br/>satquery-orchestrator"]
272
+ R -->|"outbound long-poll<br/>POST /tunnel/agent"| T{{"outbound tunnel"}}
273
+ T --> C["GitHub Codespace<br/>FastAPI inference · CPU · :8000"]
274
+ C --> S["Specialists"]
275
+ S --> M["SmolVLM · RemoteCLIP · STANet-change<br/>CROMA-fusion · MiniLM router"]
276
+ M --> E["Evidence engine<br/>+ temperature scaling"]
277
+ E --> RE["ResultEnvelope"]
278
+ RE -->|"tunnel → Render"| B
279
+ ```
280
+
281
+ ### Why a tunnel
282
+
283
+ The inference host runs in a GitHub Codespace. The forwarded-port path is not reachable for a private
284
+ repo (it returns HTTP 302), so the orchestrator keeps a **long-poll tunnel**: the Codespace dials out
285
+ to `POST /tunnel/agent` and holds the connection; Render queues work onto it. `transport_mode` is
286
+ `auto`, and the tunnel is the live transport. There is **no** `SATQUERY_UPSTREAM_URL` and **no**
287
+ `HF_TOKEN` in the live configuration — the transport is the outbound tunnel, not a forwarded port.
288
+
289
+ The live health payload (`GET /api/health`, probed 2026-09-25) records the tunnel's state directly:
290
+
291
+ ```json
292
+ {"status":"ok","service":"satquery-orchestrator",
293
+ "tunnel":{"agent_connected":true,"agent_id":"codespaces-fd1038","pending":0,"completed":97},
294
+ "config":{"codespace_name":"potential-space-trout-r4ppw969w45j2pvvw\n","codespace_port":8000,
295
+ "transport_mode":"auto","tunnel_timeout_s":150.0,"wake_timeout_s":120.0,
296
+ "upstream_timeout_s":90.0,"device":"cpu","has_github_token":true}}
297
+ ```
298
+
299
+ Note `codespace_name` still carries a trailing `\n` — that is item B-02, cosmetic, and
300
+ [still open](#deployment-caveats).
301
+
302
+ ### The nine-state controller
303
+
304
+ The inference host is a FastAPI service built by `build_space_app()`. Behind the transport sits a
305
+ deterministic controller with a **nine-state** finite state machine (`core/schemas.py::ControllerState`,
306
+ mirrored in `configs/base.yaml` §`agent.states`):
307
+
308
+ ```
309
+ RECEIVE → PARSE → VALIDATE → PLAN → PREPROCESS → EXECUTE → AGGREGATE → VERIFY → RESPOND
310
+ ```
311
+
312
+ The controller is the **only** component that dispatches. `agent.max_specialists` is 4,
313
+ `agent.timeout_seconds` is 120, and `agent.unload_after_workflow` is true — the controller unloads
314
+ models after a workflow so that `cache_max_models: 1` is honoured rather than thrashing the cache.
315
+
316
+ ### Frozen backbones, trained modules
317
+
318
+ Four backbones are pinned by `repo_id` + `revision` and fetched from the Hub on first use. Nothing is
319
+ fine-tuned end-to-end.
320
+
321
+ | Role | Repository | Revision | Size / notes |
322
+ |---|---|---|---|
323
+ | Router encoder | `sentence-transformers/all-MiniLM-L6-v2` | `1110a243fdf4` | 90.9 MB, 22,713,216 params, 384-dim embeddings |
324
+ | VLM | `HuggingFaceTB/SmolVLM-500M-Instruct` | `a7da5b986cb5` | ~1015 MB safetensors |
325
+ | Grounding | `chendelong/RemoteCLIP` (`RemoteCLIP-ViT-B-32.pt`) | `bf1d8a3ccf2d` | 605.2 MB; width 768, **projected** dim 512 |
326
+ | Optical-SAR | `antofuller/CROMA` (`CROMA_base.pt`) | `0dd28e3d633b` | 777.6 MB; `encoder_dim` 768, `image_resolution` 120 |
327
+
328
+ Two consequences follow: the system is small (the six trained artifacts total ~125 MiB; everything else
329
+ is public weights), and **no backbone weights are redistributed** by this release.
330
+
331
+ ### The evidence and confidence stages
332
+
333
+ Specialists each emit `Evidence` for what *they* computed. The `EvidenceEngine`
334
+ (`evidence/engine.py`) does not re-derive any of it; it aggregates:
335
+
336
+ ```
337
+ collect across specialists → deduplicate → order → renumber → cap
338
+ ```
339
+
340
+ The pipeline is **pure and deterministic** — no clock, no RNG, no I/O — and its purity is a test
341
+ assertion, not a hope. Three details matter:
342
+
343
+ - **Stable identity.** `Evidence.evidence_id` defaults to a random uuid, which is useless across runs.
344
+ The engine renumbers to `evidence_001`, `evidence_002`, … (zero-padded to three digits, well past the
345
+ `evidence.max_items` bound of 32) so a downstream artefact can cite one item deterministically.
346
+ - **Defined order.** Items are sorted by `(type, source_specialist, score DESC, coordinates)`. The
347
+ coordinate tie-breaker is what makes the order *total*; without it, two items sharing type,
348
+ specialist and score would fall back to Python's stable-sort insertion order, reintroducing
349
+ input-order dependence.
350
+ - **Content identity, not container identity.** Dedup keys on
351
+ `(type, source_specialist, coordinate_system, rounded coordinates, rounded score)` — `payload`,
352
+ `artifact_ref` and `evidence_id` are deliberately excluded. Two items that agree on the same
353
+ geolocation, one carrying a `crs` and one not, have made the same claim about the world; the surviving
354
+ item's payload is merged with the discarded one's so the `crs` is not lost. Two items that share a type
355
+ and score but **disagree** on coordinates are two different claims and are both kept — suppressing a
356
+ spatial disagreement would be the silent contradiction the freeze forbids.
357
+
358
+ The engine records `dropped_duplicates` (non-zero is normal and healthy — it means two specialists
359
+ agreed), `dropped_over_limit` (non-zero is a warning — a specialist's evidence did not survive), and
360
+ `truncated`. `evidence_digest()` exists so that reproducibility is an assertion:
361
+
362
+ ```
363
+ aggregate(inputs_a) is reproducible iff digest(a) == digest(b)
364
+ ```
365
+
366
+ The confidence stage is `evidence/confidence.py`, and it exists to enforce one rule: **a calibration that
367
+ claims to be fitted when it is not is a false claim of reliability.** So when no fitted artifact is
368
+ available it passes the raw score through unchanged, sets `method="uncalibrated"`, and leaves
369
+ `calibrated=None` — which is what makes it honest, because `ConfidenceBreakdown.value` then returns
370
+ `raw`, and any consumer can distinguish "we calibrated this" from "we did not". The temperature scaling
371
+ itself is `calibrated = sigmoid(logit(z) / T)`, with two stated failure modes handled explicitly: a
372
+ fitted `T` of exactly `1.0` is the identity map and is reported as uncalibrated rather than silently
373
+ pretending to have done something, and `z = 0.0` (whose log-odds diverge) is clamped at the boundary so
374
+ a hard zero cannot become a NaN.
375
+
376
+ ### Repository map
377
+
378
+ | Path | Contents |
379
+ |---|---|
380
+ | `app/` | FastAPI inference service and its composition root |
381
+ | `core/` | Config registry, evidence engine, contracts |
382
+ | `specialists/` | One module per specialist (vqa, caption, grounding, change, optical_sar) |
383
+ | `router/` | MiniLM intent router |
384
+ | `gateway/` | Render orchestration hub (`/api/*`, CORS, wake flow) |
385
+ | `frontend/` | The static console (HTML/CSS/JS) |
386
+ | `configs/base.yaml` | The frozen configuration registry — single source of truth |
387
+ | `evaluation/`, `training/` | Evaluation harnesses and training entry points |
388
+ | `artifacts/` | Trained heads, checkpoints, evaluation outputs, provenance |
389
+ | `docs/` | Architecture, models, benchmarks, deployment, limitations |
390
+
391
+ The component inventory in full (every path is the authoritative location):
392
+
393
+ | Layer | Module | Responsibility |
394
+ |---|---|---|
395
+ | **Contracts** | `core/schemas.py` | the binding typed contract: `Task`, `Intent`, `Evidence`, `SpecialistResult`, `ResultEnvelope`, `ExecutionTrace`, … |
396
+ | **Config** | `core/config.py` | load, deep-merge, validate, hash the registry; `get_config()` singleton |
397
+ | **Errors** | `core/errors.py` | the error taxonomy (`ConfigError`, `ModelLoadError`, `RoutingError`, `WorkflowPlanError`, …) |
398
+ | **Planning** | `core/planner.py` | turn an `Intent` into a concrete, ordered `ExecutionPlan` |
399
+ | **Registry** | `core/registry.py` | specialist registration / lookup |
400
+ | **Controller** | `core/controller.py` | the nine-state FSM; the only thing that dispatches |
401
+ | **Router** | `router/encoder.py` | frozen MiniLM embedding, cached |
402
+ | | `router/adapter.py` | the five-head `IntentAdapter` (the only trainable router part) |
403
+ | | `router/classifier.py` | learned classification + confidence gate + fallback selection |
404
+ | | `router/fallback.py` | deterministic lexical fallback (`lexical_route`) |
405
+ | | `router/label_space.py` | the ontology, single source of truth |
406
+ | | `router/dataset.py`, `router/train.py` | dataset generation and training |
407
+ | **Specialists** | `specialists/base.py` | the specialist interface |
408
+ | | `specialists/vqa/{model,inference,prompts}.py` | SmolVLM VQA |
409
+ | | `specialists/grounding/{remoteclip,head,inference,specialist}.py` | RemoteCLIP + head |
410
+ | | `specialists/change/{stanet,specialist,postprocess,vqa_specialist}.py` | change detection + change-VQA |
411
+ | | `specialists/optical_sar/{croma,fusion_head,inference,specialist,sensor_adapter,radiometry,prompts}.py` | CROMA fusion |
412
+ | **Evidence** | `evidence/engine.py` | aggregation: dedup → sort → renumber → cap |
413
+ | | `evidence/confidence.py` | temperature scaling, honest pass-through |
414
+ | **Inference app** | `app/space_app.py` | `build_space_app()`; the four-endpoint JSON contract |
415
+ | | `app/serving.py` | the composition root (`build_serving_controller()`) |
416
+ | | `app/deployment.py` | capability / health payload builders |
417
+ | **Gateway** | `gateway/` | the Render orchestrator, asset store, policy/error translation |
418
+ | **Frontend** | `frontend/` | the static site |
419
+
420
+ ### The four endpoints
421
+
422
+ The inference service serves exactly **four** JSON endpoints (`app/space_app.py`). No Gradio GUI exists
423
+ in code; the file serves JSON only, by design, so it cannot compete with the static frontend's contract.
424
+
425
+ | Endpoint | Method | Purpose |
426
+ |---|---|---|
427
+ | `/v1/health` | GET | liveness + capability states; **loads no model** |
428
+ | `/v1/capabilities` | GET | the six servable tasks with `requires_pair` / `max_assets` |
429
+ | `/v1/assets` | POST | accept one uploaded file, return an opaque ephemeral handle |
430
+ | `/v1/analyze` | POST | run one analysis; returns a `ResultEnvelope` |
431
+
432
+ The gateway in front of it exposes the same functionality under `/api/*` and holds the request-side
433
+ security boundary. Two details are worth recording because they cost real debugging time:
434
+
435
+ - **Framework-raised errors carry the same envelope.** An unmatched route (404) and a method mismatch
436
+ (405) are wrapped by a Starlette exception handler so they answer with the contract's error shape
437
+ (`routing_error`, `recoverable: false`) rather than FastAPI's default `{"detail": …}`. An unhandled
438
+ exception answers with `satquery_error` and a fixed, operator-safe message; the exception's own text is
439
+ logged **server-side only**, so a traceback cannot disclose internal paths to an unauthenticated caller.
440
+ - **The upload body is read bounded.** `POST /v1/assets` uses `read_body_bounded(request, cap)` rather
441
+ than `await request.body()`, so the cap is applied *while reading* rather than after the whole body has
442
+ been buffered. The measured defect this fixed: with the cap at 1 MiB, a 64 MiB body produced a peak
443
+ allocation of 128 MiB, tracking body size linearly with no ceiling, and the `413` came only after
444
+ everything had been held.
445
+
446
+ ### Documentation map
447
+
448
+ The architecture reference is a hub plus ten deep sub-documents. Every one of them is written at
449
+ long-form depth, with real signatures, schemas, numbers and file paths.
450
+
451
+ | # | Document | What it covers |
452
+ |---|---|---|
453
+ | — | [`docs/architecture/README.md`](docs/architecture/README.md) | the architecture hub: thesis, sub-document index, cross-cutting principles, what is deliberately absent |
454
+ | 01 | [`docs/architecture/01-system-overview.md`](docs/architecture/01-system-overview.md) | the thesis, the component inventory, the frozen-backbone strategy, what is deliberately absent |
455
+ | 02 | [`docs/architecture/02-deployment-topology.md`](docs/architecture/02-deployment-topology.md) | the four tiers, the gateway, the outbound tunnel, wake flow, cold start, `transport_mode` |
456
+ | 03 | [`docs/architecture/03-request-lifecycle.md`](docs/architecture/03-request-lifecycle.md) | the nine-state controller, validation rules, modality inference, tiling |
457
+ | 04 | [`docs/architecture/04-router.md`](docs/architecture/04-router.md) | frozen MiniLM, the five-head adapter, `interpret()` vs `chooseTask()`, the lexical fallback, the label space |
458
+ | 05 | [`docs/architecture/05-specialists.md`](docs/architecture/05-specialists.md) | all six tasks: entry points, preprocessing, postprocessing, outputs |
459
+ | 06 | [`docs/architecture/06-evidence-and-confidence.md`](docs/architecture/06-evidence-and-confidence.md) | the evidence schema, the aggregation pipeline, temperature scaling, the eight execution events |
460
+ | 07 | [`docs/architecture/07-configuration-freeze.md`](docs/architecture/07-configuration-freeze.md) | the registry, the enforced invariants, the config hash, why it is frozen |
461
+ | 08 | [`docs/architecture/08-api-contract.md`](docs/architecture/08-api-contract.md) | the four endpoints, the envelopes, error codes, transport headers |
462
+ | 09 | [`docs/architecture/09-frontend.md`](docs/architecture/09-frontend.md) | the static pages, the Analyze console, real-vs-preview, platform traps |
463
+ | 10 | [`docs/architecture/10-observability-and-ops.md`](docs/architecture/10-observability-and-ops.md) | health, counters, traces, what is and is not observed |
464
+
465
+ Companion references, all in this repository:
466
+
467
+ | Document | Contents |
468
+ |---|---|
469
+ | [`docs/MODELS.md`](docs/MODELS.md) | the six artifacts in detail, backbone pinning, rejected model decisions |
470
+ | [`MODEL_CARD.md`](MODEL_CARD.md) | the Hugging Face model card (intended use, out-of-scope use, measured performance) |
471
+ | [`models/manifest.json`](models/manifest.json) | machine-generated byte counts and sha256, one entry per artifact |
472
+ | [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md) | the headline metric table and the rules it follows |
473
+ | [`docs/EVALUATION.md`](docs/EVALUATION.md) | how each number was produced; evaluation-honesty rules; behavioural validation |
474
+ | [`docs/DATASETS.md`](docs/DATASETS.md) | LEVIR-CD-256, VRSBench, BigEarthNet, CDVQA/SECOND — measured corpus figures |
475
+ | [`docs/TRAINING.md`](docs/TRAINING.md) | per-artifact hyperparameters and where each was trained |
476
+ | [`docs/DEPLOYMENT.md`](docs/DEPLOYMENT.md) | live revisions, env vars, deploy mechanics, platform traps |
477
+ | [`docs/REPRODUCIBILITY.md`](docs/REPRODUCIBILITY.md) | what a third party can and cannot reproduce |
478
+ | [`docs/RESEARCH_NOTES.md`](docs/RESEARCH_NOTES.md) | findings that changed the code; the router defect; negative results |
479
+ | [`docs/LIMITATIONS.md`](docs/LIMITATIONS.md) | the honest catalogue of everything not done or done poorly |
480
+ | [`docs/CHANGELOG.md`](docs/CHANGELOG.md) | versioned record of what changed and what was verified |
481
+
482
+ ## Routing and the execution trace
483
+
484
+ Routing is deliberately two-stage, and the split matters:
485
+
486
+ 1. **`interpret()` — the reading.** A lexical pass over the query produces a *reading*: task intent,
487
+ modality, temporal requirement, spatial scope, and expected evidence kind. It is
488
+ **asset-count-blind**.
489
+ 2. **`chooseTask()` — the dispatch.** The reading is combined with the number of attached assets to
490
+ decide the task actually dispatched. This is why a reading of `change` with **one** asset dispatches
491
+ `change_vqa` — the documented quantifier upgrade.
492
+
493
+ The server-side router has the same two-part shape in Python: `router/classifier.py::IntentRouter.route()`
494
+ produces an `Intent`, and `core/planner.py` turns that `Intent` into an ordered `ExecutionPlan`. The
495
+ planner's docstring states the rule it exists to enforce: **the planner is the only component permitted
496
+ to choose what runs; its input is a router prediction, its output is a plan, and the router's `Intent` is
497
+ an input to a decision, never the decision.** A design where `intent.task` selects a specialist in one
498
+ step has collapsed *understand* into *decide*, and three things break: the router can never be overruled,
499
+ a query needing two specialists can never get both, and there is no auditable record of the decision.
500
+
501
+ ### The five router heads
502
+
503
+ The learned router is a small adapter over frozen MiniLM embeddings. `router/label_space.py` is the single
504
+ source of truth for its output ontology — both the dataset generator and the training script import from
505
+ it, so a class added in one place cannot silently desynchronise the other.
506
+
507
+ | Head | Classes / kind | Loss |
508
+ |---|---|---|
509
+ | `task` | 6 classes: `vqa, caption, grounding, change, optical_sar, unsupported` | softmax cross-entropy |
510
+ | `modality` | 4 classes: `optical, sar, optical_sar, unknown` | softmax cross-entropy |
511
+ | `temporal` | binary | single logit + `BCEWithLogitsLoss` |
512
+ | `spatial_output` | binary | single logit + `BCEWithLogitsLoss` |
513
+ | `language_output` | binary | single logit + `BCEWithLogitsLoss` |
514
+
515
+ The three binary heads use a single logit rather than a two-way softmax: a 2-way softmax would waste a
516
+ parameter and make the loss harder to weight. Because the heads are independent by construction, a
517
+ confident task label with an incoherent binary head is possible; the classifier resolves that in favour
518
+ of the task label (the controller keys off the task), and records that it did so.
519
+
520
+ Router configuration (`configs/base.yaml` §`router`):
521
+
522
+ | Key | Value |
523
+ |---|---|
524
+ | `model` | `sentence-transformers/all-MiniLM-L6-v2` |
525
+ | `revision` | `1110a243fdf4` |
526
+ | `max_length` | 128 |
527
+ | `embedding_dim` | 384 |
528
+ | `hidden_dim` | 128 |
529
+ | `dropout` | 0.10 |
530
+ | `confidence_threshold` | 0.70 |
531
+ | `num_tasks` | 6 |
532
+ | `training.epochs` / `batch_size` / `learning_rate` | 60 / 64 / 0.001 |
533
+ | `training.weight_decay` | 0.01 |
534
+ | `training.task_loss_weight` / `modality_loss_weight` / `binary_loss_weight` | 1.0 / 0.3 / 0.5 |
535
+ | `training.val_ratio` | 0.15 |
536
+ | `training.hard_negatives_to_test` | true |
537
+
538
+ **Finding F4-1 — the tokenizer ceiling.** The MiniLM tokenizer's own ceiling is **256** (verified by
539
+ probe). The project truncates to **128** — a deliberate truncation *well inside* the ceiling, not the
540
+ model limit. Satellite queries are short; halving the sequence halves attention cost for no measurable
541
+ accuracy loss. The encoder **asserts** `max_length ≤ 256`, because truncating above the ceiling is a
542
+ silent no-op.
543
+
544
+ **Finding F4-2 — the router needs no GPU.** The encoder is frozen, so embeddings are cached and the
545
+ 50,822-parameter adapter trains on cached vectors. **Measured on CPU: 20 epochs / 4,096 vectors in
546
+ 0.28 s.**
547
+
548
+ **Finding F4-3 — splits must be by group.** Splits are by **group** (template / hard-negative family),
549
+ never by example. Hard-negative families are placed in the **test** split so their accuracy measures
550
+ generalisation rather than memorisation; splitting by example would leak template variants across the
551
+ boundary.
552
+
553
+ ### The confidence gate and the fallback
554
+
555
+ `IntentRouter.route()` runs the learned router first, and falls back to the lexical rules only when the
556
+ learned router is below the confidence gate (or when the encoder cannot be loaded at all):
557
+
558
+ ```
559
+ query
560
+ → encoder.encode (frozen MiniLM, 384-d)
561
+ → adapter (5 heads, softmax / sigmoid)
562
+ → confidence gate (router.confidence_threshold = 0.70)
563
+ → lexical fallback (only if below the gate)
564
+ → Intent (validated pydantic model)
565
+ ```
566
+
567
+ The router **never refuses to answer**; `above_threshold` carries the uncertainty, and the planner — not
568
+ the router — decides what to do about it. The fallback is purely lexical: no model, no embeddings, no
569
+ randomness, ordered rules with the highest specificity first, and it **never invents capability** — if
570
+ nothing matches it returns `unsupported` with low confidence rather than guessing a task. Its precedence
571
+ is explicit, and it matters because the phrasings overlap:
572
+
573
+ ```
574
+ dual_modality > temporal > spatial > caption > vqa > unsupported
575
+ ```
576
+
577
+ Two examples of why precedence is load-bearing, both from the source:
578
+
579
+ - *"show me where the change happened"* has both a spatial term and a temporal term → `change` +
580
+ `spatial_output=True`.
581
+ - *"compare optical and radar to locate built-up areas"* has dual-modality **and** spatial → `optical_sar`
582
+ (spatial does not apply to the joint workflow, whose output is a classification).
583
+
584
+ The fallback's confidence band (0.72–0.92) **overlaps and can exceed** the trained model's — on the spec's
585
+ own examples the fallback returns 0.850–0.920 against the trained model's 0.780–1.000. So the planner
586
+ applies a **provenance discount** to its own reading of the confidence and never edits
587
+ `Intent.confidence` itself: a lexical fallback at 0.9 is not the same evidence as a learned model at 0.9,
588
+ and treating them identically would let a matched keyword outrank the model it fell back from.
589
+ `IntentRouter.adapter_source` returns `'trained'` or `'lexical_fallback'` so a caller can check this in
590
+ one place instead of inferring it from a confidence band — because `from_config` defaults `adapter_path`
591
+ to `None`, meaning the default router runs the **fallback**, not the trained adapter.
592
+
593
+ ### A router bug worth recording
594
+
595
+ An earlier revision evaluated the temporal rule before the location rule, so *"Where are the built-up
596
+ areas in this image?"* matched `\bbuilt\b` as a *change* marker and `area` inside *"areas"* as a
597
+ quantifier. With one asset it collapsed to `vqa` and answered **"River"**. Fixed on 2026-09-25 in
598
+ `frontend/assets/js/mission.js`; the fix is covered by regression tests and verified live. The same defect
599
+ existed on a second surface (`SQ.policy` in `core.js`) and was fixed the same day.
600
+
601
+ The fix is four lexical changes, each documented in the source because each was a real failure:
602
+
603
+ | Change | Why |
604
+ |---|---|
605
+ | `built` removed from the temporal term set entirely | *"built-up areas"* is land-cover vocabulary, not a change marker. While it sat in the temporal set, the location question *"Where are the built-up areas in this image?"* was read as a change request and answered with the degenerate one-word *"River"*. Measured live, 2026-09-25. |
606
+ | `\barea\b` instead of bare `area` | Without the boundary the substring matched inside *"areas"*, so the already-mis-read change question was upgraded **again** to `change_vqa`. The boundary keeps the quantifier reading for a real *"how much area changed"* while refusing the plural land-cover noun. |
607
+ | `new` counts as a change marker **only** when the query is not a `where` question | The repo ships `eo/new-airport.jpg`, so *"Where is the new airport?"* is a real question, and `new` is a place descriptor as often as a change marker. |
608
+ | the change stem is matched **without** a trailing `\b` | `\bchang\b` cannot match *"changed"*, *"changes"* or *"changing"* — there is no word boundary between the stem and its inflection. With the boundary, the page's own default question (*"What changed here?"*) fell through to the `vqa` branch, so the change path was unreachable from the UI that exists to reach it. |
609
+
610
+ Both defect queries now dispatch to `grounding` and are captured in the screenshot set below.
611
+
612
+ ### The eight execution events
613
+
614
+ The console renders an execution trace built from **eight events**, emitted by the frontend around
615
+ real network calls (`SQ.EVENT_NAMES` in `frontend/assets/js/core.js`):
616
+
617
+ | # | Event | Emitted when |
618
+ |---|---|---|
619
+ | 1 | `QUERY_RECEIVED` | The query and assets are accepted |
620
+ | 2 | `QUERY_UNDERSTOOD` | `interpret()` has produced the reading |
621
+ | 3 | `ROUTE_SELECTED` | `chooseTask()` has selected the dispatched task |
622
+ | 4 | `SPECIALIST_STARTED` | The inference request has been issued |
623
+ | 5 | `SPECIALIST_COMPLETED` | The specialist has returned |
624
+ | 6 | `EVIDENCE_GENERATED` | Evidence items are available |
625
+ | 7 | `CONFIDENCE_COMPUTED` | The calibrated confidence is available |
626
+ | 8 | `RESULT_ASSEMBLED` | The `ResultEnvelope` is complete |
627
+
628
+ These are a **frontend** vocabulary driven by observable events — not a backend protocol and not a
629
+ model's reasoning trace. The run engine is deliberately dumb: it renders whatever events it receives, and
630
+ swapping the mock driver for a websocket/SSE feed of the same event names is the entire integration
631
+ surface. On live runs the trace bar reaches **94.4444 %** (17/18) and every node is marked live; the
632
+ preview path is the only source of mock-marked nodes.
633
+
634
+ The trace is deliberately *not* chain-of-thought. `ExecutionTrace` (`core/schemas.py`) carries
635
+ `run_id`, `schema_version`, `task`, `query`, `inputs`, `modalities`, `intent`, `validation`, `workflow`,
636
+ `steps`, `selected_models`, `parameters`, `outputs`, `confidence`, `timings`, `fallbacks`, `errors`,
637
+ `contradiction`, `config_hash`, `started_at`, `finished_at` — observable facts, with no field for model
638
+ reasoning and no LLM-generated confidence. The router's own trace projection is the model of this: it
639
+ emits the task, modality, the three booleans, the rounded confidence, the source, `above_threshold`,
640
+ `used_fallback` and `fallback_rule` — and nothing else.
641
+
642
+ ## Real inference vs. the preview path
643
+
644
+ - **Real path (production).** With assets attached, the console calls
645
+ `POST /api/infer` on the Render orchestrator. Every run returns a real `run_*` identifier from the
646
+ inference service. **Live validation recorded 0 mock nodes across 24 live runs.**
647
+ - **Preview path.** With *no* files selected, the console renders a labelled illustrative preview so
648
+ the interface is explorable without the stack awake. Preview nodes are explicitly marked `is-mock`
649
+ and never appear in a live run.
650
+
651
+ The distinction is observable, not asserted: a live run shows `live · N evidence · transport …` and
652
+ zero `.trace__node.is-mock` elements.
653
+
654
+ Two related traps are worth stating because they are the kind of thing a reader will otherwise
655
+ misdiagnose:
656
+
657
+ - **A pair-requiring task with one asset is refused, not silently downgraded server-side.** With one
658
+ asset, `change` answers `invalid_request` (*"change requires exactly 2 assets (T1 and T2); got 1"*) and
659
+ the whole envelope comes back `degraded: true`. Measured live, 2026-09-25. The console's job is to
660
+ avoid asking for a pair-requiring task when only one asset exists; it does so in `chooseTask()`, and it
661
+ **names the substitution** rather than hiding it.
662
+ - **The asset set sent is per-task.** `vqa`, `grounding` and `caption` accept a single image, while
663
+ `change`, `change_vqa` and `optical_sar` require a pair. Sending the optional earlier frame to a
664
+ single-image task makes the backend reject the whole request — measured live when a pair was uploaded
665
+ and a VQA question asked. The fix isolates the file set so the pair is only ever sent to the tasks that
666
+ declared it.
667
+
668
+ ## Models
669
+
670
+ Six trained artifacts are released. **Four are task heads and two are adapters** — none is a complete
671
+ standalone model, and each documents its backbone dependency. Full detail: [`docs/MODELS.md`](docs/MODELS.md),
672
+ [`MODEL_CARD.md`](MODEL_CARD.md), and the generated [`models/manifest.json`](models/manifest.json).
673
+
674
+ | Task | Backbone (pinned) | Custom component | Artifact | Size | Eval data | Metric | Status |
675
+ |---|---|---|---|---|---|---|---|
676
+ | `change` | STANet-style, ResNet-18 encoder, PAM | change head | `head.pt` | 63,231,009 B | LEVIR-CD-256, test n=2048 | pooled IoU **0.8122** · macro IoU **0.8457** · pooled F1 **0.8964** | **VERIFIED** |
677
+ | `grounding` | `chendelong/RemoteCLIP` ViT-B/32 @ `bf1d8a3ccf2d` (frozen) | trainable head (2048→512) | `head.pt` | 12,639,041 B | VRSBench, n=16159 | mean_best_IoU **0.2838** · recall@0.5 **0.2198** (canonical) | measured — two protocols |
678
+ | `optical_sar` | `antofuller/CROMA` base @ `0dd28e3d633b` | fusion head (2318→512→19) | `head.pt` | 14,427,457 B | BigEarthNet, 19 CLC classes, test n=4000 | accuracy **0.931** · macro_F1 **0.434161** | measured — **ruling OPEN** |
679
+ | `change_vqa` | as `change` | change-VQA head | `head.pt` | 5,822,809 B | test n=39686 | accuracy **0.697626** · macro_F1 **0.378373** | measured — **ruling OPEN** |
680
+ | `router` | `sentence-transformers/all-MiniLM-L6-v2` @ `1110a243fdf4` | intent adapter | `adapter.pt` | 211,961 B | val n=86 | accuracy **0.965116** | **TEST NOT RUN** |
681
+ | `vqa` / `caption` | `HuggingFaceTB/SmolVLM-500M-Instruct` @ `a7da5b986cb5` | **LoRA** (r=16, α=32, dropout 0.05) | `adapter_model.safetensors` | 34,798,048 B | frozen 1000-Q subset | exact_match **0.963** · F1 **0.96432** | **ACCEPTANCE-REJECTED** |
682
+
683
+ Backbones are third-party and pinned by `repo_id` + `revision` in `configs/base.yaml`; they are
684
+ fetched from the Hugging Face Hub, not redistributed here.
685
+
686
+ ### The six artifacts, byte-for-byte
687
+
688
+ `models/manifest.json` is **generated by reading the files** — no byte count or hash is typed by hand. Its
689
+ schema is `satquery_model_manifest_v1`, generated 2026-09-25T18:15:38+00:00, `artifact_count: 6`, and
690
+ every entry carries the frozen `config_hash` `78f1e3700da15aa1`.
691
+
692
+ | # | `id` | Task | Kind | Local path | HF path | Bytes | sha256 (first 16) |
693
+ |---|---|---|---|---|---|---|---|
694
+ | 1 | `change_head` | `change` | trained head | `artifacts/change/levir_change_v001/head.pt` | `change/head.pt` | 63,231,009 | `c5ef31277b67aa01` |
695
+ | 2 | `change_vqa_head` | `change_vqa` | trained head | `artifacts/change_vqa/run/head.pt` | `change_vqa/head.pt` | 5,822,809 | `cfae5e43b97ca930` |
696
+ | 3 | `optical_sar_fusion_head` | `optical_sar` | trained head | `artifacts/optical_sar/fusion_head_production_v001/head.pt` | `optical_sar/head.pt` | 14,427,457 | `785815729a3a39fc` |
697
+ | 4 | `grounding_head` | `grounding` | trained head | `artifacts/grounding/remoteclip_grounding_v001/head.pt` | `grounding/head.pt` | 12,639,041 | `93432f7034be91a8` |
698
+ | 5 | `router_adapter` | `router` | trained adapter | `artifacts/router/router_adapter_v001/adapter.pt` | `router/adapter.pt` | 211,961 | `8527c3ed28a293e1` |
699
+ | 6 | `vlm_lora_adapter` | `vlm` | LoRA adapter | `.scratch/phase6_real_adapter/phase6_adapter/adapter_model.safetensors` | `vlm/adapter_model.safetensors` | 34,798,048 | `07c76a75fa046248` |
700
+
701
+ Total released weight payload: **131,130,325 bytes (~125 MiB)**. Two of the six hashes are cross-checked
702
+ against values recorded **independently** elsewhere in the project — `change_vqa_head` against
703
+ `artifacts/change_vqa/run/PROMOTION.json`, and `vlm_lora_adapter` against the adapter's own provenance
704
+ manifest — and both agree. That is an external cross-check, not a self-consistency claim.
705
+
706
+ The manifest also records each artifact's architecture and the metric artifact it came from:
707
+
708
+ | `id` | Architecture | Source metric artifact |
709
+ |---|---|---|
710
+ | `change_head` | STANet-style Siamese change detector (ResNet-18 + PAM) | `artifacts/change/eval_test/eval_result.json` |
711
+ | `change_vqa_head` | `change_vqa_head_v1` (1,453,912 parameters) | `artifacts/change_vqa/run/PROMOTION.json` |
712
+ | `optical_sar_fusion_head` | CROMA-base fusion head (input 2318 → hidden 512 �� 19 classes) | `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json` |
713
+ | `grounding_head` | RemoteCLIP ViT-B/32 grounding head (feature 2048, hidden 512) | `artifacts/grounding/remoteclip_grounding_v001/eval_result_canonical.json` |
714
+ | `router_adapter` | task/modality adapter over frozen MiniLM embeddings (~50,822 params) | `artifacts/router/threshold_sweep_val.json` |
715
+ | `vlm_lora_adapter` | PEFT LoRA (r=16, α=32, dropout 0.05) on text-model projections | `artifacts/vlm/phase6_closure.json` |
716
+
717
+ Training checkpoints also exist (`checkpoint_last.pt` at 189,291,829 B for change; `checkpoint_last.pt` at
718
+ 12,640,331 B for grounding; `checkpoint-1500` / `checkpoint-2000` for the LoRA adapter) and are
719
+ **not** the released artifacts — they are archived as provenance.
720
+
721
+ ### The Hugging Face release
722
+
723
+ The six artifacts are published at **https://huggingface.co/thundercode/SatQuery** (public,
724
+ `private: false`, `gated: false`), HEAD `55681e0cddb91a4a5655da98a49bc025e537b657`, 22 files, last
725
+ modified `2026-09-25T18:21:52Z`.
726
+
727
+ The release was verified by **re-downloading each artifact over direct HTTPS and hashing the bytes
728
+ received**, rather than trusting the upload step: 6/6 `MATCH`, 0 failed, and the four support files
729
+ (`README.md`, `MODEL_CARD.md`, `models/manifest.json`, `models/checksums.sha256`) confirmed present. The
730
+ pre-existing content was a 25-byte stub README (literally `---\nlicense: unknown\n---`) which was
731
+ replaced, and the standard HF LFS routing `.gitattributes`, which was left untouched.
732
+
733
+ > **A verification method that was itself wrong (recorded).** The *first* verification attempt reported
734
+ > all six artifacts `DIFFER`, with every remote hash equal to
735
+ > `e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855` — the sha256 of **empty content**.
736
+ > The cause was not the upload: `hf_hub_download` returned an empty file in this environment, so the
737
+ > verifier hashed nothing. It was caught by a second, independent method (a direct `curl` download), which
738
+ > produced the correct hash for `router/adapter.pt` and confirmed it to be a real PyTorch zip (`PK\x03\x04`).
739
+ > The verifier was then rewritten to use direct HTTPS with proxies disabled. The failed first attempt is
740
+ > recorded because a verifier that silently hashes an empty file would have produced a **false failure** —
741
+ > and, with a different bug, could just as easily have produced a **false pass**.
742
+
743
+ **No secret was uploaded.** The uploaded set is the model card, the manifest, the checksums, the docs, and
744
+ the six weight files; no tokens, keys, environment files or credentials exist in any uploaded file, and
745
+ the token used for the upload is not written into any released file.
746
+
747
+ ## Measured results
748
+
749
+ Every number below traces to an artifact, a test, or a live run. **Nothing here is a system-level
750
+ benchmark — no such benchmark exists** (see [Known limitations](#known-limitations)).
751
+
752
+ | Metric | Value | Split / protocol | Source key | Status |
753
+ |---|---|---|---|---|
754
+ | Change pooled IoU | 0.8122 | LEVIR-CD-256 test, n=2048, thr 0.50 | `metrics.pooled.iou` | **VERIFIED** |
755
+ | Change macro IoU | 0.8457 | same | `metrics.macro.miou` | **VERIFIED** |
756
+ | Change pooled F1 | 0.8964 | same | `metrics.pooled.f1` | **VERIFIED** |
757
+ | Grounding mean_best_IoU (canonical, `head_threshold`) | 0.2838 | VRSBench, n=16159 | `results.head_threshold.mean_best_iou` | measured |
758
+ | Grounding recall@0.5 (canonical, `head_threshold`) | 0.2198 | same | `results.head_threshold.recall.0.50` | measured |
759
+ | Grounding mean_best_IoU (matched6, `head_threshold`) | 0.2566 | VRSBench, n=16159 | `results.head_threshold.mean_best_iou` | measured |
760
+ | Grounding recall@0.5 (matched6, `head_threshold`) | 0.1938 | same | `results.head_threshold.recall.0.50` | measured |
761
+ | Grounding `head_argmax` decode (both protocols) | 0.1215 | same | `results.head_argmax.mean_best_iou` | measured — **worse** |
762
+ | Grounding zero-shot baseline | 0.0972 | same | `results.zero_shot_matched.mean_best_iou` | measured |
763
+ | Optical-SAR accuracy | 0.931 | BigEarthNet, held-out test n=4000 | `accuracy` | measured — ruling **OPEN** |
764
+ | Optical-SAR macro_F1 | 0.434161 | same | `macro_f1` | measured — ruling **OPEN** |
765
+ | Change-VQA accuracy (**test**) | 0.697626 | test n=39686 | `verification.test_accuracy` | measured — ruling **OPEN** |
766
+ | Change-VQA macro_F1 (**test**) | 0.378373 | same | `verification.test_macro_f1` | measured — ruling **OPEN** |
767
+ | Change-VQA accuracy (**test2**) | 0.651469 | second test set | `verification.test2_accuracy` | measured — **lower** |
768
+ | Change-VQA macro_F1 (**test2**) | 0.372309 | second test set | `verification.test2_macro_f1` | measured — **lower** |
769
+ | VLM adapter exact_match | 0.963 | frozen 1000-Q subset | `artifacts/vlm/phase6_closure.json` | USABLE_VERIFIED — **ACCEPTANCE-REJECTED** |
770
+ | VLM adapter F1 | 0.96432 | same | same | USABLE_VERIFIED — **ACCEPTANCE-REJECTED** |
771
+ | Router **overall ungated** accuracy | 0.965116 | val, n=86, corpus-limited | `overall_ungated_accuracy` | **TEST NOT RUN** |
772
+ | System-level end-to-end benchmark | — | — | — | **NOT RUN — none exists** |
773
+
774
+ ### Source artifacts and the rules the table follows
775
+
776
+ Each metric family has exactly one source artifact:
777
+
778
+ | Metric family | Artifact |
779
+ |---|---|
780
+ | change | `artifacts/change/eval_test/eval_result.json` |
781
+ | grounding | `artifacts/grounding/remoteclip_grounding_v001/eval_result_canonical.json`, `…_matched6.json` |
782
+ | optical-SAR | `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json` |
783
+ | change-VQA | `artifacts/change_vqa/run/PROMOTION.json` |
784
+ | router | `artifacts/router/threshold_sweep_val.json` |
785
+ | calibration | `artifacts/calibration_v001.json` |
786
+ | VLM | `artifacts/vlm/phase6_closure.json` |
787
+
788
+ All **20** quoted metrics are checked against these files by
789
+ [`release/tools/verify_readme_metrics.py`](https://github.com/Anish-lab-blip/SatQuery-AI), which resolves
790
+ nested artifact keys (including keys that themselves contain dots — the `recall` dict is keyed
791
+ `"0.10"/"0.25"/"0.50"`, so a naive `split(".")` walk would break) and compares each value at the precision
792
+ printed here. It exits non-zero if any claim fails and prints `ALL CLAIMS VERIFIED` only when everything
793
+ matches. The table obeys six rules:
794
+
795
+ 1. **Two protocols are never collapsed.** Grounding is reported under *both* the canonical and matched6
796
+ protocols. Quoting 0.2838 alone would be selective.
797
+ 2. **Two test sets are never collapsed.** Change-VQA is reported on `test` **and** `test2`.
798
+ 3. **accuracy never travels without macro-F1.** For imbalanced multi-class heads (optical-SAR,
799
+ change-VQA) the macro-F1 is reported alongside accuracy, always.
800
+ 4. **Validation is not test.** The router number is labelled "overall **ungated** accuracy", val, n = 86.
801
+ 5. **A negative result stays negative.** Calibration ECE worsened and is shown worsening.
802
+ 6. **USABLE ≠ ACCEPTED.** The VLM metrics are real; the artifact is nevertheless acceptance-rejected.
803
+
804
+ There is also **no composite or vanity score**: no single headline accuracy for the system, and none
805
+ invented by averaging the per-task numbers.
806
+
807
+ ### Grounding: three decode variants, two protocols
808
+
809
+ The grounding head is evaluated under **two matching protocols** (canonical, matched6) and **three**
810
+ decode variants. Quoting a single number would misrepresent the result, so all of them are listed:
811
+
812
+ | Decode | canonical mean_best_IoU | matched6 mean_best_IoU |
813
+ |---|---|---|
814
+ | `head_threshold` (the headline number) | **0.2838** | **0.2566** |
815
+ | `head_argmax` | 0.1215 | 0.1215 |
816
+ | `zero_shot_matched` (baseline, no head) | 0.0972 | 0.0972 |
817
+
818
+ The head clears the zero-shot baseline, but only the threshold decode is meaningfully above it — the
819
+ argmax decode (0.1215) is barely better than zero-shot. The absolute level is modest either way:
820
+ **grounding is useful, not solved.**
821
+
822
+ The recall@0.5 numbers travel with the IoU numbers: canonical **0.2198**, matched6 **0.1938**. The box
823
+ convention is a common source of silent error, which is why the project converts VRSBench's 0–100 boxes
824
+ to its own 0–1 convention through a *declared* `benchmark_box_scale: 100.0` — so the conversion cannot be
825
+ applied twice or forgotten — and reports both protocols.
826
+
827
+ The head itself is a trainable head over the frozen RemoteCLIP ViT-B/32 encoder. Its per-cell feature is
828
+ `concat([patch, text, patch·text, global_pool])` = `4 × 512 = 2048` (finding P7-1: the transformer width
829
+ is 768, but `visual.proj` maps to a projected dim of **512**). Cells are assigned by ground-truth box
830
+ centre (`cell_relative` decode). The objectness BCE is weighted **20×** because only ~1 of 49 cells is
831
+ positive; unweighted, the optimum collapses to "no object" everywhere. Image resolution is frozen at
832
+ **224** — 448 was evaluated and **rejected** (see below).
833
+
834
+ ### The grounding resolution decision — a pre-registered rejection
835
+
836
+ **Question:** should grounding decode at 448 or 224? **Answer: 224. 448 was rejected** — and the rejection
837
+ is notable because it was *pre-registered* and then *confirmed* by a paired test over identical samples
838
+ (n = 16,159):
839
+
840
+ | Comparison (448 vs 224) | Value |
841
+ |---|---|
842
+ | mean best IoU | **−0.0147** |
843
+ | recall@0.5 | −0.0022 |
844
+ | recall@0.10 | −0.0699 |
845
+ | recall@0.25 | −0.0243 |
846
+ | latency | **1.59×** |
847
+ | paired 95 % CI | [−0.0160, −0.0134] |
848
+ | paired t | **−22.63** |
849
+ | 448 better on | 8.5 % of records |
850
+ | 448 worse on | **20.9 %** of records |
851
+
852
+ 448 lost on **every** axis. The pre-registered decision rule and the paired test **agree** on 224. This is
853
+ a model of how a resolution decision should be made: declared in advance, then tested. Recorded as
854
+ `RESOLVED 2026-09-16` in `configs/base.yaml` and in
855
+ [`docs/RESEARCH_NOTES.md`](docs/RESEARCH_NOTES.md) §2.
856
+
857
+ ### Fusion: measured but the ruling is open
858
+
859
+ Optical-SAR fusion reaches **0.931 accuracy** on a 19-class held-out set of 4,000 — but only
860
+ **0.434 macro_F1**. Those two numbers describe very different things: the model is accurate on
861
+ frequent classes and weak on rare ones. The acceptance ruling for this head is **OPEN**, and the
862
+ headline accuracy must never be quoted without the macro_F1 beside it.
863
+
864
+ The metric JSON records why the macro score is low by construction: of the 19 classes in the label space,
865
+ **14 are present** and **5 are absent** in the scored split, and the **macro-F1 denominator is all 19**
866
+ (absent classes contribute 0.0). It also records `is_deciding_statistic: False` — this is a reported
867
+ measurement, not a decision statistic. The feature concatenation is the frozen one:
868
+
869
+ ```
870
+ input_dim = 3 × 768 + 12 + 2 = 2318 → hidden 512 → num_classes 19 (BigEarthNet CLC)
871
+ ```
872
+
873
+ and the **availability mask is consumed by the fusion head, not by CROMA** (finding C-1) — CROMA always
874
+ sees the canonical channel counts (12 optical, 2 SAR).
875
+
876
+ Two further caveats on this head, stated rather than hidden:
877
+
878
+ - **The live service returns a bare class index** (`class_18`), not a human-readable label. The modality
879
+ accounting in the response confirms the right channels reached the fusion head, but the presentation is
880
+ not user-facing.
881
+ - **The local BigEarthNet subset is 100 % single-label**, against the official 1–11 multi-label scheme, so
882
+ its metrics are **not comparable** to published BigEarthNet numbers. Any statement of the form
883
+ "BigEarthNet mAP = X" is false for this subset.
884
+
885
+ ### Change-VQA: two test sets, and they disagree
886
+
887
+ `artifacts/change_vqa/run/PROMOTION.json` records **two** test evaluations:
888
+
889
+ | Split | accuracy | macro_F1 |
890
+ |---|---|---|
891
+ | `test` | 0.697626 | 0.378373 |
892
+ | `test2` | **0.651469** | 0.372309 |
893
+
894
+ The `test` numbers are the higher pair. Both are reported here; quoting only `test` would overstate
895
+ the result. The acceptance ruling is **OPEN**.
896
+
897
+ The head was trained **outside this repository**, on an external GPU (Kaggle), and promoted through a
898
+ byte-identity gate: sha256 `cfae5e43b97ca930f568dc5b8ae4f36b24e9ff717af226159802206ffd63a82a`, 5,822,809
899
+ bytes, architecture `change_vqa_head_v1`, 1,453,912 parameters, **0 non-finite tensors**, weights not
900
+ modified during promotion, byte-identical to source, hash agreeing across `model_metadata.json`,
901
+ `run_record.json` and `hashes.json`. Selection was epoch **8**, chosen on **val answer accuracy
902
+ 0.700018**, stopped by early stopping; seed 42. It trains on **cached change + text features**, not on
903
+ raw imagery — the raw CDVQA loader loads examples but has no training loop of its own, and the two paths
904
+ are not conflated.
905
+
906
+ ### Phase 6 / VLM: deployment success ≠ model acceptance
907
+
908
+ The Phase-6 SmolVLM LoRA adapter reaches **exact_match 0.963** and **F1 0.96432** on a frozen
909
+ 1,000-question subset. It is marked **USABLE_VERIFIED** and **ACCEPTANCE-REJECTED**.
910
+
911
+ Those two verdicts are not in conflict, and the distinction is the point:
912
+
913
+ - **USABLE_VERIFIED** — the adapter loads, runs, and produces the measured numbers in the deployed
914
+ pipeline. The F1 is **+49.5 pp** over the unadapted baseline.
915
+ - **ACCEPTANCE-REJECTED** — the change did not clear the project's own pre-registered acceptance bar. The
916
+ record of *why* is stored in `artifacts/vlm/phase6_closure.json` under `why_acceptance_rejected`; the
917
+ artifact's `status` is `CLOSED`.
918
+
919
+ A model can be a working engineering artifact and a rejected research result at the same time.
920
+ This release keeps both labels. The deployed caption/VQA path therefore uses the **unadapted** SmolVLM.
921
+
922
+ The adapter's own shape is recorded: PEFT **0.19.1**, `r = 16`, `alpha = 32`, `dropout = 0.05`, targeting
923
+ `model.text_model.*.{q,k,v,o,gate,up,down}_proj`, precision **fp16** (finding C-6: the T4 is compute
924
+ capability 7.5, so fp16 — **not** bf16), batch size 2, gradient accumulation 8, learning rate 0.0002,
925
+ 1 epoch, gradient checkpointing on, `save_every_steps` 500.
926
+
927
+ ### Calibration: it got worse, and we say so
928
+
929
+ The `change_vqa` confidence path applies temperature scaling (`T = 0.9773`). Measured on the
930
+ validation split (n=16441):
931
+
932
+ | | ECE | NLL |
933
+ |---|---|---|
934
+ | Before temperature scaling | **0.013755** | 0.689741 |
935
+ | After temperature scaling | **0.014929** | 0.689631 |
936
+
937
+ **Calibration did not improve — it moved slightly worse.** The fitted temperature is `0.9772732` and
938
+ `ece_improvement` is **−0.001174**: negative. The scaling is retained because it is part
939
+ of the frozen configuration, not because it helped. The reliability curve plotted on the Benchmark page
940
+ is explicitly labelled as the **pre-scaling** diagram so a reader cannot mistake it for the calibrated
941
+ result. The calibrated curve is **not plotted**.
942
+
943
+ ### What is NOT benchmarked
944
+
945
+ | Benchmark | Status | Note |
946
+ |---|---|---|
947
+ | **System-level end-to-end accuracy** | **NOT RUN — none exists** | There is no measured end-to-end benchmark of the full router → specialist → envelope pipeline. No such number is claimed anywhere. |
948
+ | **Router test split** | **NOT RUN** | Only the validation split (n = 86) was scored. |
949
+ | **Benchmark adapters** | **NOT RUN** | Adapter-based benchmark runs were not executed. |
950
+ | **Efficiency / latency benchmark** | not systematically measured | Per-specialist latency is recorded incidentally in artifacts (e.g. grounding `latency_ms_per_image` 2.205 ms for the head), but there is no end-to-end latency benchmark. |
951
+ | **Cross-dataset generalisation** | **NOT RUN** | Each specialist is evaluated only on its own training-family test split. |
952
+ | **Human evaluation** | **NOT RUN** | — |
953
+ | **Adversarial / robustness evaluation** | **NOT RUN** | — |
954
+
955
+ ## Live validation
956
+
957
+ Validation drove the **production site** in a headed browser, one upload per case, with per-case
958
+ screenshots and recorded run identifiers. It is **behavioural** evidence — that the pipeline runs and
959
+ routes correctly — and it is **not** an accuracy claim; accuracy and behaviour are evaluated separately.
960
+
961
+ | Property | Result |
962
+ |---|---|
963
+ | Independent full passes | **3** |
964
+ | Cases per pass | 8 (6 regression + 2 defect) |
965
+ | Passes at 8/8 | **3 of 3** |
966
+ | Live runs executed | **24** |
967
+ | Correct dispatches | **24** |
968
+ | Mock-node contamination | **0** on every live run |
969
+ | Trace fill | 94.4444 % on every live run |
970
+ | Frontend regression suite | **106 passed** (`tests/unit/test_frontend_live_wiring.py`) |
971
+
972
+ Each pass produced **fresh run identifiers** — no run id is shared between passes. The three passes ran
973
+ against two frontend revisions:
974
+
975
+ | Pass | Deployed HEAD | Result |
976
+ |---|---|---|
977
+ | 1 | `ff46eba42b18` + `d413d3672311` | 8/8 |
978
+ | 2 | `2d7ae53b482d` | 8/8 |
979
+ | 3 | `2d7ae53b482d` | 8/8 |
980
+
981
+ The harness asserts the form state **before** dispatch — that the query box really holds the intended
982
+ query, that `#obsTail` reads `ready`, and that both frames are attached for pair tasks. This matters:
983
+ an earlier harness revision typed with synthetic key events that Chrome silently drops when the window
984
+ lacks OS focus, so it dispatched the page's *default* query and still recorded a "result". The
985
+ assertions exist because of that failure. The earlier 8/8 run was independently re-examined and confirmed
986
+ **not** to have been infected (its answers were query-specific and the query text was embedded in the
987
+ answers), but the failure mode is recorded because it is exactly the kind of silent false-positive an
988
+ evaluation harness must never have.
989
+
990
+ Deployed-artifact integrity was checked separately: **9 files** were re-read from the GitHub API and
991
+ compared byte-for-byte against local copies, and all 9 were **sha256 byte-identical**; the deployed HEAD
992
+ was re-read from the API.
993
+
994
+ ### Representative real run IDs
995
+
996
+ One full pass (pass 3 of 3) — the same pass the screenshots below are drawn from. Run identifiers
997
+ are fresh on every pass; the other two passes recorded different ids.
998
+
999
+ | Case | Query | Dispatched | Run ID |
1000
+ |---|---|---|---|
1001
+ | vqa | What type of terrain dominates this scene? | `vqa` | `run_fef26e91e7e6` |
1002
+ | caption | Describe the main visual characteristics of this scene. | `caption` | `run_96281bdfcc08` |
1003
+ | grounding | Where are the visible buildings in this image? | `grounding` | `run_e49adc8d319f` |
1004
+ | change | What changed between the earlier and later image? | `change` | `run_aedc59cbcdc9` |
1005
+ | change_vqa | Did the coastline advance between the two observations? | `change_vqa` | `run_62ca98d510be` |
1006
+ | optical_sar | …combining the optical and SAR observations? | `optical_sar` | `run_beacf6aa4e21` |
1007
+ | **grounding** | **Where are the built-up areas in this image?** | **`grounding`** | **`run_467ffa406f22`** |
1008
+ | **grounding** | **Where is the new airport?** | **`grounding`** | **`run_46980ba55c62`** |
1009
+
1010
+ The last two are the router-defect queries. Both previously collapsed to `vqa` and answered "River".
1011
+
1012
+ Note the fifth row: *"Did the coastline advance between the two observations?"* is read as `change` and
1013
+ **dispatches `change_vqa`** — the documented quantifier upgrade, because the page's Answer block promises
1014
+ an answer and the server's `change` returns a spatial map with no language output. With one asset
1015
+ attached, `change` would instead be refused outright (*"requires exactly 2 assets"*); the console avoids
1016
+ asking for a pair-requiring task when only one asset exists.
1017
+
1018
+ ### Screenshots
1019
+
1020
+ Eight captures from the post-fix live run (headed browser, 1384×855, one upload per case). Each
1021
+ panel shows the run identifier, the frozen config hash `78f1e3700da15aa1`, and the evidence list
1022
+ returned by the specialist — nothing is mocked.
1023
+
1024
+ | | |
1025
+ |---|---|
1026
+ | ![Grounding — built-up areas](screenshots/analyze-grounding.png) | ![Optical-SAR](screenshots/analyze-optical-sar.png) |
1027
+ | **Grounding** — "Where are the built-up areas in this image?" — the fixed router defect (`run_467ffa406f22`, dispatched `grounding`, not `vqa`) | **Optical-SAR** fusion on a real optical/SAR GeoTIFF pair (`run_beacf6aa4e21`, fused class 18) |
1028
+ | ![Grounding — new airport](screenshots/analyze-grounding-new-airport.png) | ![Grounding — visible buildings](screenshots/analyze-grounding-buildings.png) |
1029
+ | **Grounding** — "Where is the new airport?" — second defect query (`run_46980ba55c62`, dispatched `grounding`) | **Grounding** — "Where are the visible buildings in this image?" (`run_e49adc8d319f`) |
1030
+ | ![Change](screenshots/analyze-change.png) | ![Caption](screenshots/analyze-caption.png) |
1031
+ | **Change** detection on a same-shape temporal pair | **Caption** of a single scene (`run_96281bdfcc08`, calibrated confidence 1.000) |
1032
+ | ![VQA](screenshots/analyze-vqa.png) | ![Change-VQA](screenshots/analyze-change-vqa.png) |
1033
+ | **VQA** — "What type of terrain dominates this scene?" | **Change-VQA** — "Did the coastline advance between the two observations?" |
1034
+
1035
+ All eight are reproduced byte-for-byte in the evidence archive (Phase 7) with SHA-256 recorded in
1036
+ `RELEASE_MANIFEST.md`.
1037
+
1038
+ ## Installation
1039
+
1040
+ Python 3.11+ and a CPU are sufficient. No CUDA requirement — device is selected via
1041
+ `SATQUERY_DEVICE`; all placement is `.to(device)`, never `.cuda()`.
1042
+
1043
+ ```bash
1044
+ git clone https://github.com/Anish-lab-blip/SatQuery-AI
1045
+ cd SatQuery-AI
1046
+ python -m venv .venv
1047
+ source .venv/Scripts/activate # Windows git-bash; use .venv/bin/activate on Linux/macOS
1048
+ pip install -r requirements.txt
1049
+ ```
1050
+
1051
+ Backbones are fetched from the Hugging Face Hub on first use, pinned by revision in
1052
+ `configs/base.yaml`. The frozen config hash is **`78f1e3700da15aa1`** — the loader refuses to run a
1053
+ config that violates the recorded invariants (for example `fusion.input_dim == 3*encoder_dim + 12 + 2`).
1054
+
1055
+ ### The invariants the loader enforces
1056
+
1057
+ `core/config.py` validates the registry at load time and raises `ConfigError` — naming every violation —
1058
+ rather than letting a bad value reach runtime. The invariants are not documentation; they are checks:
1059
+
1060
+ | Invariant | Why it exists |
1061
+ |---|---|
1062
+ | `croma.image_resolution % 8 == 0` | CROMA asserts this (finding C-7); native 120 → 225 patches |
1063
+ | `training.precision ∈ {fp16, bf16, fp32}` | the T4 is SM 7.5, so bf16 is unavailable (finding C-6) |
1064
+ | `deployment.torch_compile is not true` | ZeroGPU does not support `torch.compile` (finding C-8) |
1065
+ | `vlm.processor_longest_edge ≤ image.tile_size` | the processor's default `longest_edge` is 2048, which upscales a 512 px tile 4× and then splits it into **17** sub-images — a ~17× overrun, not the 4× the plan estimated (finding F5-2). Tying the pin to `image.tile_size` makes it a *control*, so the processor cannot silently start upscaling again. |
1066
+ | `vlm.prompt_must_use_chat_template is true` | SmolVLM raises `ValueError` on prompts lacking one `<image>` token per image (finding F5-3) |
1067
+ | `fusion.input_dim == 3*encoder_dim + optical_channels + sar_channels` (= 2318) | CROMA emits optical/SAR/joint GAP vectors; the availability mask is consumed by the head (finding C-1) |
1068
+ | `croma.optical_channels == 12` and `croma.sar_channels == 2` | CROMA's `s2_channels` / `s1_channels` are fixed |
1069
+ | `grounding_head.feature_dim == 4 * grounding.encoder_projected_dim` (= 2048) | a mismatch is a **silent** shape error — torch raises only at the similarity step, after patch features are already cached (finding P7-1) |
1070
+ | `router.tasks` includes `unsupported` and `router.num_tasks == len(router.tasks)` | the ontology and its declared size cannot drift apart |
1071
+ | `change.sa_mode ∈ {BAM, PAM}` and `change.encoder` is set | the change architecture is not implicit |
1072
+ | `image.top_k_tiles ≤ image.max_tiles` | the dispatch ceiling cannot exceed the examination ceiling |
1073
+
1074
+ Because a config edit moves `Config.hash` and invalidates every artifact keyed to it, deployment state
1075
+ that must not move the hash (asset-store capacity, TTL, the per-file cap) is read from the **environment**
1076
+ rather than from `configs/base.yaml` — the same reasoning that keeps the config hash frozen.
1077
+
1078
+ ## Local development
1079
+
1080
+ ```bash
1081
+ # Inference service, CPU (this is the launcher the Codespace runs)
1082
+ PORT=8000 python deploy/codespace/serve.py
1083
+
1084
+ # Health
1085
+ curl localhost:8000/v1/health
1086
+ ```
1087
+
1088
+ The frontend is fully static and needs no build step to serve locally:
1089
+
1090
+ ```bash
1091
+ python -m http.server 5500 --directory frontend
1092
+ ```
1093
+
1094
+ Run the frontend regression suite:
1095
+
1096
+ ```bash
1097
+ python -m pytest tests/unit/test_frontend_live_wiring.py -q
1098
+ ```
1099
+
1100
+ ### The test suites
1101
+
1102
+ | Suite | Command | Expected |
1103
+ |---|---|---|
1104
+ | Frontend live-wiring | `pytest tests/unit/test_frontend_live_wiring.py` | **106 passed** |
1105
+ | Doc/frontend suite | `pytest` on the 5 doc/frontend files | **183 passed** |
1106
+ | Full unit suite | `pytest tests/unit` | 5–6 **environmental** failures (sandbox delete guard × 4, 1 ordering flake, 1 stale adapter test) |
1107
+
1108
+ The full-suite failures are **not hidden**, and they are not regressions: 4 are the sandbox's bulk-delete
1109
+ guard (`test_safe_delete_shim`), 1 is an ordering flake that passes in isolation, and 1 is a stale adapter
1110
+ test (CROMA is now shipped). Re-running the affected files together gives **137 passed**, confirming the
1111
+ failures are attributable to the sandbox environment and test ordering rather than the code under test.
1112
+
1113
+ ### Reproduce a live run
1114
+
1115
+ The deployed stack is reachable. Note the authoring sandbox has a dead proxy, so outbound calls need
1116
+ `--noproxy '*'` (curl) or `ProxyHandler({})` (Python):
1117
+
1118
+ ```bash
1119
+ curl --noproxy '*' https://satquery-backend-m4yv.onrender.com/api/health
1120
+ curl --noproxy '*' https://satquery-backend-m4yv.onrender.com/api/capabilities
1121
+ ```
1122
+
1123
+ `/api/capabilities` returns six tasks, all `available: true`. A live run requires the tunnel agent to be
1124
+ connected (`agent_connected: true`); if the Codespace is stopped, the request parks until the tunnel
1125
+ timeout. See [`docs/DEPLOYMENT.md`](docs/DEPLOYMENT.md) §5–6.
1126
+
1127
+ ## Deployment
1128
+
1129
+ The live topology is **Cloudflare Pages → Render → outbound tunnel → GitHub Codespace**.
1130
+
1131
+ | Layer | Role | Source |
1132
+ |---|---|---|
1133
+ | Cloudflare Pages | Static frontend at **https://satquery.pages.dev** | `frontend/` |
1134
+ | Render | Orchestrator / API gateway, `/api/*`, CORS, wake flow | `gateway/` |
1135
+ | GitHub Codespace | FastAPI inference host, CPU, port 8000 | `app/` |
1136
+ | Hugging Face | Model cards, released artifacts, checksums | this release |
1137
+
1138
+ Deployment sources are **separate repositories** from this release. The wake flow is: Cloudflare →
1139
+ Render → start the Codespace if stopped → poll `/v1/health` → surface *"Waking inference engine…"* →
1140
+ `POST /infer` → result.
1141
+
1142
+ ### Live revisions at this release
1143
+
1144
+ | Component | Repository | Visibility | Revision | Host |
1145
+ |---|---|---|---|---|
1146
+ | Frontend | `Anish-lab-blip/SatQuery-Frontend` | private | **`2d7ae53b482d`** | Cloudflare Pages → `satquery.pages.dev` |
1147
+ | Backend / orchestrator | `Anish-lab-blip/SatQuery-Backend` | private | **`89d80eaddec5`** | Render → `satquery-backend-m4yv.onrender.com` |
1148
+ | Inference | `Anish-lab-blip/SatQuery-Inference` | private | **`5a0936ace491`** | Codespace, port 8000, via outbound tunnel |
1149
+ | Public umbrella | `Anish-lab-blip/SatQuery-AI` | **public** | `3dcabd32da41` | this release home |
1150
+
1151
+ > **Trap.** `deploy/` inside the monorepo is **stale and untracked**. It is **not** the deployed source.
1152
+ > Edits must go to the three real repositories. Recorded in
1153
+ > [`docs/DEPLOYMENT.md`](docs/DEPLOYMENT.md) §1.
1154
+
1155
+ ### Environment variables
1156
+
1157
+ Render (gateway), measured live:
1158
+
1159
+ | Variable | Value (live) |
1160
+ |---|---|
1161
+ | `CODESPACE_NAME` | `potential-space-trout-r4ppw969w45j2pvvw` |
1162
+ | `CODESPACE_PORT` | `8000` |
1163
+ | `SATQUERY_ALLOWED_ORIGINS` | `https://satquery.pages.dev` |
1164
+ | `SATQUERY_DEVICE` | `cpu` |
1165
+ | `SATQUERY_TRANSPORT` | `auto` |
1166
+ | `SATQUERY_TUNNEL_TIMEOUT_S` | `150` |
1167
+ | `SATQUERY_WAKE_TIMEOUT_S` | `120` |
1168
+ | `SATQUERY_UPSTREAM_TIMEOUT_S` | `90` |
1169
+ | `GITHUB_TOKEN` | present |
1170
+
1171
+ Codespace (inference):
1172
+
1173
+ | Variable | Purpose |
1174
+ |---|---|
1175
+ | `PORT` | platform-assigned; **must be read** |
1176
+ | `SATQUERY_DEVICE` | `cpu` \| `cuda` \| `mps` \| `null`; read **without importing torch** |
1177
+ | `SATQUERY_MAX_FILE_BYTES` | per-file cap (shared with Render) |
1178
+ | `SATQUERY_ASSET_ENABLED` / `SATQUERY_ASSET_DIR` | both required for `/v1/assets`; fails closed otherwise |
1179
+ | `SATQUERY_ASSET_MAX_FILES` / `SATQUERY_ASSET_TTL_S` | optional handle capacity / lifetime |
1180
+
1181
+ Deploy mechanics: the frontend is staged by `scripts/stage_pages.mjs` and deployed with
1182
+ `npx wrangler pages deploy`; the backend is a `render.yaml` blueprint whose `main.py` exposes `app`; the
1183
+ inference host runs `deploy/codespace/serve.py` on `$PORT` and the devcontainer forwards port 8000 and
1184
+ starts the tunnel agent on `postStartCommand`. Repository writes are performed through the GitHub Git
1185
+ Data API (blob → tree → commit → `PATCH` ref) with **sha256 byte-verification** of every uploaded blob,
1186
+ rather than `git push`, so each deployed file is verified by content hash.
1187
+
1188
+ ### Deployment caveats
1189
+
1190
+ - **Cold start.** The inference host may be stopped when idle. The first request after a cold start
1191
+ can exceed the client timeout while weights are fetched; a retry a few seconds later normally
1192
+ succeeds. Warm the stack before any demonstration and confirm
1193
+ `GET /api/health` reports `tunnel.agent_connected: true`. Cold start is **tens of seconds** and is
1194
+ documented rather than papered over.
1195
+ - **Tunnel gaps.** The tunnel agent can be briefly absent. A request issued during such a gap may hang
1196
+ or return HTTP 504. **This is not fixed in production** — a prepared patch
1197
+ (`forward_unavailable` 503 / `upstream_timeout` 504 plus a `codespace_name` fix) exists and is
1198
+ documented, but it was deliberately not deployed. Root cause: in `auto` transport mode a tunnel
1199
+ timeout falls through to the forwarded-port path (`main.py:546`), which then spends the 120 s wake
1200
+ timeout on an HTTP 302 — the observed ~249 s failure (150 + 120).
1201
+ - **`codespace_name`** is still reported with a trailing newline by `/api/health` (cosmetic; the wake
1202
+ path strips it).
1203
+ - **Never retry `POST /api/infer` at the gateway** — a retry consumes inference twice.
1204
+ - **Platform traps, recorded so they are not rediscovered.** Cloudflare `_headers` rules **concatenate**
1205
+ rather than override, and Chromium takes the first `max-age` it encounters, so a later rule cannot "fix"
1206
+ an earlier one. Cloudflare 308-redirects `X.html` → `/X`, so the extensionless path must be referenced.
1207
+ A forwarded Codespace port returns 302 for a private repo — which is *why* the tunnel exists. And the
1208
+ tunnel agent must be started by the devcontainer's `postStartCommand`, or a restarted Codespace comes up
1209
+ with `agent_connected: false`.
1210
+
1211
+ ### Historical context
1212
+
1213
+ The superseded design ran inference on an **HF Space with ZeroGPU** behind a **Railway** gateway. The
1214
+ active design moves to **Render + Codespace**, CPU-first, with an outbound tunnel. The four-endpoint
1215
+ contract, the gateway responsibility table, the env-var vocabulary and the config freeze are unchanged —
1216
+ only host names moved. `configs/deploy.yaml` still describes the old HF-Space/ZeroGPU target and is left
1217
+ **undisturbed as frozen paperwork** (editing it would move the config hash); no Gradio runtime exists in
1218
+ code. The declared ZeroGPU durations are transcribed, not invented — `app/space_app.py` carries
1219
+ `GPU_DURATIONS` = `vqa` 20, `caption` 20, `grounding` 45, `change` 30, `optical_sar` 45, `change_vqa` 30,
1220
+ and a task with no declared duration is a programming error rather than a default, because silently
1221
+ picking one would reserve the wrong amount of the 5 GPU-minute daily budget. The decoration has **never
1222
+ executed** here (`spaces` is not installed in this environment); `configs/deploy.yaml` sets
1223
+ `cpu_mode_required: true`, so a CPU run must work, and it does.
1224
+
1225
+ ## Reproducibility
1226
+
1227
+ 1. **Configuration.** `configs/base.yaml` is the single registry; no magic numbers in Python. Its hash
1228
+ is recorded in every execution trace. Frozen hash: `78f1e3700da15aa1`. A config edit moves the hash and
1229
+ invalidates every artifact keyed to it.
1230
+ 2. **Backbones.** Pinned by `repo_id` + `revision`, never by floating tag:
1231
+ `HuggingFaceTB/SmolVLM-500M-Instruct@a7da5b986cb5`,
1232
+ `chendelong/RemoteCLIP@bf1d8a3ccf2d`,
1233
+ `sentence-transformers/all-MiniLM-L6-v2@1110a243fdf4`,
1234
+ `antofuller/CROMA@0dd28e3d633b`.
1235
+ 3. **Released artifacts.** `models/manifest.json` and `models/checksums.sha256` are **generated from
1236
+ the actual files** — never hand-typed. Verify with `sha256sum -c models/checksums.sha256`.
1237
+ 4. **Splits.** LEVIR-CD-256: train 7120 / val 1024 / test 2048. Grounding: VRSBench n=16159.
1238
+ Fusion: held-out test n=4000. Change-VQA: test n=39686. Leakage isolation is by `scene_id`.
1239
+ 5. **Prompts** are versioned files, frozen before benchmark evaluation.
1240
+ 6. **Negative results are preserved.** Rejected and open rulings are recorded, not removed.
1241
+
1242
+ The frozen contract, in the registry's own vocabulary:
1243
+
1244
+ | Guarantee | How it is enforced |
1245
+ |---|---|
1246
+ | Frozen configuration | all tunables live in `configs/base.yaml`; the loader validates invariants and computes a hash |
1247
+ | Frozen config hash | `78f1e3700da15aa1`; every artifact records the hash it was produced against |
1248
+ | Pinned backbones | every backbone is pinned by revision; the Hub resolves the exact commit |
1249
+ | Seed | `project.seed: 42` |
1250
+ | Immutable public test | `evaluation.immutable_public_test: true`; `hidden_data_access: false` |
1251
+ | Byte-verified artifacts | every released artifact ships with a sha256 in `models/checksums.sha256` |
1252
+ | Verified metrics | every quoted number is checked against its artifact by the metric-verification tool |
1253
+
1254
+ ### Reproduce the metric check (cheap, no GPU)
1255
+
1256
+ ```bash
1257
+ python release/tools/verify_readme_metrics.py
1258
+ ```
1259
+
1260
+ It **reads** the artifacts under `artifacts/`, **compares** each of the 20 quoted metrics at the precision
1261
+ printed in this README, and **also asserts** the statuses (that the VLM headline contains
1262
+ `ACCEPTANCE-REJECTED`; the router's `corpus_limited` / `n_val`; the calibration temperature and
1263
+ `ece_improvement`). It **exits 0** and prints `ALL CLAIMS VERIFIED` only when everything matches.
1264
+
1265
+ ### What "reproduce" means, and what it does not
1266
+
1267
+ | Artifact | Where it trains | Reproducible from this release? |
1268
+ |---|---|---|
1269
+ | router adapter | local CPU | yes — `configs/base.yaml` §`router.training` |
1270
+ | grounding head | local | yes — `configs/base.yaml` §`grounding_training` |
1271
+ | change head | local | yes — `configs/base.yaml` §`change` |
1272
+ | optical_sar fusion head | local, seed sweep | yes (see [`docs/TRAINING.md`](docs/TRAINING.md) §5) |
1273
+ | change_vqa head | **external GPU (Kaggle)** | **partly** — the promotion gate, evaluation and serving wiring are reproducible; there is no one-command retrain |
1274
+ | vlm LoRA adapter | **external GPU** | **partly** — same |
1275
+
1276
+ For the two externally-trained artifacts, the repository reproduces the **promotion gate**
1277
+ (byte-identity, sha256, zero non-finite tensors), the **evaluation**, and the **serving wiring**; it does
1278
+ **not** ship a one-command retrain. That is stated rather than implied.
1279
+
1280
+ ### What is NOT reproducible from this release
1281
+
1282
+ | Item | Reason |
1283
+ |---|---|
1284
+ | The private deployment repos | they are private; the deployed sources are not in this release |
1285
+ | System-level end-to-end benchmark | **no such benchmark exists** |
1286
+ | Router test-split number | **not run** |
1287
+ | CDVQA / SECOND imagery | public but large; the release documents the acquisition + name-verification procedure, not the data |
1288
+ | BigEarthNet full corpus | **not downloaded** (only a 28k S2 subset was used) |
1289
+ | The historical ZeroGPU/Gradio deploy target | frozen paperwork only; no runtime exists in code |
1290
+
1291
+ Environment traps worth recording for anyone reproducing: the authoring sandbox has a **dead proxy**
1292
+ (outbound calls need `--noproxy '*'` for curl or `ProxyHandler({})` for Python); pytest is installed only
1293
+ in the repository virtualenv (`.venv/Scripts/python.exe`); the full-suite run trips the sandbox's
1294
+ bulk-delete guard; Cloudflare 308-redirects `X.html` → `/X`; and Chrome drops synthetic CDP key events
1295
+ when the window lacks OS focus, which is relevant to any browser-driven reproduction of the live
1296
+ validation.
1297
+
1298
+ ## Known limitations
1299
+
1300
+ 1. **No system-level end-to-end benchmark exists.** Per-specialist metrics are real; a single
1301
+ end-to-end number is **NOT RUN**.
1302
+ 2. **Router accuracy is validation-only** (n=86, corpus-limited). Its **test set was never run**.
1303
+ 3. **Optical-SAR fusion returns a bare class index** (`class_18`), not a human-readable label. The
1304
+ modality accounting in the response confirms the right channels reached the fusion head, but the
1305
+ presentation is not user-facing.
1306
+ 4. **VQA is weak-but-related.** Asked what terrain dominates a scene, it answers "Grassland".
1307
+ 5. **Fusion macro_F1 is low (0.434161)** against 0.931 accuracy — rare classes are poorly handled. Of
1308
+ the 19 classes, 5 are absent from the scored split and contribute 0.0 to macro-F1 by construction.
1309
+ 6. **Grounding IoU is modest** (0.2838 canonical, 0.2566 matched6) — useful, not solved — and it is
1310
+ protocol-sensitive: the argmax decode (0.1215) is barely above the zero-shot baseline (0.0972).
1311
+ 7. **Calibration makes ECE slightly worse** (0.013755 → 0.014929), and is retained only because it is
1312
+ part of the frozen configuration. The calibrated reliability curve is not plotted.
1313
+ 8. **Router lexical residuals.** *"What is the new runway?"* reads `change` rather than `vqa` (the
1314
+ `new`-as-change heuristic fires outside `where` questions), and *"How much built-up area was
1315
+ added?"* reads `vqa` (under-trigger). A lexical router cannot cleanly separate "the new X" from
1316
+ "what's new"; a trained intent router exists in `artifacts/router/` but is not attached.
1317
+ 9. **B-07 tunnel gaps are not fixed in production** (see [Deployment caveats](#deployment-caveats)).
1318
+ 10. **No license has been selected** for this repository. Until one is, the artifacts carry
1319
+ `license: unknown` and no reuse rights should be assumed. This is an open owner decision.
1320
+ 11. **The Anatomy of a Run page** renders a recorded run whose plate uses the 720×720 variant of an
1321
+ image analysed at 730×730 — identical content, scaled by the canvas, but the "actual analysed
1322
+ image" wording is slightly loose.
1323
+ 12. **The BigEarthNet local subset is single-label** (100 %) against the official 1–11 multi-label
1324
+ scheme, so its metrics are **not comparable** to published numbers.
1325
+ 13. **The VLM adapter is not accepted** — metrics usable (exact_match 0.963), status
1326
+ acceptance-rejected; the deployed path uses the unadapted model.
1327
+ 14. **The deployment repos are private**, so their links 404 for an outside audience — by design.
1328
+
1329
+ ### Explicit non-claims
1330
+
1331
+ - **No claim of state-of-the-art performance** on any benchmark.
1332
+ - **No claim of production readiness** for model quality — the deployment runs, but the models carry the
1333
+ limitations above.
1334
+ - **No claim that the trained heads generalise** beyond their training-family test splits.
1335
+ - **No claim that calibration improves confidence.**
1336
+ - **No claim that the VLM adapter is accepted** for production use.
1337
+ - **No system-level accuracy** is claimed anywhere, and none is produced by averaging the per-task
1338
+ numbers.
1339
+
1340
+ ## Links
1341
+
1342
+ | | |
1343
+ |---|---|
1344
+ | Live demo | https://satquery.pages.dev |
1345
+ | GitHub | https://github.com/Anish-lab-blip/SatQuery-AI |
1346
+ | Hugging Face | https://huggingface.co/thundercode/SatQuery |
1347
+
1348
+ Third-party models this work builds on (pinned, not redistributed):
1349
+
1350
+ | Model | Revision | Role |
1351
+ |---|---|---|
1352
+ | [`HuggingFaceTB/SmolVLM-500M-Instruct`](https://huggingface.co/HuggingFaceTB/SmolVLM-500M-Instruct) | `a7da5b986cb5` | VQA + captioning backbone |
1353
+ | [`chendelong/RemoteCLIP`](https://huggingface.co/chendelong/RemoteCLIP) | `bf1d8a3ccf2d` | remote-sensing grounding encoder |
1354
+ | [`sentence-transformers/all-MiniLM-L6-v2`](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) | `1110a243fdf4` | router embedding |
1355
+ | [`antofuller/CROMA`](https://huggingface.co/antofuller/CROMA) | `0dd28e3d633b` | optical/SAR fusion encoder |
1356
+
1357
+ Datasets referenced by the evaluations: LEVIR-CD-256 (change), VRSBench (grounding),
1358
+ BigEarthNet (optical-SAR fusion, 19 CLC classes), CDVQA + SECOND (change-VQA). No dataset is
1359
+ redistributed here.
1360
+
1361
+ Documentation in this repository:
1362
+
1363
+ | Document | Contents |
1364
+ |---|---|
1365
+ | [`docs/architecture/README.md`](docs/architecture/README.md) | architecture hub + index of the ten sub-documents |
1366
+ | [`docs/architecture/01-system-overview.md`](docs/architecture/01-system-overview.md) | thesis, component inventory, frozen backbones |
1367
+ | [`docs/architecture/02-deployment-topology.md`](docs/architecture/02-deployment-topology.md) | four tiers, tunnel, wake flow, cold start |
1368
+ | [`docs/architecture/03-request-lifecycle.md`](docs/architecture/03-request-lifecycle.md) | nine-state controller, validation, modality inference, tiling |
1369
+ | [`docs/architecture/04-router.md`](docs/architecture/04-router.md) | MiniLM, five heads, `interpret()` vs `chooseTask()`, the lexical fallback |
1370
+ | [`docs/architecture/05-specialists.md`](docs/architecture/05-specialists.md) | all six tasks end to end |
1371
+ | [`docs/architecture/06-evidence-and-confidence.md`](docs/architecture/06-evidence-and-confidence.md) | evidence schema, aggregation, temperature scaling, the eight events |
1372
+ | [`docs/architecture/07-configuration-freeze.md`](docs/architecture/07-configuration-freeze.md) | the registry, invariants, the config hash |
1373
+ | [`docs/architecture/08-api-contract.md`](docs/architecture/08-api-contract.md) | four endpoints, envelopes, error codes |
1374
+ | [`docs/architecture/09-frontend.md`](docs/architecture/09-frontend.md) | static pages, the Analyze console, real-vs-preview |
1375
+ | [`docs/architecture/10-observability-and-ops.md`](docs/architecture/10-observability-and-ops.md) | health, counters, traces |
1376
+ | [`docs/MODELS.md`](docs/MODELS.md) | the six artifacts in detail; rejected decisions |
1377
+ | [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md) | headline metrics and the rules they follow |
1378
+ | [`docs/EVALUATION.md`](docs/EVALUATION.md) | per-task protocols; evaluation-honesty rules |
1379
+ | [`docs/DATASETS.md`](docs/DATASETS.md) | measured corpus figures and caveats |
1380
+ | [`docs/TRAINING.md`](docs/TRAINING.md) | per-artifact hyperparameters |
1381
+ | [`docs/DEPLOYMENT.md`](docs/DEPLOYMENT.md) | live revisions, env vars, traps |
1382
+ | [`docs/REPRODUCIBILITY.md`](docs/REPRODUCIBILITY.md) | what a third party can reproduce |
1383
+ | [`docs/RESEARCH_NOTES.md`](docs/RESEARCH_NOTES.md) | findings, the router defect, negative results |
1384
+ | [`docs/LIMITATIONS.md`](docs/LIMITATIONS.md) | the full limitation catalogue |
1385
+ | [`docs/CHANGELOG.md`](docs/CHANGELOG.md) | versioned record |
1386
+ | [`MODEL_CARD.md`](MODEL_CARD.md) | the Hugging Face model card |
1387
+
1388
+ Supporting and subsystem documentation:
1389
+
1390
+ | Document | Contents |
1391
+ |---|---|
1392
+ | [`docs/SECURITY.md`](docs/SECURITY.md) | trust model, secret custody, CORS allowlist, limits, what is *not* defended |
1393
+ | [`docs/TESTING.md`](docs/TESTING.md) | suite inventory, the doc-guard tests, the harness false-positive lesson |
1394
+ | [`docs/GEOSPATIAL.md`](docs/GEOSPATIAL.md) | raster contract, CRS, the coordinate-system rule, band inference, normalisation |
1395
+ | [`docs/DATA_PIPELINE.md`](docs/DATA_PIPELINE.md) | asset → specialist input, cached features, sensor adapter |
1396
+ | [`docs/OPERATIONS.md`](docs/OPERATIONS.md) | the runbook, warming, incident triage, known operational gaps |
1397
+ | [`docs/DEVELOPMENT.md`](docs/DEVELOPMENT.md) | local setup, the config-hash rule, how to add a specialist, dev traps |
1398
+ | [`docs/PERFORMANCE.md`](docs/PERFORMANCE.md) | model footprint, measured component timings, cost traps |
1399
+ | [`docs/FRONTEND.md`](docs/FRONTEND.md) | the 11 pages, the Analyze console, the eight events, Cloudflare traps |
1400
+ | [`docs/SERVING.md`](docs/SERVING.md) | `build_space_app()`, the four endpoints, lazy loading, the G-1 trap |
1401
+ | [`docs/GLOSSARY.md`](docs/GLOSSARY.md) | every domain term and status vocabulary, defined |
1402
+ | [`docs/MASTER_ARCHITECTURE_PLAN.md`](docs/MASTER_ARCHITECTURE_PLAN.md) | the **original** master plan (historical; superseded in parts) |
1403
+ | [`HF_RELEASE_VERIFICATION.md`](HF_RELEASE_VERIFICATION.md) | the Hugging Face release verification record |
1404
+ | [`RELEASE_MANIFEST.md`](RELEASE_MANIFEST.md) | every released file with its size and sha256 |
1405
+ | [`tools/`](tools/) | the verification tooling (`verify_readme_metrics.py`, manifest generators, archive tools) |
1406
+
1407
+ > **Honesty rule.** Every document in this repository uses one status vocabulary —
1408
+ > `IMPLEMENTED` · `VERIFIED` · `MEASURED` · `ATTEMPTED` · `NOT RUN` · `BLOCKED` · `DEFERRED` ·
1409
+ > `REJECTED` · `OPEN` · `RESOLVED` · `CLOSED` — and negative results are recorded rather than
1410
+ > omitted. Where evidence was missing, the text says
1411
+ > `UNKNOWN — not established from the available evidence` instead of guessing.
1412
+
1413
+ ## Citation
1414
+
1415
+ No paper accompanies this release. Until one exists, cite the repository:
1416
+
1417
+ ```bibtex
1418
+ @misc{satqueryai2026,
1419
+ title = {SatQuery AI: an interactive vision-language assistant for
1420
+ multimodal remote-sensing image analysis},
1421
+ author = {SatQuery AI contributors},
1422
+ year = {2026},
1423
+ url = {https://github.com/Anish-lab-blip/SatQuery-AI}
1424
+ }
1425
+ ```
1426
+
1427
+ ## License
1428
+
1429
+ **Not yet selected.** See limitation 10. Backbone models remain under their own upstream licenses.
1430
+
1431
+ No `LICENSE` file exists in this repository. Until one is selected, the released artifacts carry
1432
+ `license: unknown` and **no reuse rights should be assumed**. This is an open owner decision, recorded
1433
+ as `OPEN` in [`docs/LIMITATIONS.md`](docs/LIMITATIONS.md) §5 and
1434
+ [`docs/CHANGELOG.md`](docs/CHANGELOG.md). The six trained artifacts are small modules over frozen
1435
+ backbones; the backbones are not redistributed here and remain under their own upstream licences —
1436
+ consult each backbone's Hugging Face page.
1437
+