SatQuery / docs /MODELS.md
thundercode's picture
release: add docs/MODELS.md
df0d288 verified
|
Raw History Blame
10.4 kB
# Models
**Status tags:** `IMPLEMENTED` Β· `VERIFIED` Β· `MEASURED` Β· `ATTEMPTED` Β· `NOT RUN` Β· `BLOCKED` Β·
`DEFERRED` Β· `REJECTED`.
SatQuery AI trains **six** artifacts. Four are task heads, one is a router adapter, one is a LoRA
adapter. Every backbone is **frozen** and publicly pinned by revision in `configs/base.yaml` β€” the
project trains small modules on top of frozen encoders, not end-to-end networks.
> **This file is the human-readable companion to the machine-generated
> [`../models/manifest.json`](../models/manifest.json) and
> [`../models/checksums.sha256`](../models/checksums.sha256) (Phase 3).** Where the two disagree,
> the generated manifest wins β€” it is computed from the files, this document is written by hand.
---
## 1. The six trained artifacts
| # | Task | Artifact path | Bytes | Kind | Backbone (frozen) |
|---|---|---|---|---|---|
| 1 | `change` | `artifacts/change/levir_change_v001/head.pt` | 63,231,009 | trained head | STANet-style, ResNet-18 + PAM |
| 2 | `change_vqa` | `artifacts/change_vqa/run/head.pt` | 5,822,809 | trained head | over the change encoder's features |
| 3 | `optical_sar` | `artifacts/optical_sar/fusion_head_production_v001/head.pt` | 14,427,457 | trained head (production) | CROMA-base (frozen), 19-class head |
| 4 | `grounding` | `artifacts/grounding/remoteclip_grounding_v001/head.pt` | 12,639,041 | trained head | RemoteCLIP ViT-B/32 (frozen) |
| 5 | `router` | `artifacts/router/router_adapter_v001/adapter.pt` | 211,961 | trained adapter | `all-MiniLM-L6-v2` (frozen) |
| 6 | `vlm` | `.scratch/phase6_real_adapter/phase6_adapter/adapter_model.safetensors` | 34,798,048 | **LoRA adapter** (PEFT) | `HuggingFaceTB/SmolVLM-500M-Instruct` (frozen) |
Training checkpoints also exist (`checkpoint_last.pt`, `checkpoint-1500`, `checkpoint-2000`) and are
**not** the released artifacts β€” they are archived as provenance.
## 2. Backbones β€” pinned, frozen, never retrained
| Role | Repository | Revision | Notes |
|---|---|---|---|
| Router encoder | `sentence-transformers/all-MiniLM-L6-v2` | `1110a243fdf4` | 90.9 MB, 22,713,216 params, 384-dim embeddings, tokenizer ceiling 256; truncation set to 128 |
| VLM | `HuggingFaceTB/SmolVLM-500M-Instruct` | `a7da5b986cb5` | ~1015 MB safetensors |
| Grounding | `chendelong/RemoteCLIP` (`RemoteCLIP-ViT-B-32.pt`) | `bf1d8a3ccf2d` | 605.2 MB; transformer width 768, **projected** dim 512 |
| Optical-SAR | `antofuller/CROMA` (`CROMA_base.pt`) | `0dd28e3d633b` | 777.6 MB; `encoder_dim` 768; `image_resolution` 120 |
These are resolved from the Hugging Face Hub on first use. **No backbone weights are redistributed**
by this project's release β€” see Β§6.
## 3. Per-artifact detail
### 3.1 Change (`change`) β€” `IMPLEMENTED`, `VERIFIED`
STANet-style Siamese change detector. Encoder ResNet-18, self-attention mode **PAM**, tile size 256,
threshold 0.50, minimum component 32 px. Loss is BCE (0.5) + Dice (0.5). Trained on LEVIR-CD-256
(train 7120 / val 1024 / test 2048).
**Measured** on the LEVIR-CD-256 test split (n = 2048, threshold 0.50):
| Metric | Value |
|---|---|
| pooled IoU | **0.8122** |
| macro IoU | **0.8457** |
| pooled F1 | **0.8964** |
This is the only task whose headline number carries the `VERIFIED` tag, because it is the only one
measured against a single, immutable public test split with a frozen threshold.
### 3.2 Change-VQA (`change_vqa`) β€” `IMPLEMENTED`, `MEASURED`, ruling **OPEN**
A head that answers natural-language change questions over a temporal pair. It is the dispatch
target for change-style questions when only one asset is attached (see
[`ARCHITECTURE.md`](ARCHITECTURE.md) Β§4).
**Measured on two test sets β€” both are reported; quoting only the better one would be selective:**
| Test set | accuracy | macro F1 |
|---|---|---|
| `test` (n = 39,686) | **0.697626** | **0.378373** |
| `test2` | **0.651469** | **0.372309** |
The wide gap between accuracy and macro-F1 means the head is carried by common classes and performs
poorly on rare ones. The ruling is **OPEN** β€” no promotion/acceptance decision has been recorded.
### 3.3 Optical-SAR fusion (`optical_sar`) β€” `IMPLEMENTED`, `MEASURED`, ruling **OPEN**
Uses frozen CROMA-base to produce optical (768), SAR (768) and joint (768) embeddings, concatenates
them with the 12 optical and 2 SAR channel descriptors, and feeds a 19-class head:
```
input_dim = 3 * 768 + 12 + 2 = 2318 β†’ hidden 512 β†’ num_classes 19 (BigEarthNet CLC)
```
The **availability mask is consumed by the fusion head, not by CROMA** β€” CROMA always sees the
canonical channel counts (12 optical, 2 SAR).
**Measured** (production head `fusion_head_production_v001`, pre-registered 115-class protocol,
held-out test n = 4000):
| Metric | Value |
|---|---|
| accuracy | **0.931** |
| macro F1 | **0.434161** |
**Both numbers must travel together.** The high accuracy with a low macro-F1 reflects class
imbalance across 19 classes. The pre-registered metric JSON records `macro_f1_denominator` and
`classes_present`/`classes_absent` so the denominator is auditable. The ruling is **OPEN**.
**Limitation:** the live service returns a bare class index (`class_18`), not a human-readable label.
### 3.4 Grounding (`grounding`) β€” `IMPLEMENTED`, `MEASURED` β€” **two protocols**
A trainable head over the frozen RemoteCLIP ViT-B/32 encoder. Per-cell feature is
`concat([patch, text, patchΒ·text, global_pool])` = `4 Γ— 512 = 2048` (enforced at config load). Cells
are assigned by ground-truth box centre (`cell_relative` decode). Objectness BCE is weighted **20Γ—**
because only ~1 of 49 cells is positive; unweighted, the optimum collapses to "no object".
Image resolution is frozen at **224** β€” 448 was evaluated and **rejected** (see Β§4).
**Measured on VRSBench (n = 16,159), reported under two protocols:**
| Protocol | mean best IoU | recall@0.5 |
|---|---|---|
| canonical (head threshold decode) | **0.2838** | **0.2198** |
| matched6 | **0.2566** | **0.1938** |
Two further decode variants exist and are reported for completeness β€” a reviewer must be able to see
the whole grid, not one cell of it:
| Variant | mean best IoU |
|---|---|
| head argmax decode (canonical) | **0.1215** |
| zero-shot matched (no trained head) | **0.0972** |
The trained head beats the zero-shot baseline by a wide margin, which is the point of the head; the
absolute IoU is low, which is an honest limitation.
### 3.5 Router (`router`) β€” `IMPLEMENTED`, `MEASURED`, **TEST NOT RUN**
A **50,822-parameter adapter** over the frozen MiniLM encoder. Because the encoder is frozen,
embeddings are cached and the adapter trains on cached vectors β€” no GPU required (measured: 20
epochs / 4,096 vectors in 0.28 s on CPU). Splits are by **group** (template / hard-negative family),
never by example; hard-negative families are placed in the test split so their accuracy measures
generalisation, not memorisation.
**Measured:** overall **ungated** task accuracy **0.965116** on the validation split, **n = 86**,
corpus-limited. This number is (a) validation-only, (b) ungated, and (c) small. The router **test**
split was **NOT RUN**. Do not read 0.965116 as a test result.
### 3.6 VLM LoRA (`vlm`) β€” `IMPLEMENTED`, `MEASURED`, **ACCEPTANCE-REJECTED**
A PEFT LoRA adapter on frozen SmolVLM-500M-Instruct: `peft_type=LORA`, `r=16`, `alpha=32`,
`dropout=0.05`, targeting `model.text_model.*.{q,k,v,o,gate,up,down}_proj`. PEFT 0.19.1.
**Measured** on a frozen 1,000-question subset: exact_match **0.963**, F1 **0.96432** (+49.5 pp over
the unadapted baseline).
**Yet the artifact's status is `CLOSED` with headline `ACCEPTANCE-REJECTED`.** This is not a
contradiction β€” it is the project's central truthfulness distinction:
- **`USABLE_VERIFIED`** β€” the adapter demonstrably works (the metrics are real and reproducible).
- **`ACCEPTANCE-REJECTED`** β€” the adapter is *not accepted* for production promotion, on grounds
recorded in `artifacts/vlm/phase6_closure.json` (`why_acceptance_rejected`).
The deployed caption/VQA path therefore uses the **unadapted** SmolVLM. Deployment success and model
acceptance are different claims, and this document keeps them apart.
## 4. Rejected and deferred model decisions
| Decision | Outcome | Evidence |
|---|---|---|
| Grounding image resolution 448 vs 224 | **224 chosen; 448 REJECTED** | 448 lost on every axis: mean best IoU βˆ’0.0147, recall@0.5 βˆ’0.0022, all recall thresholds lower, at 1.59Γ— latency. Paired test: mean diff βˆ’0.0147, 95 % CI [βˆ’0.0160, βˆ’0.0134], t = βˆ’22.63; 448 better on 8.5 % of records, worse on 20.9 %. Pre-registered rule and the paired test **agree**. |
| VLM adapter promotion | **REJECTED** | metrics usable, acceptance rejected (Β§3.6) |
| Calibration | **kept but ineffective** | see Β§5 |
| optical-SAR / change-VQA rulings | **OPEN** | no decision recorded |
## 5. Calibration β€” `MEASURED`, **not an improvement**
Temperature scaling is enabled (`confidence.temperature_scaling: true`) with
`calibration_v001.json`. Fitted temperature **T = 0.9772732** on the validation split (n = 16,441).
| Metric | Before | After |
|---|---|---|
| ECE | 0.013755 | **0.014929** |
| NLL | 0.689741 | 0.689631 |
**ECE got worse** (`ece_improvement = βˆ’0.001174`). The scaling is retained because it is part of the
frozen configuration, **not** because it helped. The reliability diagram on the Benchmark page is
explicitly labelled **pre-scaling** so a reader cannot mistake it for the calibrated result. This is
recorded as a negative result, not smoothed over.
## 6. Distribution and licensing
- **Backbones are not redistributed.** They are fetched from the Hugging Face Hub at run time, pinned
by revision. Their licences are their own (see each model's HF page).
- **The six trained artifacts are published by this project** on the Hugging Face Hub under
`thundercode/SatQuery`, labelled by kind, each with its backbone dependency documented and each
accompanied by a checksum. See [`../HF_RELEASE_VERIFICATION.md`](../HF_RELEASE_VERIFICATION.md).
- **No licence file exists in the source repository.** This is an **OPEN** item flagged in
[`LIMITATIONS.md`](LIMITATIONS.md); the repository README instructs the owner to select one before
any public release of *code*. Model weights carry the terms of their backbone licences.