SatQuery / docs /MODELS.md
thundercode's picture
release: add docs/MODELS.md
df0d288 verified
|
Raw
History Blame
10.4 kB

Models

Status tags: IMPLEMENTED · VERIFIED · MEASURED · ATTEMPTED · NOT RUN · BLOCKED · DEFERRED · REJECTED.

SatQuery AI trains six artifacts. Four are task heads, one is a router adapter, one is a LoRA adapter. Every backbone is frozen and publicly pinned by revision in configs/base.yaml — the project trains small modules on top of frozen encoders, not end-to-end networks.

This file is the human-readable companion to the machine-generated ../models/manifest.json and ../models/checksums.sha256 (Phase 3). Where the two disagree, the generated manifest wins — it is computed from the files, this document is written by hand.


1. The six trained artifacts

# Task Artifact path Bytes Kind Backbone (frozen)
1 change artifacts/change/levir_change_v001/head.pt 63,231,009 trained head STANet-style, ResNet-18 + PAM
2 change_vqa artifacts/change_vqa/run/head.pt 5,822,809 trained head over the change encoder's features
3 optical_sar artifacts/optical_sar/fusion_head_production_v001/head.pt 14,427,457 trained head (production) CROMA-base (frozen), 19-class head
4 grounding artifacts/grounding/remoteclip_grounding_v001/head.pt 12,639,041 trained head RemoteCLIP ViT-B/32 (frozen)
5 router artifacts/router/router_adapter_v001/adapter.pt 211,961 trained adapter all-MiniLM-L6-v2 (frozen)
6 vlm .scratch/phase6_real_adapter/phase6_adapter/adapter_model.safetensors 34,798,048 LoRA adapter (PEFT) HuggingFaceTB/SmolVLM-500M-Instruct (frozen)

Training checkpoints also exist (checkpoint_last.pt, checkpoint-1500, checkpoint-2000) and are not the released artifacts — they are archived as provenance.

2. Backbones — pinned, frozen, never retrained

Role Repository Revision Notes
Router encoder sentence-transformers/all-MiniLM-L6-v2 1110a243fdf4 90.9 MB, 22,713,216 params, 384-dim embeddings, tokenizer ceiling 256; truncation set to 128
VLM HuggingFaceTB/SmolVLM-500M-Instruct a7da5b986cb5 ~1015 MB safetensors
Grounding chendelong/RemoteCLIP (RemoteCLIP-ViT-B-32.pt) bf1d8a3ccf2d 605.2 MB; transformer width 768, projected dim 512
Optical-SAR antofuller/CROMA (CROMA_base.pt) 0dd28e3d633b 777.6 MB; encoder_dim 768; image_resolution 120

These are resolved from the Hugging Face Hub on first use. No backbone weights are redistributed by this project's release — see §6.

3. Per-artifact detail

3.1 Change (change) — IMPLEMENTED, VERIFIED

STANet-style Siamese change detector. Encoder ResNet-18, self-attention mode PAM, tile size 256, threshold 0.50, minimum component 32 px. Loss is BCE (0.5) + Dice (0.5). Trained on LEVIR-CD-256 (train 7120 / val 1024 / test 2048).

Measured on the LEVIR-CD-256 test split (n = 2048, threshold 0.50):

Metric Value
pooled IoU 0.8122
macro IoU 0.8457
pooled F1 0.8964

This is the only task whose headline number carries the VERIFIED tag, because it is the only one measured against a single, immutable public test split with a frozen threshold.

3.2 Change-VQA (change_vqa) — IMPLEMENTED, MEASURED, ruling OPEN

A head that answers natural-language change questions over a temporal pair. It is the dispatch target for change-style questions when only one asset is attached (see ARCHITECTURE.md §4).

Measured on two test sets — both are reported; quoting only the better one would be selective:

Test set accuracy macro F1
test (n = 39,686) 0.697626 0.378373
test2 0.651469 0.372309

The wide gap between accuracy and macro-F1 means the head is carried by common classes and performs poorly on rare ones. The ruling is OPEN — no promotion/acceptance decision has been recorded.

3.3 Optical-SAR fusion (optical_sar) — IMPLEMENTED, MEASURED, ruling OPEN

Uses frozen CROMA-base to produce optical (768), SAR (768) and joint (768) embeddings, concatenates them with the 12 optical and 2 SAR channel descriptors, and feeds a 19-class head:

input_dim = 3 * 768 + 12 + 2 = 2318   →   hidden 512   →   num_classes 19   (BigEarthNet CLC)

The availability mask is consumed by the fusion head, not by CROMA — CROMA always sees the canonical channel counts (12 optical, 2 SAR).

Measured (production head fusion_head_production_v001, pre-registered 115-class protocol, held-out test n = 4000):

Metric Value
accuracy 0.931
macro F1 0.434161

Both numbers must travel together. The high accuracy with a low macro-F1 reflects class imbalance across 19 classes. The pre-registered metric JSON records macro_f1_denominator and classes_present/classes_absent so the denominator is auditable. The ruling is OPEN.

Limitation: the live service returns a bare class index (class_18), not a human-readable label.

3.4 Grounding (grounding) — IMPLEMENTED, MEASURED — two protocols

A trainable head over the frozen RemoteCLIP ViT-B/32 encoder. Per-cell feature is concat([patch, text, patch·text, global_pool]) = 4 × 512 = 2048 (enforced at config load). Cells are assigned by ground-truth box centre (cell_relative decode). Objectness BCE is weighted 20× because only ~1 of 49 cells is positive; unweighted, the optimum collapses to "no object".

Image resolution is frozen at 224 — 448 was evaluated and rejected (see §4).

Measured on VRSBench (n = 16,159), reported under two protocols:

Protocol mean best IoU recall@0.5
canonical (head threshold decode) 0.2838 0.2198
matched6 0.2566 0.1938

Two further decode variants exist and are reported for completeness — a reviewer must be able to see the whole grid, not one cell of it:

Variant mean best IoU
head argmax decode (canonical) 0.1215
zero-shot matched (no trained head) 0.0972

The trained head beats the zero-shot baseline by a wide margin, which is the point of the head; the absolute IoU is low, which is an honest limitation.

3.5 Router (router) — IMPLEMENTED, MEASURED, TEST NOT RUN

A 50,822-parameter adapter over the frozen MiniLM encoder. Because the encoder is frozen, embeddings are cached and the adapter trains on cached vectors — no GPU required (measured: 20 epochs / 4,096 vectors in 0.28 s on CPU). Splits are by group (template / hard-negative family), never by example; hard-negative families are placed in the test split so their accuracy measures generalisation, not memorisation.

Measured: overall ungated task accuracy 0.965116 on the validation split, n = 86, corpus-limited. This number is (a) validation-only, (b) ungated, and (c) small. The router test split was NOT RUN. Do not read 0.965116 as a test result.

3.6 VLM LoRA (vlm) — IMPLEMENTED, MEASURED, ACCEPTANCE-REJECTED

A PEFT LoRA adapter on frozen SmolVLM-500M-Instruct: peft_type=LORA, r=16, alpha=32, dropout=0.05, targeting model.text_model.*.{q,k,v,o,gate,up,down}_proj. PEFT 0.19.1.

Measured on a frozen 1,000-question subset: exact_match 0.963, F1 0.96432 (+49.5 pp over the unadapted baseline).

Yet the artifact's status is CLOSED with headline ACCEPTANCE-REJECTED. This is not a contradiction — it is the project's central truthfulness distinction:

  • USABLE_VERIFIED — the adapter demonstrably works (the metrics are real and reproducible).
  • ACCEPTANCE-REJECTED — the adapter is not accepted for production promotion, on grounds recorded in artifacts/vlm/phase6_closure.json (why_acceptance_rejected).

The deployed caption/VQA path therefore uses the unadapted SmolVLM. Deployment success and model acceptance are different claims, and this document keeps them apart.

4. Rejected and deferred model decisions

Decision Outcome Evidence
Grounding image resolution 448 vs 224 224 chosen; 448 REJECTED 448 lost on every axis: mean best IoU −0.0147, recall@0.5 −0.0022, all recall thresholds lower, at 1.59× latency. Paired test: mean diff −0.0147, 95 % CI [−0.0160, −0.0134], t = −22.63; 448 better on 8.5 % of records, worse on 20.9 %. Pre-registered rule and the paired test agree.
VLM adapter promotion REJECTED metrics usable, acceptance rejected (§3.6)
Calibration kept but ineffective see §5
optical-SAR / change-VQA rulings OPEN no decision recorded

5. Calibration — MEASURED, not an improvement

Temperature scaling is enabled (confidence.temperature_scaling: true) with calibration_v001.json. Fitted temperature T = 0.9772732 on the validation split (n = 16,441).

Metric Before After
ECE 0.013755 0.014929
NLL 0.689741 0.689631

ECE got worse (ece_improvement = −0.001174). The scaling is retained because it is part of the frozen configuration, not because it helped. The reliability diagram on the Benchmark page is explicitly labelled pre-scaling so a reader cannot mistake it for the calibrated result. This is recorded as a negative result, not smoothed over.

6. Distribution and licensing

  • Backbones are not redistributed. They are fetched from the Hugging Face Hub at run time, pinned by revision. Their licences are their own (see each model's HF page).
  • The six trained artifacts are published by this project on the Hugging Face Hub under thundercode/SatQuery, labelled by kind, each with its backbone dependency documented and each accompanied by a checksum. See ../HF_RELEASE_VERIFICATION.md.
  • No licence file exists in the source repository. This is an OPEN item flagged in LIMITATIONS.md; the repository README instructs the owner to select one before any public release of code. Model weights carry the terms of their backbone licences.