Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download docs/MODELS.md from thundercode/SatQuery: direct link, hf CLI and curl.
- Browser
- Download file 10.4 kB
-
https://huggingface.co/thundercode/SatQuery/resolve/394471d520abaa95b85b18eddbdbd65d55f7d996/docs/MODELS.md
- Command line
-
hf download hf://thundercode/SatQuery@394471d520abaa95b85b18eddbdbd65d55f7d996/docs/MODELS.md
-
curl -L -o MODELS.md https://huggingface.co/thundercode/SatQuery/resolve/394471d520abaa95b85b18eddbdbd65d55f7d996/docs/MODELS.md
10.4 kB
| # Models | |
| **Status tags:** `IMPLEMENTED` Β· `VERIFIED` Β· `MEASURED` Β· `ATTEMPTED` Β· `NOT RUN` Β· `BLOCKED` Β· | |
| `DEFERRED` Β· `REJECTED`. | |
| SatQuery AI trains **six** artifacts. Four are task heads, one is a router adapter, one is a LoRA | |
| adapter. Every backbone is **frozen** and publicly pinned by revision in `configs/base.yaml` β the | |
| project trains small modules on top of frozen encoders, not end-to-end networks. | |
| > **This file is the human-readable companion to the machine-generated | |
| > [`../models/manifest.json`](../models/manifest.json) and | |
| > [`../models/checksums.sha256`](../models/checksums.sha256) (Phase 3).** Where the two disagree, | |
| > the generated manifest wins β it is computed from the files, this document is written by hand. | |
| --- | |
| ## 1. The six trained artifacts | |
| | # | Task | Artifact path | Bytes | Kind | Backbone (frozen) | | |
| |---|---|---|---|---|---| | |
| | 1 | `change` | `artifacts/change/levir_change_v001/head.pt` | 63,231,009 | trained head | STANet-style, ResNet-18 + PAM | | |
| | 2 | `change_vqa` | `artifacts/change_vqa/run/head.pt` | 5,822,809 | trained head | over the change encoder's features | | |
| | 3 | `optical_sar` | `artifacts/optical_sar/fusion_head_production_v001/head.pt` | 14,427,457 | trained head (production) | CROMA-base (frozen), 19-class head | | |
| | 4 | `grounding` | `artifacts/grounding/remoteclip_grounding_v001/head.pt` | 12,639,041 | trained head | RemoteCLIP ViT-B/32 (frozen) | | |
| | 5 | `router` | `artifacts/router/router_adapter_v001/adapter.pt` | 211,961 | trained adapter | `all-MiniLM-L6-v2` (frozen) | | |
| | 6 | `vlm` | `.scratch/phase6_real_adapter/phase6_adapter/adapter_model.safetensors` | 34,798,048 | **LoRA adapter** (PEFT) | `HuggingFaceTB/SmolVLM-500M-Instruct` (frozen) | | |
| Training checkpoints also exist (`checkpoint_last.pt`, `checkpoint-1500`, `checkpoint-2000`) and are | |
| **not** the released artifacts β they are archived as provenance. | |
| ## 2. Backbones β pinned, frozen, never retrained | |
| | Role | Repository | Revision | Notes | | |
| |---|---|---|---| | |
| | Router encoder | `sentence-transformers/all-MiniLM-L6-v2` | `1110a243fdf4` | 90.9 MB, 22,713,216 params, 384-dim embeddings, tokenizer ceiling 256; truncation set to 128 | | |
| | VLM | `HuggingFaceTB/SmolVLM-500M-Instruct` | `a7da5b986cb5` | ~1015 MB safetensors | | |
| | Grounding | `chendelong/RemoteCLIP` (`RemoteCLIP-ViT-B-32.pt`) | `bf1d8a3ccf2d` | 605.2 MB; transformer width 768, **projected** dim 512 | | |
| | Optical-SAR | `antofuller/CROMA` (`CROMA_base.pt`) | `0dd28e3d633b` | 777.6 MB; `encoder_dim` 768; `image_resolution` 120 | | |
| These are resolved from the Hugging Face Hub on first use. **No backbone weights are redistributed** | |
| by this project's release β see Β§6. | |
| ## 3. Per-artifact detail | |
| ### 3.1 Change (`change`) β `IMPLEMENTED`, `VERIFIED` | |
| STANet-style Siamese change detector. Encoder ResNet-18, self-attention mode **PAM**, tile size 256, | |
| threshold 0.50, minimum component 32 px. Loss is BCE (0.5) + Dice (0.5). Trained on LEVIR-CD-256 | |
| (train 7120 / val 1024 / test 2048). | |
| **Measured** on the LEVIR-CD-256 test split (n = 2048, threshold 0.50): | |
| | Metric | Value | | |
| |---|---| | |
| | pooled IoU | **0.8122** | | |
| | macro IoU | **0.8457** | | |
| | pooled F1 | **0.8964** | | |
| This is the only task whose headline number carries the `VERIFIED` tag, because it is the only one | |
| measured against a single, immutable public test split with a frozen threshold. | |
| ### 3.2 Change-VQA (`change_vqa`) β `IMPLEMENTED`, `MEASURED`, ruling **OPEN** | |
| A head that answers natural-language change questions over a temporal pair. It is the dispatch | |
| target for change-style questions when only one asset is attached (see | |
| [`ARCHITECTURE.md`](ARCHITECTURE.md) Β§4). | |
| **Measured on two test sets β both are reported; quoting only the better one would be selective:** | |
| | Test set | accuracy | macro F1 | | |
| |---|---|---| | |
| | `test` (n = 39,686) | **0.697626** | **0.378373** | | |
| | `test2` | **0.651469** | **0.372309** | | |
| The wide gap between accuracy and macro-F1 means the head is carried by common classes and performs | |
| poorly on rare ones. The ruling is **OPEN** β no promotion/acceptance decision has been recorded. | |
| ### 3.3 Optical-SAR fusion (`optical_sar`) β `IMPLEMENTED`, `MEASURED`, ruling **OPEN** | |
| Uses frozen CROMA-base to produce optical (768), SAR (768) and joint (768) embeddings, concatenates | |
| them with the 12 optical and 2 SAR channel descriptors, and feeds a 19-class head: | |
| ``` | |
| input_dim = 3 * 768 + 12 + 2 = 2318 β hidden 512 β num_classes 19 (BigEarthNet CLC) | |
| ``` | |
| The **availability mask is consumed by the fusion head, not by CROMA** β CROMA always sees the | |
| canonical channel counts (12 optical, 2 SAR). | |
| **Measured** (production head `fusion_head_production_v001`, pre-registered 115-class protocol, | |
| held-out test n = 4000): | |
| | Metric | Value | | |
| |---|---| | |
| | accuracy | **0.931** | | |
| | macro F1 | **0.434161** | | |
| **Both numbers must travel together.** The high accuracy with a low macro-F1 reflects class | |
| imbalance across 19 classes. The pre-registered metric JSON records `macro_f1_denominator` and | |
| `classes_present`/`classes_absent` so the denominator is auditable. The ruling is **OPEN**. | |
| **Limitation:** the live service returns a bare class index (`class_18`), not a human-readable label. | |
| ### 3.4 Grounding (`grounding`) β `IMPLEMENTED`, `MEASURED` β **two protocols** | |
| A trainable head over the frozen RemoteCLIP ViT-B/32 encoder. Per-cell feature is | |
| `concat([patch, text, patchΒ·text, global_pool])` = `4 Γ 512 = 2048` (enforced at config load). Cells | |
| are assigned by ground-truth box centre (`cell_relative` decode). Objectness BCE is weighted **20Γ** | |
| because only ~1 of 49 cells is positive; unweighted, the optimum collapses to "no object". | |
| Image resolution is frozen at **224** β 448 was evaluated and **rejected** (see Β§4). | |
| **Measured on VRSBench (n = 16,159), reported under two protocols:** | |
| | Protocol | mean best IoU | recall@0.5 | | |
| |---|---|---| | |
| | canonical (head threshold decode) | **0.2838** | **0.2198** | | |
| | matched6 | **0.2566** | **0.1938** | | |
| Two further decode variants exist and are reported for completeness β a reviewer must be able to see | |
| the whole grid, not one cell of it: | |
| | Variant | mean best IoU | | |
| |---|---| | |
| | head argmax decode (canonical) | **0.1215** | | |
| | zero-shot matched (no trained head) | **0.0972** | | |
| The trained head beats the zero-shot baseline by a wide margin, which is the point of the head; the | |
| absolute IoU is low, which is an honest limitation. | |
| ### 3.5 Router (`router`) β `IMPLEMENTED`, `MEASURED`, **TEST NOT RUN** | |
| A **50,822-parameter adapter** over the frozen MiniLM encoder. Because the encoder is frozen, | |
| embeddings are cached and the adapter trains on cached vectors β no GPU required (measured: 20 | |
| epochs / 4,096 vectors in 0.28 s on CPU). Splits are by **group** (template / hard-negative family), | |
| never by example; hard-negative families are placed in the test split so their accuracy measures | |
| generalisation, not memorisation. | |
| **Measured:** overall **ungated** task accuracy **0.965116** on the validation split, **n = 86**, | |
| corpus-limited. This number is (a) validation-only, (b) ungated, and (c) small. The router **test** | |
| split was **NOT RUN**. Do not read 0.965116 as a test result. | |
| ### 3.6 VLM LoRA (`vlm`) β `IMPLEMENTED`, `MEASURED`, **ACCEPTANCE-REJECTED** | |
| A PEFT LoRA adapter on frozen SmolVLM-500M-Instruct: `peft_type=LORA`, `r=16`, `alpha=32`, | |
| `dropout=0.05`, targeting `model.text_model.*.{q,k,v,o,gate,up,down}_proj`. PEFT 0.19.1. | |
| **Measured** on a frozen 1,000-question subset: exact_match **0.963**, F1 **0.96432** (+49.5 pp over | |
| the unadapted baseline). | |
| **Yet the artifact's status is `CLOSED` with headline `ACCEPTANCE-REJECTED`.** This is not a | |
| contradiction β it is the project's central truthfulness distinction: | |
| - **`USABLE_VERIFIED`** β the adapter demonstrably works (the metrics are real and reproducible). | |
| - **`ACCEPTANCE-REJECTED`** β the adapter is *not accepted* for production promotion, on grounds | |
| recorded in `artifacts/vlm/phase6_closure.json` (`why_acceptance_rejected`). | |
| The deployed caption/VQA path therefore uses the **unadapted** SmolVLM. Deployment success and model | |
| acceptance are different claims, and this document keeps them apart. | |
| ## 4. Rejected and deferred model decisions | |
| | Decision | Outcome | Evidence | | |
| |---|---|---| | |
| | Grounding image resolution 448 vs 224 | **224 chosen; 448 REJECTED** | 448 lost on every axis: mean best IoU β0.0147, recall@0.5 β0.0022, all recall thresholds lower, at 1.59Γ latency. Paired test: mean diff β0.0147, 95 % CI [β0.0160, β0.0134], t = β22.63; 448 better on 8.5 % of records, worse on 20.9 %. Pre-registered rule and the paired test **agree**. | | |
| | VLM adapter promotion | **REJECTED** | metrics usable, acceptance rejected (Β§3.6) | | |
| | Calibration | **kept but ineffective** | see Β§5 | | |
| | optical-SAR / change-VQA rulings | **OPEN** | no decision recorded | | |
| ## 5. Calibration β `MEASURED`, **not an improvement** | |
| Temperature scaling is enabled (`confidence.temperature_scaling: true`) with | |
| `calibration_v001.json`. Fitted temperature **T = 0.9772732** on the validation split (n = 16,441). | |
| | Metric | Before | After | | |
| |---|---|---| | |
| | ECE | 0.013755 | **0.014929** | | |
| | NLL | 0.689741 | 0.689631 | | |
| **ECE got worse** (`ece_improvement = β0.001174`). The scaling is retained because it is part of the | |
| frozen configuration, **not** because it helped. The reliability diagram on the Benchmark page is | |
| explicitly labelled **pre-scaling** so a reader cannot mistake it for the calibrated result. This is | |
| recorded as a negative result, not smoothed over. | |
| ## 6. Distribution and licensing | |
| - **Backbones are not redistributed.** They are fetched from the Hugging Face Hub at run time, pinned | |
| by revision. Their licences are their own (see each model's HF page). | |
| - **The six trained artifacts are published by this project** on the Hugging Face Hub under | |
| `thundercode/SatQuery`, labelled by kind, each with its backbone dependency documented and each | |
| accompanied by a checksum. See [`../HF_RELEASE_VERIFICATION.md`](../HF_RELEASE_VERIFICATION.md). | |
| - **No licence file exists in the source repository.** This is an **OPEN** item flagged in | |
| [`LIMITATIONS.md`](LIMITATIONS.md); the repository README instructs the owner to select one before | |
| any public release of *code*. Model weights carry the terms of their backbone licences. | |