Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
release: add docs/BENCHMARKS.md
Browse files- docs/BENCHMARKS.md +100 -0
docs/BENCHMARKS.md
ADDED
|
@@ -0,0 +1,100 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Benchmarks
|
| 2 |
+
|
| 3 |
+
**Every number here is artifact-backed.** Where a metric exists, its source file is named. Where a
|
| 4 |
+
benchmark does **not** exist, that is stated explicitly and tagged `NOT RUN` — an absent number is
|
| 5 |
+
never silently omitted or replaced with an estimate.
|
| 6 |
+
|
| 7 |
+
**Status tags:** `VERIFIED` · `MEASURED` · `NOT RUN` · `OPEN` · `REJECTED`.
|
| 8 |
+
|
| 9 |
+
---
|
| 10 |
+
|
| 11 |
+
## 1. Headline table
|
| 12 |
+
|
| 13 |
+
| Capability | Metric | Value | Split / protocol | Status |
|
| 14 |
+
|---|---|---|---|---|
|
| 15 |
+
| Change detection | pooled IoU | **0.8122** | LEVIR-CD-256 test, n = 2048, thr 0.50 | **VERIFIED** |
|
| 16 |
+
| Change detection | macro IoU | **0.8457** | same | **VERIFIED** |
|
| 17 |
+
| Change detection | pooled F1 | **0.8964** | same | **VERIFIED** |
|
| 18 |
+
| Grounding | mean best IoU | **0.2838** / **0.2566** | VRSBench n = 16159, canonical / matched6 | MEASURED (2 protocols) |
|
| 19 |
+
| Grounding | recall@0.5 | **0.2198** / **0.1938** | same | MEASURED (2 protocols) |
|
| 20 |
+
| Grounding | head-argmax IoU | **0.1215** | canonical | MEASURED |
|
| 21 |
+
| Grounding | zero-shot baseline IoU | **0.0972** | canonical | MEASURED (baseline) |
|
| 22 |
+
| Optical-SAR fusion | accuracy | **0.931** | held-out test n = 4000, 19 classes | MEASURED, ruling **OPEN** |
|
| 23 |
+
| Optical-SAR fusion | macro F1 | **0.434161** | same | MEASURED, ruling **OPEN** |
|
| 24 |
+
| Change-VQA | accuracy | **0.697626** | test n = 39686 | MEASURED, ruling **OPEN** |
|
| 25 |
+
| Change-VQA | macro F1 | **0.378373** | same | MEASURED, ruling **OPEN** |
|
| 26 |
+
| Change-VQA | accuracy (2nd set) | **0.651469** | test2 | MEASURED, ruling **OPEN** |
|
| 27 |
+
| Change-VQA | macro F1 (2nd set) | **0.372309** | test2 | MEASURED, ruling **OPEN** |
|
| 28 |
+
| VLM (adapted) | exact_match | **0.963** | frozen 1000-question subset | MEASURED, **ACCEPTANCE-REJECTED** |
|
| 29 |
+
| VLM (adapted) | F1 | **0.96432** | same | MEASURED, **ACCEPTANCE-REJECTED** |
|
| 30 |
+
| Router | overall **ungated** accuracy | **0.965116** | val n = 86, corpus-limited | MEASURED — **TEST NOT RUN** |
|
| 31 |
+
| Calibration | ECE before / after | **0.013755 → 0.014929** | val n = 16441, T = 0.9773 | MEASURED — **worse** |
|
| 32 |
+
|
| 33 |
+
### Source artifacts
|
| 34 |
+
|
| 35 |
+
| Metric family | Artifact |
|
| 36 |
+
|---|---|
|
| 37 |
+
| change | `artifacts/change/eval_test/eval_result.json` |
|
| 38 |
+
| grounding | `artifacts/grounding/remoteclip_grounding_v001/eval_result_canonical.json`, `…_matched6.json` |
|
| 39 |
+
| optical-SAR | `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json` |
|
| 40 |
+
| change-VQA | `artifacts/change_vqa/run/PROMOTION.json` |
|
| 41 |
+
| router | `artifacts/router/threshold_sweep_val.json` |
|
| 42 |
+
| calibration | `artifacts/calibration_v001.json` |
|
| 43 |
+
| VLM | `artifacts/vlm/phase6_closure.json` |
|
| 44 |
+
|
| 45 |
+
> All 20 values in the table above are checked against these files by
|
| 46 |
+
> [`../tools/verify_readme_metrics.py`](../tools/verify_readme_metrics.py). Its output
|
| 47 |
+
> (`ALL CLAIMS VERIFIED`) is committed as `../tools/readme_metrics_report.txt`.
|
| 48 |
+
|
| 49 |
+
## 2. Rules this table follows
|
| 50 |
+
|
| 51 |
+
1. **Two protocols are never collapsed.** Grounding is reported under *both* the canonical and
|
| 52 |
+
matched6 protocols. Quoting 0.2838 alone would be selective.
|
| 53 |
+
2. **Two test sets are never collapsed.** Change-VQA is reported on `test` **and** `test2`.
|
| 54 |
+
3. **accuracy never travels without macro-F1.** For imbalanced multi-class heads (optical-SAR,
|
| 55 |
+
change-VQA) the macro-F1 is reported alongside accuracy, always.
|
| 56 |
+
4. **Validation is not test.** The router number is labelled "overall **ungated** accuracy", val,
|
| 57 |
+
n = 86. It is not a test result.
|
| 58 |
+
5. **A negative result stays negative.** Calibration ECE worsened and is shown worsening.
|
| 59 |
+
6. **USABLE ≠ ACCEPTED.** The VLM metrics are real; the artifact is nevertheless acceptance-rejected.
|
| 60 |
+
|
| 61 |
+
## 3. What is NOT benchmarked
|
| 62 |
+
|
| 63 |
+
| Benchmark | Status | Note |
|
| 64 |
+
|---|---|---|
|
| 65 |
+
| **System-level end-to-end accuracy** | **NOT RUN — none exists** | There is no measured end-to-end benchmark of the full router→specialist→envelope pipeline. No such number is claimed anywhere. |
|
| 66 |
+
| **Router test split** | **NOT RUN** | Only the validation split (n = 86) was scored. |
|
| 67 |
+
| **Benchmark adapters** | **NOT RUN** | Adapter-based benchmark runs were not executed. |
|
| 68 |
+
| **Efficiency / latency benchmark** | not systematically measured | Per-specialist latency is recorded incidentally in artifacts (e.g. grounding `latency_ms_per_image` 2.205 ms for the head), but there is no end-to-end latency benchmark. |
|
| 69 |
+
| **Cross-dataset generalisation** | **NOT RUN** | Each specialist is evaluated only on its own training-family test split. |
|
| 70 |
+
|
| 71 |
+
## 4. Live validation (behavioural, not accuracy)
|
| 72 |
+
|
| 73 |
+
Accuracy is separate from **behavioural** validation. The deployed stack was driven end-to-end in a
|
| 74 |
+
headed browser, one upload per case, with per-case screenshots and recorded run ids:
|
| 75 |
+
|
| 76 |
+
| Property | Result |
|
| 77 |
+
|---|---|
|
| 78 |
+
| Independent full passes | **3** |
|
| 79 |
+
| Cases per pass | 8 (6 regression + 2 router-defect) |
|
| 80 |
+
| Passes at 8/8 | **3 of 3** |
|
| 81 |
+
| Live runs executed | **24** |
|
| 82 |
+
| Correct dispatches | **24** |
|
| 83 |
+
| Mock-node contamination | **0** on every live run |
|
| 84 |
+
| Trace fill | **94.4444 %** on every live run |
|
| 85 |
+
| Frontend regression suite | **106 passed** (`tests/unit/test_frontend_live_wiring.py`) |
|
| 86 |
+
|
| 87 |
+
Every pass produced **fresh run identifiers** — no run id is shared between passes. This is
|
| 88 |
+
behavioural evidence that the pipeline runs and routes correctly; it is **not** an accuracy claim.
|
| 89 |
+
|
| 90 |
+
## 5. How to reproduce
|
| 91 |
+
|
| 92 |
+
```bash
|
| 93 |
+
# verify every README/benchmark number against its artifact
|
| 94 |
+
python release/tools/verify_readme_metrics.py
|
| 95 |
+
```
|
| 96 |
+
|
| 97 |
+
The script resolves nested artifact keys (including keys that themselves contain dots, such as the
|
| 98 |
+
`recall` dict keyed `"0.10"/"0.25"/"0.50"`) and compares each value at the precision printed in the
|
| 99 |
+
README. It exits non-zero if any claim fails. See [`REPRODUCIBILITY.md`](REPRODUCIBILITY.md) for the
|
| 100 |
+
full reproduction guide.
|