Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download docs/EVALUATION.md from thundercode/SatQuery: direct link, hf CLI and curl.
- Browser
- Download file 7.19 kB
-
https://huggingface.co/thundercode/SatQuery/resolve/00a146ce419c6c7c109650e26cf45353fffb594e/docs/EVALUATION.md
- Command line
-
hf download hf://thundercode/SatQuery@00a146ce419c6c7c109650e26cf45353fffb594e/docs/EVALUATION.md
-
curl -L -o EVALUATION.md https://huggingface.co/thundercode/SatQuery/resolve/00a146ce419c6c7c109650e26cf45353fffb594e/docs/EVALUATION.md
7.19 kB
| # Evaluation | |
| This document describes **how** each number in [`BENCHMARKS.md`](BENCHMARKS.md) was produced, and | |
| enforces the project's evaluation-honesty rules. It is deliberately conservative: a metric that was | |
| not measured is tagged `NOT RUN`, and a metric that got worse is shown getting worse. | |
| **Status tags:** `VERIFIED` Β· `MEASURED` Β· `NOT RUN` Β· `OPEN` Β· `REJECTED`. | |
| --- | |
| ## 1. Evaluation-honesty rules (enforced by convention and by tooling) | |
| 1. **Evidence before claims.** Every reported number has an artifact path. A number with no artifact | |
| does not appear in the docs. | |
| 2. **Two protocols are never collapsed.** Grounding is reported under canonical **and** matched6. | |
| 3. **Two test sets are never collapsed.** Change-VQA is reported on `test` **and** `test2`. | |
| 4. **accuracy never travels without macro-F1** for imbalanced multi-class heads. | |
| 5. **Validation is not test.** The router figure is labelled validation, ungated, n = 86. | |
| 6. **A negative result stays negative.** Calibration ECE worsened; it is shown worsening. | |
| 7. **USABLE β ACCEPTED.** The VLM adapter's metrics are real; its status is acceptance-rejected. | |
| 8. **No composite/vanity score.** There is no single headline accuracy for the system, and none is | |
| invented by averaging the per-task numbers. | |
| ## 2. Per-task evaluation protocols | |
| ### 2.1 Change detection β `VERIFIED` | |
| - **Split:** LEVIR-CD-256 test, n = 2,048 (immutable public split). | |
| - **Threshold:** 0.50, frozen in config. | |
| - **Metrics:** pooled IoU / macro IoU / pooled F1, plus the full confusion counts so any metric can | |
| be recomputed. | |
| - **Why both pooled and macro:** the test split is only β 5 % changed pixels | |
| (mean change fraction 0.0509). Pooled IoU (0.8122) and macro IoU (0.8457) answer different | |
| questions about that imbalance. | |
| - **Artifact:** `artifacts/change/eval_test/eval_result.json`. | |
| ### 2.2 Grounding β `MEASURED`, two protocols Γ two decode variants | |
| - **Split:** VRSBench, n = 16,159. | |
| - **Box convention:** VRSBench 0β100 β project 0β1, via declared `benchmark_box_scale: 100.0`. | |
| - **Protocols:** *canonical* and *matched6* β both reported. | |
| - **Decode variants:** *head threshold* (the shipped decode), *head argmax*, and *zero-shot | |
| matched* (baseline). | |
| - **Resolution:** frozen at 224; 448 rejected by a pre-registered paired test. | |
| | Protocol | decode | mean best IoU | recall@0.5 | | |
| |---|---|---|---| | |
| | canonical | head threshold | 0.2838 | 0.2198 | | |
| | canonical | head argmax | 0.1215 | β | | |
| | canonical | zero-shot matched | 0.0972 | β | | |
| | matched6 | head threshold | 0.2566 | 0.1938 | | |
| **Artifacts:** `β¦/eval_result_canonical.json`, `β¦/eval_result_matched6.json`. | |
| ### 2.3 Optical-SAR fusion β `MEASURED`, ruling `OPEN` | |
| - **Split:** held-out test, n = 4,000. | |
| - **Label space:** 19 CLC classes; 14 present, 5 absent in the scored split. | |
| - **macro-F1 denominator:** all 19 classes (absent classes contribute 0.0) β recorded explicitly. | |
| - **Pre-registration:** the metric is a pre-registered protocol (`pre_registered_11.5`); | |
| `is_deciding_statistic: False`. | |
| - **Artifact:** `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`. | |
| ### 2.4 Change-VQA β `MEASURED`, ruling `OPEN` | |
| - **Splits:** `test` (n = 39,686) and `test2`. | |
| - **Selection:** epoch 8, chosen on val answer accuracy 0.700018. | |
| - **Artifact:** `artifacts/change_vqa/run/PROMOTION.json`. | |
| ### 2.5 VLM adapter β `MEASURED`, `ACCEPTANCE-REJECTED` | |
| - **Split:** frozen 1,000-question subset. | |
| - **Metrics:** exact_match 0.963, F1 0.96432 (+49.5 pp over unadapted baseline). | |
| - **Decision:** the artifact is **not accepted** for promotion; the record of *why* is stored in | |
| `artifacts/vlm/phase6_closure.json` under `why_acceptance_rejected`. | |
| - **Artifact:** `artifacts/vlm/phase6_closure.json` (`status: CLOSED`). | |
| ### 2.6 Router β `MEASURED`, `TEST NOT RUN` | |
| - **Split scored:** validation, n = 86, corpus-limited. | |
| - **Metric:** overall **ungated** accuracy 0.965116. | |
| - **Not done:** the test split was never scored. The val number is indicative only. | |
| ### 2.7 Calibration β `MEASURED`, not an improvement | |
| - **Fit split:** val, n = 16,441. Temperature **T = 0.9772732**. | |
| - **Result:** ECE 0.013755 β **0.014929** (`ece_improvement = β0.001174`). | |
| - **Retention rationale:** part of the frozen config, **not** because it helped. | |
| - **Artifact:** `artifacts/calibration_v001.json`. | |
| ## 3. Behavioural evaluation (live validation) | |
| Accuracy and behaviour are evaluated separately. The deployed stack was driven in a headed browser, | |
| one upload per case, with per-case screenshots and recorded run ids: | |
| | Property | Result | | |
| |---|---| | |
| | Independent full passes | **3** | | |
| | Cases per pass | 8 (6 regression + 2 router-defect) | | |
| | Passes at 8/8 | **3 of 3** | | |
| | Live runs | **24** | | |
| | Correct dispatches | **24** | | |
| | Mock-node contamination | **0** | | |
| | Trace fill | **94.4444 %** | | |
| | Frontend regression suite | **106 passed** | | |
| ### 3.1 Harness integrity β a bug that was caught | |
| An earlier harness revision typed queries with **synthetic CDP key events**, which Chrome **silently | |
| drops when the window lacks OS focus**. The harness therefore dispatched the page's *default* query | |
| and still recorded a "result" β a false pass. | |
| The current harness **asserts form state before dispatch**: that the query box really holds the | |
| intended query, that `#obsTail` reads `ready`, and that both frames are attached for pair tasks. The | |
| earlier 8/8 run was independently checked and confirmed **not** to have been infected (its answers | |
| were query-specific and the query text was embedded in the answers). This failure mode is recorded | |
| here because it is exactly the kind of silent false-positive an evaluation harness must not have. | |
| ## 4. Test suite | |
| | Suite | Result | Notes | | |
| |---|---|---| | |
| | `tests/unit/test_frontend_live_wiring.py` | **106 passed** | re-run this session | | |
| | Doc/frontend suite (5 files) | **183 passed** | | | |
| | Full `tests/unit` | 5β6 failures, **all environmental/ordering** | 4Γ `test_safe_delete_shim` (sandbox delete guard), 1 ordering flake (passes in isolation), 1 stale adapter test (CROMA now shipped) | | |
| | Affected files re-run together | **137 passed** | confirms the failures are not real regressions | | |
| **The environmental failures are not hidden.** They are attributable to the sandbox delete guard and | |
| test ordering, not to the code under test. | |
| ## 5. What is NOT evaluated | |
| | Item | State | | |
| |---|---| | |
| | System-level end-to-end accuracy | **NOT RUN β none exists** | | |
| | Router test split | **NOT RUN** | | |
| | Benchmark adapters | **NOT RUN** | | |
| | End-to-end latency benchmark | **NOT RUN** | | |
| | Cross-dataset generalisation | **NOT RUN** | | |
| | Human evaluation | **NOT RUN** | | |
| | Adversarial / robustness evaluation | **NOT RUN** | | |
| ## 6. Reproducing the metric check | |
| ```bash | |
| python release/tools/verify_readme_metrics.py | |
| ``` | |
| This walks every claim in [`../README.md`](../README.md) and [`BENCHMARKS.md`](BENCHMARKS.md) to its | |
| source artifact and compares values at the printed precision. It prints `ALL CLAIMS VERIFIED` (exit 0) | |
| only when **all 20** numeric claims match and the status assertions hold. Committed output: | |
| `release/tools/readme_metrics_report.txt`. | |