Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Download docs/EVALUATION.md from thundercode/SatQuery: direct link, hf CLI and curl.
- Browser
- Download file 7.19 kB
-
https://huggingface.co/thundercode/SatQuery/resolve/00a146ce419c6c7c109650e26cf45353fffb594e/docs/EVALUATION.md
- Command line
-
hf download hf://thundercode/SatQuery@00a146ce419c6c7c109650e26cf45353fffb594e/docs/EVALUATION.md
-
curl -L -o EVALUATION.md https://huggingface.co/thundercode/SatQuery/resolve/00a146ce419c6c7c109650e26cf45353fffb594e/docs/EVALUATION.md
Evaluation
This document describes how each number in BENCHMARKS.md was produced, and
enforces the project's evaluation-honesty rules. It is deliberately conservative: a metric that was
not measured is tagged NOT RUN, and a metric that got worse is shown getting worse.
Status tags: VERIFIED · MEASURED · NOT RUN · OPEN · REJECTED.
1. Evaluation-honesty rules (enforced by convention and by tooling)
- Evidence before claims. Every reported number has an artifact path. A number with no artifact does not appear in the docs.
- Two protocols are never collapsed. Grounding is reported under canonical and matched6.
- Two test sets are never collapsed. Change-VQA is reported on
testandtest2. - accuracy never travels without macro-F1 for imbalanced multi-class heads.
- Validation is not test. The router figure is labelled validation, ungated, n = 86.
- A negative result stays negative. Calibration ECE worsened; it is shown worsening.
- USABLE ≠ ACCEPTED. The VLM adapter's metrics are real; its status is acceptance-rejected.
- No composite/vanity score. There is no single headline accuracy for the system, and none is invented by averaging the per-task numbers.
2. Per-task evaluation protocols
2.1 Change detection — VERIFIED
- Split: LEVIR-CD-256 test, n = 2,048 (immutable public split).
- Threshold: 0.50, frozen in config.
- Metrics: pooled IoU / macro IoU / pooled F1, plus the full confusion counts so any metric can be recomputed.
- Why both pooled and macro: the test split is only ≈ 5 % changed pixels (mean change fraction 0.0509). Pooled IoU (0.8122) and macro IoU (0.8457) answer different questions about that imbalance.
- Artifact:
artifacts/change/eval_test/eval_result.json.
2.2 Grounding — MEASURED, two protocols × two decode variants
- Split: VRSBench, n = 16,159.
- Box convention: VRSBench 0–100 → project 0–1, via declared
benchmark_box_scale: 100.0. - Protocols: canonical and matched6 — both reported.
- Decode variants: head threshold (the shipped decode), head argmax, and zero-shot matched (baseline).
- Resolution: frozen at 224; 448 rejected by a pre-registered paired test.
| Protocol | decode | mean best IoU | recall@0.5 |
|---|---|---|---|
| canonical | head threshold | 0.2838 | 0.2198 |
| canonical | head argmax | 0.1215 | — |
| canonical | zero-shot matched | 0.0972 | — |
| matched6 | head threshold | 0.2566 | 0.1938 |
Artifacts: …/eval_result_canonical.json, …/eval_result_matched6.json.
2.3 Optical-SAR fusion — MEASURED, ruling OPEN
- Split: held-out test, n = 4,000.
- Label space: 19 CLC classes; 14 present, 5 absent in the scored split.
- macro-F1 denominator: all 19 classes (absent classes contribute 0.0) — recorded explicitly.
- Pre-registration: the metric is a pre-registered protocol (
pre_registered_11.5);is_deciding_statistic: False. - Artifact:
artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json.
2.4 Change-VQA — MEASURED, ruling OPEN
- Splits:
test(n = 39,686) andtest2. - Selection: epoch 8, chosen on val answer accuracy 0.700018.
- Artifact:
artifacts/change_vqa/run/PROMOTION.json.
2.5 VLM adapter — MEASURED, ACCEPTANCE-REJECTED
- Split: frozen 1,000-question subset.
- Metrics: exact_match 0.963, F1 0.96432 (+49.5 pp over unadapted baseline).
- Decision: the artifact is not accepted for promotion; the record of why is stored in
artifacts/vlm/phase6_closure.jsonunderwhy_acceptance_rejected. - Artifact:
artifacts/vlm/phase6_closure.json(status: CLOSED).
2.6 Router — MEASURED, TEST NOT RUN
- Split scored: validation, n = 86, corpus-limited.
- Metric: overall ungated accuracy 0.965116.
- Not done: the test split was never scored. The val number is indicative only.
2.7 Calibration — MEASURED, not an improvement
- Fit split: val, n = 16,441. Temperature T = 0.9772732.
- Result: ECE 0.013755 → 0.014929 (
ece_improvement = −0.001174). - Retention rationale: part of the frozen config, not because it helped.
- Artifact:
artifacts/calibration_v001.json.
3. Behavioural evaluation (live validation)
Accuracy and behaviour are evaluated separately. The deployed stack was driven in a headed browser, one upload per case, with per-case screenshots and recorded run ids:
| Property | Result |
|---|---|
| Independent full passes | 3 |
| Cases per pass | 8 (6 regression + 2 router-defect) |
| Passes at 8/8 | 3 of 3 |
| Live runs | 24 |
| Correct dispatches | 24 |
| Mock-node contamination | 0 |
| Trace fill | 94.4444 % |
| Frontend regression suite | 106 passed |
3.1 Harness integrity — a bug that was caught
An earlier harness revision typed queries with synthetic CDP key events, which Chrome silently drops when the window lacks OS focus. The harness therefore dispatched the page's default query and still recorded a "result" — a false pass.
The current harness asserts form state before dispatch: that the query box really holds the
intended query, that #obsTail reads ready, and that both frames are attached for pair tasks. The
earlier 8/8 run was independently checked and confirmed not to have been infected (its answers
were query-specific and the query text was embedded in the answers). This failure mode is recorded
here because it is exactly the kind of silent false-positive an evaluation harness must not have.
4. Test suite
| Suite | Result | Notes |
|---|---|---|
tests/unit/test_frontend_live_wiring.py |
106 passed | re-run this session |
| Doc/frontend suite (5 files) | 183 passed | |
Full tests/unit |
5–6 failures, all environmental/ordering | 4× test_safe_delete_shim (sandbox delete guard), 1 ordering flake (passes in isolation), 1 stale adapter test (CROMA now shipped) |
| Affected files re-run together | 137 passed | confirms the failures are not real regressions |
The environmental failures are not hidden. They are attributable to the sandbox delete guard and test ordering, not to the code under test.
5. What is NOT evaluated
| Item | State |
|---|---|
| System-level end-to-end accuracy | NOT RUN — none exists |
| Router test split | NOT RUN |
| Benchmark adapters | NOT RUN |
| End-to-end latency benchmark | NOT RUN |
| Cross-dataset generalisation | NOT RUN |
| Human evaluation | NOT RUN |
| Adversarial / robustness evaluation | NOT RUN |
6. Reproducing the metric check
python release/tools/verify_readme_metrics.py
This walks every claim in ../README.md and BENCHMARKS.md to its
source artifact and compares values at the printed precision. It prints ALL CLAIMS VERIFIED (exit 0)
only when all 20 numeric claims match and the status assertions hold. Committed output:
release/tools/readme_metrics_report.txt.