SatQuery / docs /EVALUATION.md
thundercode's picture
release: add docs/EVALUATION.md
c2283ee verified
|
Raw History Blame
7.19 kB
# Evaluation
This document describes **how** each number in [`BENCHMARKS.md`](BENCHMARKS.md) was produced, and
enforces the project's evaluation-honesty rules. It is deliberately conservative: a metric that was
not measured is tagged `NOT RUN`, and a metric that got worse is shown getting worse.
**Status tags:** `VERIFIED` Β· `MEASURED` Β· `NOT RUN` Β· `OPEN` Β· `REJECTED`.
---
## 1. Evaluation-honesty rules (enforced by convention and by tooling)
1. **Evidence before claims.** Every reported number has an artifact path. A number with no artifact
does not appear in the docs.
2. **Two protocols are never collapsed.** Grounding is reported under canonical **and** matched6.
3. **Two test sets are never collapsed.** Change-VQA is reported on `test` **and** `test2`.
4. **accuracy never travels without macro-F1** for imbalanced multi-class heads.
5. **Validation is not test.** The router figure is labelled validation, ungated, n = 86.
6. **A negative result stays negative.** Calibration ECE worsened; it is shown worsening.
7. **USABLE β‰  ACCEPTED.** The VLM adapter's metrics are real; its status is acceptance-rejected.
8. **No composite/vanity score.** There is no single headline accuracy for the system, and none is
invented by averaging the per-task numbers.
## 2. Per-task evaluation protocols
### 2.1 Change detection β€” `VERIFIED`
- **Split:** LEVIR-CD-256 test, n = 2,048 (immutable public split).
- **Threshold:** 0.50, frozen in config.
- **Metrics:** pooled IoU / macro IoU / pooled F1, plus the full confusion counts so any metric can
be recomputed.
- **Why both pooled and macro:** the test split is only β‰ˆ 5 % changed pixels
(mean change fraction 0.0509). Pooled IoU (0.8122) and macro IoU (0.8457) answer different
questions about that imbalance.
- **Artifact:** `artifacts/change/eval_test/eval_result.json`.
### 2.2 Grounding β€” `MEASURED`, two protocols Γ— two decode variants
- **Split:** VRSBench, n = 16,159.
- **Box convention:** VRSBench 0–100 β†’ project 0–1, via declared `benchmark_box_scale: 100.0`.
- **Protocols:** *canonical* and *matched6* β€” both reported.
- **Decode variants:** *head threshold* (the shipped decode), *head argmax*, and *zero-shot
matched* (baseline).
- **Resolution:** frozen at 224; 448 rejected by a pre-registered paired test.
| Protocol | decode | mean best IoU | recall@0.5 |
|---|---|---|---|
| canonical | head threshold | 0.2838 | 0.2198 |
| canonical | head argmax | 0.1215 | β€” |
| canonical | zero-shot matched | 0.0972 | β€” |
| matched6 | head threshold | 0.2566 | 0.1938 |
**Artifacts:** `…/eval_result_canonical.json`, `…/eval_result_matched6.json`.
### 2.3 Optical-SAR fusion β€” `MEASURED`, ruling `OPEN`
- **Split:** held-out test, n = 4,000.
- **Label space:** 19 CLC classes; 14 present, 5 absent in the scored split.
- **macro-F1 denominator:** all 19 classes (absent classes contribute 0.0) β€” recorded explicitly.
- **Pre-registration:** the metric is a pre-registered protocol (`pre_registered_11.5`);
`is_deciding_statistic: False`.
- **Artifact:** `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`.
### 2.4 Change-VQA β€” `MEASURED`, ruling `OPEN`
- **Splits:** `test` (n = 39,686) and `test2`.
- **Selection:** epoch 8, chosen on val answer accuracy 0.700018.
- **Artifact:** `artifacts/change_vqa/run/PROMOTION.json`.
### 2.5 VLM adapter β€” `MEASURED`, `ACCEPTANCE-REJECTED`
- **Split:** frozen 1,000-question subset.
- **Metrics:** exact_match 0.963, F1 0.96432 (+49.5 pp over unadapted baseline).
- **Decision:** the artifact is **not accepted** for promotion; the record of *why* is stored in
`artifacts/vlm/phase6_closure.json` under `why_acceptance_rejected`.
- **Artifact:** `artifacts/vlm/phase6_closure.json` (`status: CLOSED`).
### 2.6 Router β€” `MEASURED`, `TEST NOT RUN`
- **Split scored:** validation, n = 86, corpus-limited.
- **Metric:** overall **ungated** accuracy 0.965116.
- **Not done:** the test split was never scored. The val number is indicative only.
### 2.7 Calibration β€” `MEASURED`, not an improvement
- **Fit split:** val, n = 16,441. Temperature **T = 0.9772732**.
- **Result:** ECE 0.013755 β†’ **0.014929** (`ece_improvement = βˆ’0.001174`).
- **Retention rationale:** part of the frozen config, **not** because it helped.
- **Artifact:** `artifacts/calibration_v001.json`.
## 3. Behavioural evaluation (live validation)
Accuracy and behaviour are evaluated separately. The deployed stack was driven in a headed browser,
one upload per case, with per-case screenshots and recorded run ids:
| Property | Result |
|---|---|
| Independent full passes | **3** |
| Cases per pass | 8 (6 regression + 2 router-defect) |
| Passes at 8/8 | **3 of 3** |
| Live runs | **24** |
| Correct dispatches | **24** |
| Mock-node contamination | **0** |
| Trace fill | **94.4444 %** |
| Frontend regression suite | **106 passed** |
### 3.1 Harness integrity β€” a bug that was caught
An earlier harness revision typed queries with **synthetic CDP key events**, which Chrome **silently
drops when the window lacks OS focus**. The harness therefore dispatched the page's *default* query
and still recorded a "result" β€” a false pass.
The current harness **asserts form state before dispatch**: that the query box really holds the
intended query, that `#obsTail` reads `ready`, and that both frames are attached for pair tasks. The
earlier 8/8 run was independently checked and confirmed **not** to have been infected (its answers
were query-specific and the query text was embedded in the answers). This failure mode is recorded
here because it is exactly the kind of silent false-positive an evaluation harness must not have.
## 4. Test suite
| Suite | Result | Notes |
|---|---|---|
| `tests/unit/test_frontend_live_wiring.py` | **106 passed** | re-run this session |
| Doc/frontend suite (5 files) | **183 passed** | |
| Full `tests/unit` | 5–6 failures, **all environmental/ordering** | 4Γ— `test_safe_delete_shim` (sandbox delete guard), 1 ordering flake (passes in isolation), 1 stale adapter test (CROMA now shipped) |
| Affected files re-run together | **137 passed** | confirms the failures are not real regressions |
**The environmental failures are not hidden.** They are attributable to the sandbox delete guard and
test ordering, not to the code under test.
## 5. What is NOT evaluated
| Item | State |
|---|---|
| System-level end-to-end accuracy | **NOT RUN β€” none exists** |
| Router test split | **NOT RUN** |
| Benchmark adapters | **NOT RUN** |
| End-to-end latency benchmark | **NOT RUN** |
| Cross-dataset generalisation | **NOT RUN** |
| Human evaluation | **NOT RUN** |
| Adversarial / robustness evaluation | **NOT RUN** |
## 6. Reproducing the metric check
```bash
python release/tools/verify_readme_metrics.py
```
This walks every claim in [`../README.md`](../README.md) and [`BENCHMARKS.md`](BENCHMARKS.md) to its
source artifact and compares values at the printed precision. It prints `ALL CLAIMS VERIFIED` (exit 0)
only when **all 20** numeric claims match and the status assertions hold. Committed output:
`release/tools/readme_metrics_report.txt`.