thundercode commited on
Commit
ad9cc6a
·
verified ·
1 Parent(s): 724b09e

release: add docs/BENCHMARKS.md

Browse files
Files changed (1) hide show
  1. docs/BENCHMARKS.md +100 -0
docs/BENCHMARKS.md ADDED
@@ -0,0 +1,100 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Benchmarks
2
+
3
+ **Every number here is artifact-backed.** Where a metric exists, its source file is named. Where a
4
+ benchmark does **not** exist, that is stated explicitly and tagged `NOT RUN` — an absent number is
5
+ never silently omitted or replaced with an estimate.
6
+
7
+ **Status tags:** `VERIFIED` · `MEASURED` · `NOT RUN` · `OPEN` · `REJECTED`.
8
+
9
+ ---
10
+
11
+ ## 1. Headline table
12
+
13
+ | Capability | Metric | Value | Split / protocol | Status |
14
+ |---|---|---|---|---|
15
+ | Change detection | pooled IoU | **0.8122** | LEVIR-CD-256 test, n = 2048, thr 0.50 | **VERIFIED** |
16
+ | Change detection | macro IoU | **0.8457** | same | **VERIFIED** |
17
+ | Change detection | pooled F1 | **0.8964** | same | **VERIFIED** |
18
+ | Grounding | mean best IoU | **0.2838** / **0.2566** | VRSBench n = 16159, canonical / matched6 | MEASURED (2 protocols) |
19
+ | Grounding | recall@0.5 | **0.2198** / **0.1938** | same | MEASURED (2 protocols) |
20
+ | Grounding | head-argmax IoU | **0.1215** | canonical | MEASURED |
21
+ | Grounding | zero-shot baseline IoU | **0.0972** | canonical | MEASURED (baseline) |
22
+ | Optical-SAR fusion | accuracy | **0.931** | held-out test n = 4000, 19 classes | MEASURED, ruling **OPEN** |
23
+ | Optical-SAR fusion | macro F1 | **0.434161** | same | MEASURED, ruling **OPEN** |
24
+ | Change-VQA | accuracy | **0.697626** | test n = 39686 | MEASURED, ruling **OPEN** |
25
+ | Change-VQA | macro F1 | **0.378373** | same | MEASURED, ruling **OPEN** |
26
+ | Change-VQA | accuracy (2nd set) | **0.651469** | test2 | MEASURED, ruling **OPEN** |
27
+ | Change-VQA | macro F1 (2nd set) | **0.372309** | test2 | MEASURED, ruling **OPEN** |
28
+ | VLM (adapted) | exact_match | **0.963** | frozen 1000-question subset | MEASURED, **ACCEPTANCE-REJECTED** |
29
+ | VLM (adapted) | F1 | **0.96432** | same | MEASURED, **ACCEPTANCE-REJECTED** |
30
+ | Router | overall **ungated** accuracy | **0.965116** | val n = 86, corpus-limited | MEASURED — **TEST NOT RUN** |
31
+ | Calibration | ECE before / after | **0.013755 → 0.014929** | val n = 16441, T = 0.9773 | MEASURED — **worse** |
32
+
33
+ ### Source artifacts
34
+
35
+ | Metric family | Artifact |
36
+ |---|---|
37
+ | change | `artifacts/change/eval_test/eval_result.json` |
38
+ | grounding | `artifacts/grounding/remoteclip_grounding_v001/eval_result_canonical.json`, `…_matched6.json` |
39
+ | optical-SAR | `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json` |
40
+ | change-VQA | `artifacts/change_vqa/run/PROMOTION.json` |
41
+ | router | `artifacts/router/threshold_sweep_val.json` |
42
+ | calibration | `artifacts/calibration_v001.json` |
43
+ | VLM | `artifacts/vlm/phase6_closure.json` |
44
+
45
+ > All 20 values in the table above are checked against these files by
46
+ > [`../tools/verify_readme_metrics.py`](../tools/verify_readme_metrics.py). Its output
47
+ > (`ALL CLAIMS VERIFIED`) is committed as `../tools/readme_metrics_report.txt`.
48
+
49
+ ## 2. Rules this table follows
50
+
51
+ 1. **Two protocols are never collapsed.** Grounding is reported under *both* the canonical and
52
+ matched6 protocols. Quoting 0.2838 alone would be selective.
53
+ 2. **Two test sets are never collapsed.** Change-VQA is reported on `test` **and** `test2`.
54
+ 3. **accuracy never travels without macro-F1.** For imbalanced multi-class heads (optical-SAR,
55
+ change-VQA) the macro-F1 is reported alongside accuracy, always.
56
+ 4. **Validation is not test.** The router number is labelled "overall **ungated** accuracy", val,
57
+ n = 86. It is not a test result.
58
+ 5. **A negative result stays negative.** Calibration ECE worsened and is shown worsening.
59
+ 6. **USABLE ≠ ACCEPTED.** The VLM metrics are real; the artifact is nevertheless acceptance-rejected.
60
+
61
+ ## 3. What is NOT benchmarked
62
+
63
+ | Benchmark | Status | Note |
64
+ |---|---|---|
65
+ | **System-level end-to-end accuracy** | **NOT RUN — none exists** | There is no measured end-to-end benchmark of the full router→specialist→envelope pipeline. No such number is claimed anywhere. |
66
+ | **Router test split** | **NOT RUN** | Only the validation split (n = 86) was scored. |
67
+ | **Benchmark adapters** | **NOT RUN** | Adapter-based benchmark runs were not executed. |
68
+ | **Efficiency / latency benchmark** | not systematically measured | Per-specialist latency is recorded incidentally in artifacts (e.g. grounding `latency_ms_per_image` 2.205 ms for the head), but there is no end-to-end latency benchmark. |
69
+ | **Cross-dataset generalisation** | **NOT RUN** | Each specialist is evaluated only on its own training-family test split. |
70
+
71
+ ## 4. Live validation (behavioural, not accuracy)
72
+
73
+ Accuracy is separate from **behavioural** validation. The deployed stack was driven end-to-end in a
74
+ headed browser, one upload per case, with per-case screenshots and recorded run ids:
75
+
76
+ | Property | Result |
77
+ |---|---|
78
+ | Independent full passes | **3** |
79
+ | Cases per pass | 8 (6 regression + 2 router-defect) |
80
+ | Passes at 8/8 | **3 of 3** |
81
+ | Live runs executed | **24** |
82
+ | Correct dispatches | **24** |
83
+ | Mock-node contamination | **0** on every live run |
84
+ | Trace fill | **94.4444 %** on every live run |
85
+ | Frontend regression suite | **106 passed** (`tests/unit/test_frontend_live_wiring.py`) |
86
+
87
+ Every pass produced **fresh run identifiers** — no run id is shared between passes. This is
88
+ behavioural evidence that the pipeline runs and routes correctly; it is **not** an accuracy claim.
89
+
90
+ ## 5. How to reproduce
91
+
92
+ ```bash
93
+ # verify every README/benchmark number against its artifact
94
+ python release/tools/verify_readme_metrics.py
95
+ ```
96
+
97
+ The script resolves nested artifact keys (including keys that themselves contain dots, such as the
98
+ `recall` dict keyed `"0.10"/"0.25"/"0.50"`) and compares each value at the precision printed in the
99
+ README. It exits non-zero if any claim fails. See [`REPRODUCIBILITY.md`](REPRODUCIBILITY.md) for the
100
+ full reproduction guide.