constructelligence commited on
Commit
e43c8ca
·
verified ·
1 Parent(s): 2a65210

Add evidence caveat, calibration/abstention (91% precision at 50% coverage) and improvement roadmap

Browse files
Files changed (2) hide show
  1. IMPROVEMENT-PLAN.md +126 -0
  2. README.md +34 -0
IMPROVEMENT-PLAN.md ADDED
@@ -0,0 +1,126 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # MasterFormat classifier — improvement plan
2
+
3
+ **Date:** 2026-10-05 · **Owner:** Constructelligence
4
+ **Published:** [`constructelligence/masterformat-classifier`](https://huggingface.co/constructelligence/masterformat-classifier)
5
+ **Current:** `mf-0.2` transformer + `ensemble/` (word+char TF-IDF blend).
6
+
7
+ | metric | published value | 95 % CI | n |
8
+ |---|---:|---|---:|
9
+ | line-item top-1 (hand-labelled estimates) | 0.702 | **[0.633, 0.762]** | 194 (191 in-taxonomy) |
10
+ | whole-section top-1 | 0.620 | **[0.540, 0.694]** | 153 |
11
+ | whole-section division | 0.804 | ±~6 pts | 153 |
12
+ | UFGS val top-1 | 0.588 | [0.579, 0.596] | 11,598 |
13
+ | manual-chunk top-1 | 0.330 | [0.319, 0.341] | 6,890 |
14
+
15
+ ---
16
+
17
+ ## 1. The binding constraint is evidence, not modelling
18
+
19
+ The two metrics that matter for the product — **hand-labelled line items** and **whole real spec sections** —
20
+ come from tiny samples. Their confidence intervals are **±6–8 points**, which means:
21
+
22
+ - the recent `0.686 → 0.702` ensemble change is **statistically indistinguishable from noise**;
23
+ - any model change smaller than ~10 points on line items cannot be validated on this eval;
24
+ - the tight, large-sample metric (`val`, ±0.9 pts) is **in-domain UFGS text**, not the target distribution —
25
+ a model can improve `val` a lot (see `mf-0.3`: 0.466 → 0.556) while getting *worse* on real manuals.
26
+
27
+ **Consequence:** the next gains are gated on **a bigger, real-world evaluation set**, not on architecture.
28
+ Chasing sub-CI improvements (more TF-IDF features, more seeds) is a trap we have already fallen into once.
29
+
30
+ ## 2. Where the headroom actually is
31
+
32
+ 1. **Line items (estimating).** TF-IDF dominates; the transformer adds little. True ceiling unknown because
33
+ the eval is underpowered. Needs *real* labelled items, not synthetic augmentation (`mf-0.3` proved
34
+ synthetic terse augmentation does not transfer: line items stayed 0.597 and manuals regressed).
35
+ 2. **Section-level *top-1*** (0.62) is the weak spot, while *division* (0.81) is strong → most residual error
36
+ is **confusion within a division** (e.g. `03 30 00` vs `03 50 00`). A coarse-to-fine or contrastive
37
+ approach targets exactly this.
38
+ 3. **Out-of-taxonomy** sections only score at division level; there is no `other`/hierarchical fallback.
39
+ 4. **No level-3 / full 8-digit section codes** — real spec books use them.
40
+
41
+ ## 3. Definition of done (gates for any new model)
42
+
43
+ | Gate | Target | How to verify |
44
+ |---|---|---|
45
+ | Powered eval | line items ≥ 1,000; sections ≥ 400 | Wilson CI ≤ ±3 pts |
46
+ | Measurable lift | ≥ +3 pts line-item top-1 vs published, 95 % CI non-overlapping | paired bootstrap / McNemar |
47
+ | No regression | manuals + section division within 1 pt | fixed eval sets |
48
+ | Calibrated | ECE ≤ 0.05 after temperature scaling | reliability curve |
49
+ | Useful abstention | ≥ 90 % precision at ≥ 50 % coverage on line items | coverage–precision curve |
50
+ | Reproducible | one `cloud/run_kaggle.py --run <name>` + seed | identical metrics ± noise |
51
+
52
+ ## 4. Plan
53
+
54
+ ### P0 — Make the measurement trustworthy, then feed it real data
55
+
56
+ **P0.1 — Enlarge the held-out eval (do first).**
57
+ *Why:* everything downstream is unverifiable until this exists. *How:* assemble ≥1,000 real estimate line
58
+ items and ≥400 real spec sections with true MasterFormat labels from public sources (§5); store as
59
+ `eval/line_items_v2.tsv`, `eval/manuals_v2.jsonl`; report Wilson CIs + paired bootstrap for every comparison.
60
+ *Done when:* CIs ≤ ±3 pts and the published vs candidate comparison reports a p-value. *Effort:* M.
61
+
62
+ **P0.2 — Acquire real line-item training labels (not just eval).**
63
+ *Why:* the transformer's line-item ceiling is a data problem, not a capacity problem. *How:* mine public bid
64
+ tabulations / agency item catalogs; map agency item codes to MasterFormat where a crosswalk exists; treat the
65
+ rest as weak/self-supervised signal. Keep a strict train/eval split by project. *Done when:* ≥10k labelled
66
+ items, none from the eval projects. *Effort:* L.
67
+
68
+ **P0.3 — Domain-adaptive pretraining corpus.**
69
+ *Why:* the encoder only ever saw UFGS. *How:* collect public construction prose (specs, bid tabs, RFIs, submittal
70
+ logs) beyond UFGS — VA design manuals, state DOT standard specs, Corps/UFC, NASA. Continued-MLM the encoder,
71
+ then fine-tune. *Done when:* mf-0.5 beats the current model on the *enlarged* eval with non-overlapping CI.
72
+ *Effort:* L (compute: 1–2 GPU-days).
73
+
74
+ **P0.4 — Label space: level-3 sections + explicit out-of-taxonomy handling.**
75
+ *Why:* real specs use 8-digit codes and many sections outside the 171 groups. *How:* extend the taxonomy to
76
+ full sections; add an `other/<division>` fallback and score it honestly; revisit `group_of()`. *Effort:* M.
77
+
78
+ ### P1 — Modelling (only once P0.1 is in place)
79
+
80
+ - **P1.1 Distillation** from the TF-IDF/ensemble teacher using **out-of-fold** soft targets (never on-data
81
+ teacher labels) into the transformer. Replaces the current brittle log-prob blend with one model.
82
+ - **P1.2 Coarse-to-fine / hierarchical head** (division → group), directly attacking within-division confusion
83
+ and giving graceful out-of-taxonomy fallback.
84
+ - **P1.3 Retrieval + rerank:** kNN over labelled embeddings for the top-k, plus a cross-encoder to reorder
85
+ them. Often the cheapest way to absorb new labelled data as it arrives.
86
+ - **P1.4 Calibration + abstention:** temperature scaling on `val`; per-class thresholds; a `REVIEW` queue.
87
+ Ship this regardless of accuracy — it converts a probabilistic classifier into a usable one.
88
+ - **P1.5 Multi-task:** shared encoder with group + division + section heads; consistency regularisation between
89
+ them.
90
+
91
+ ### P2 — Serving / ops
92
+
93
+ - **P2.1 Routing policy:** transformer for whole sections / long text, ensemble for terse items (already the
94
+ evidence-backed split).
95
+ - **P2.2 Quantization-aware / int8-calibrated ONNX** (the int8 export currently drifts; add QAT or
96
+ per-channel calibration and a browser benchmark).
97
+ - **P2.3 Golden-file CI** on the enlarged eval; nightly regression that fails on a gate breach.
98
+ - **P2.4 Metrics dashboard** with coverage/precision so abstention is tuned in production.
99
+
100
+ ## 5. Data sourcing shortlist (public unless noted)
101
+
102
+ | Source | Use | Note |
103
+ |---|---|---|
104
+ | UFGS + UFC | train (in use) | US federal, public domain |
105
+ | State DOT standard specs (Caltrans, TxDOT, WSDOT, MnDOT…) | pretrain / train | mostly public |
106
+ | VA design manuals & specifications | pretrain / train | US federal, public |
107
+ | NASA / DoD / Corps specs | pretrain | public |
108
+ | Public **bid tabulations** (agency open-data portals; city/state capital projects) | line-item text + weak labels | item number + description; codes vary by agency |
109
+ | SAM.gov / FPDS contract line items | weak labels | huge, noisy |
110
+ | Public BOQ / cost datasets (e.g. Kaggle construction cost) | line items | check licence |
111
+ | RSMeans / Gordian cost codes | labels | **proprietary** — licence required |
112
+
113
+ ## 6. Sequenced milestones
114
+
115
+ 1. **M1 (now):** publish the evaluation caveat + `cloud/stats.py` (CIs, abstention). Freeze `mf-0.2`+ensemble
116
+ as the reference. ✅ started in this change.
117
+ 2. **M2:** enlarged eval (P0.1) from public bid tabs + manuals; re-score the reference.
118
+ 3. **M3:** domain-adaptive pretrain + retrain (`mf-0.5`) on P0.2/P0.3 data; validate with CIs (P0.3).
119
+ 4. **M4:** distillation + hierarchical head (`mf-0.6`); pick the best single model (P1.1/P1.2).
120
+ 5. **M5:** calibration/abstention + routing in serving; publish with a coverage–precision card (P1.4, P2.1).
121
+
122
+ ## 7. Explicitly not doing
123
+
124
+ - More synthetic line-item augmentation (tested: `mf-0.3` — no line-item transfer, manual regression).
125
+ - More TF-IDF feature engineering (already saturated; gains below the CI).
126
+ - Reporting sub-CI improvements as wins on the model card.
README.md CHANGED
@@ -155,6 +155,11 @@ sets are **fixed**, so those columns are directly comparable for every scorer. `
155
  commercial project manuals; out-of-taxonomy gold labels (e.g. `22 40 00`) count only toward the division
156
  score, so top-1 is over in-taxonomy items only.
157
 
 
 
 
 
 
158
  **Reading it:** `mf-0.2` is the best scorer overall at section-level *division* accuracy (0.810). On short
159
  line items a lexical TF-IDF model is stronger (0.696 alone), so the shipped **transformer + TF-IDF ensemble**
160
  reaches **0.702** and is the production candidate; see *Ensemble*.
@@ -194,6 +199,34 @@ the weights on the 194-item test itself and does not generalise. The transformer
194
  section *division* accuracy (0.810). A reasonable production split is the transformer for whole sections, the
195
  ensemble for terse line items.
196
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
197
  ## Intended use
198
 
199
  - **Use it for:** tagging construction text with a candidate MasterFormat group (top-3 shown), an
@@ -268,6 +301,7 @@ manual-chunk and whole-section accuracy. Constructelligence's proprietary models
268
  - `onnx/` — ONNX fp32 and int8 exports.
269
  - `ensemble/` — `tfidf.joblib` (word + char n-grams, ~112 MB), `blend.json`, `predict_ensemble.py`: the
270
  transformer + TF-IDF line-item ensemble.
 
271
  - `CITATION.cff`, `requirements.txt`.
272
 
273
  ## Citation
 
155
  commercial project manuals; out-of-taxonomy gold labels (e.g. `22 40 00`) count only toward the division
156
  score, so top-1 is over in-taxonomy items only.
157
 
158
+ **Evidence caveat.** The line-item and section sets are small. At n = 191 the 95 % Wilson interval on the
159
+ line-item top-1 is **[0.633, 0.762] (±6.4 pts)**, and whole-section is ±~8 pts — so differences of a few
160
+ points between models here are **noise**. Treat the `val` and `manual-chunk` rows (±1 pt) as the measurable
161
+ ones, and the line-item/section rows as indicative. Enlarging these eval sets is the top item on the roadmap.
162
+
163
  **Reading it:** `mf-0.2` is the best scorer overall at section-level *division* accuracy (0.810). On short
164
  line items a lexical TF-IDF model is stronger (0.696 alone), so the shipped **transformer + TF-IDF ensemble**
165
  reaches **0.702** and is the production candidate; see *Ensemble*.
 
199
  section *division* accuracy (0.810). A reasonable production split is the transformer for whole sections, the
200
  ensemble for terse line items.
201
 
202
+ ## Calibration and abstention
203
+
204
+ The ensemble is already close to calibrated; a single **temperature** `T = 0.90` (fit on `val`) lowers the
205
+ expected calibration error on `val` from **0.060 to 0.022**. More useful in production: send the
206
+ lowest-confidence predictions to a human instead of filing them.
207
+
208
+ Coverage → precision on the 194 line items (predictions sorted by confidence):
209
+
210
+ | auto-filed (coverage) | 100 % | 90 % | 80 % | 70 % | 50 % | 30 % |
211
+ |---|---:|---:|---:|---:|---:|---:|
212
+ | precision | 0.69 | 0.75 | 0.81 | **0.85** | **0.91** | 0.97 |
213
+
214
+ Routing the least-confident **30 %** to review lifts precision to **91 %**; routing 50 % gives **97 %**. So
215
+ the model is usable today as an **assist that flags its own uncertainty**, even though it is not accurate
216
+ enough to file unattended. Reproduce with `python cloud/stats.py`.
217
+
218
+ ## Roadmap
219
+
220
+ Further gains are **evidence-bound**, not architecture-bound (see the repo's `IMPROVEMENT-PLAN.md`):
221
+
222
+ 1. **Enlarge the eval** (≥1,000 line items, ≥400 sections) so changes are measurable at ±3 pts — currently
223
+ line-item changes below ~10 pts cannot be validated.
224
+ 2. **Real line-item training data** (public bid tabulations, agency item catalogs) — the transformer's
225
+ line-item ceiling is a data problem; synthetic augmentation was tested and did not transfer (`mf-0.3`).
226
+ 3. **Domain-adaptive pretraining** on construction prose beyond UFGS, then retrain.
227
+ 4. **Distillation + hierarchical (division → group) head** to beat the TF-IDF blend with a single model.
228
+ 5. **Level-3 / full-section codes** and an explicit out-of-taxonomy fallback.
229
+
230
  ## Intended use
231
 
232
  - **Use it for:** tagging construction text with a candidate MasterFormat group (top-3 shown), an
 
301
  - `onnx/` — ONNX fp32 and int8 exports.
302
  - `ensemble/` — `tfidf.joblib` (word + char n-grams, ~112 MB), `blend.json`, `predict_ensemble.py`: the
303
  transformer + TF-IDF line-item ensemble.
304
+ - `IMPROVEMENT-PLAN.md` — the prioritised roadmap (enlarged eval, data, domain-adaptive pretraining, distillation).
305
  - `CITATION.cff`, `requirements.txt`.
306
 
307
  ## Citation