masterformat-classifier / IMPROVEMENT-PLAN.md
constructelligence's picture
Add evidence caveat, calibration/abstention (91% precision at 50% coverage) and improvement roadmap
e43c8ca verified
|
Raw History Blame Contribute Delete
7.79 kB

MasterFormat classifier — improvement plan

Date: 2026-10-05 · Owner: Constructelligence Published: constructelligence/masterformat-classifier Current: mf-0.2 transformer + ensemble/ (word+char TF-IDF blend).

metric published value 95 % CI n
line-item top-1 (hand-labelled estimates) 0.702 [0.633, 0.762] 194 (191 in-taxonomy)
whole-section top-1 0.620 [0.540, 0.694] 153
whole-section division 0.804 ±~6 pts 153
UFGS val top-1 0.588 [0.579, 0.596] 11,598
manual-chunk top-1 0.330 [0.319, 0.341] 6,890

1. The binding constraint is evidence, not modelling

The two metrics that matter for the product — hand-labelled line items and whole real spec sections — come from tiny samples. Their confidence intervals are ±6–8 points, which means:

  • the recent 0.686 → 0.702 ensemble change is statistically indistinguishable from noise;
  • any model change smaller than ~10 points on line items cannot be validated on this eval;
  • the tight, large-sample metric (val, ±0.9 pts) is in-domain UFGS text, not the target distribution — a model can improve val a lot (see mf-0.3: 0.466 → 0.556) while getting worse on real manuals.

Consequence: the next gains are gated on a bigger, real-world evaluation set, not on architecture. Chasing sub-CI improvements (more TF-IDF features, more seeds) is a trap we have already fallen into once.

2. Where the headroom actually is

  1. Line items (estimating). TF-IDF dominates; the transformer adds little. True ceiling unknown because the eval is underpowered. Needs real labelled items, not synthetic augmentation (mf-0.3 proved synthetic terse augmentation does not transfer: line items stayed 0.597 and manuals regressed).
  2. Section-level top-1 (0.62) is the weak spot, while division (0.81) is strong → most residual error is confusion within a division (e.g. 03 30 00 vs 03 50 00). A coarse-to-fine or contrastive approach targets exactly this.
  3. Out-of-taxonomy sections only score at division level; there is no other/hierarchical fallback.
  4. No level-3 / full 8-digit section codes — real spec books use them.

3. Definition of done (gates for any new model)

Gate Target How to verify
Powered eval line items ≥ 1,000; sections ≥ 400 Wilson CI ≤ ±3 pts
Measurable lift ≥ +3 pts line-item top-1 vs published, 95 % CI non-overlapping paired bootstrap / McNemar
No regression manuals + section division within 1 pt fixed eval sets
Calibrated ECE ≤ 0.05 after temperature scaling reliability curve
Useful abstention ≥ 90 % precision at ≥ 50 % coverage on line items coverage–precision curve
Reproducible one cloud/run_kaggle.py --run <name> + seed identical metrics ± noise

4. Plan

P0 — Make the measurement trustworthy, then feed it real data

P0.1 — Enlarge the held-out eval (do first). Why: everything downstream is unverifiable until this exists. How: assemble ≥1,000 real estimate line items and ≥400 real spec sections with true MasterFormat labels from public sources (§5); store as eval/line_items_v2.tsv, eval/manuals_v2.jsonl; report Wilson CIs + paired bootstrap for every comparison. Done when: CIs ≤ ±3 pts and the published vs candidate comparison reports a p-value. Effort: M.

P0.2 — Acquire real line-item training labels (not just eval). Why: the transformer's line-item ceiling is a data problem, not a capacity problem. How: mine public bid tabulations / agency item catalogs; map agency item codes to MasterFormat where a crosswalk exists; treat the rest as weak/self-supervised signal. Keep a strict train/eval split by project. Done when: ≥10k labelled items, none from the eval projects. Effort: L.

P0.3 — Domain-adaptive pretraining corpus. Why: the encoder only ever saw UFGS. How: collect public construction prose (specs, bid tabs, RFIs, submittal logs) beyond UFGS — VA design manuals, state DOT standard specs, Corps/UFC, NASA. Continued-MLM the encoder, then fine-tune. Done when: mf-0.5 beats the current model on the enlarged eval with non-overlapping CI. Effort: L (compute: 1–2 GPU-days).

P0.4 — Label space: level-3 sections + explicit out-of-taxonomy handling. Why: real specs use 8-digit codes and many sections outside the 171 groups. How: extend the taxonomy to full sections; add an other/<division> fallback and score it honestly; revisit group_of(). Effort: M.

P1 — Modelling (only once P0.1 is in place)

  • P1.1 Distillation from the TF-IDF/ensemble teacher using out-of-fold soft targets (never on-data teacher labels) into the transformer. Replaces the current brittle log-prob blend with one model.
  • P1.2 Coarse-to-fine / hierarchical head (division → group), directly attacking within-division confusion and giving graceful out-of-taxonomy fallback.
  • P1.3 Retrieval + rerank: kNN over labelled embeddings for the top-k, plus a cross-encoder to reorder them. Often the cheapest way to absorb new labelled data as it arrives.
  • P1.4 Calibration + abstention: temperature scaling on val; per-class thresholds; a REVIEW queue. Ship this regardless of accuracy — it converts a probabilistic classifier into a usable one.
  • P1.5 Multi-task: shared encoder with group + division + section heads; consistency regularisation between them.

P2 — Serving / ops

  • P2.1 Routing policy: transformer for whole sections / long text, ensemble for terse items (already the evidence-backed split).
  • P2.2 Quantization-aware / int8-calibrated ONNX (the int8 export currently drifts; add QAT or per-channel calibration and a browser benchmark).
  • P2.3 Golden-file CI on the enlarged eval; nightly regression that fails on a gate breach.
  • P2.4 Metrics dashboard with coverage/precision so abstention is tuned in production.

5. Data sourcing shortlist (public unless noted)

Source Use Note
UFGS + UFC train (in use) US federal, public domain
State DOT standard specs (Caltrans, TxDOT, WSDOT, MnDOT…) pretrain / train mostly public
VA design manuals & specifications pretrain / train US federal, public
NASA / DoD / Corps specs pretrain public
Public bid tabulations (agency open-data portals; city/state capital projects) line-item text + weak labels item number + description; codes vary by agency
SAM.gov / FPDS contract line items weak labels huge, noisy
Public BOQ / cost datasets (e.g. Kaggle construction cost) line items check licence
RSMeans / Gordian cost codes labels proprietary — licence required

6. Sequenced milestones

  1. M1 (now): publish the evaluation caveat + cloud/stats.py (CIs, abstention). Freeze mf-0.2+ensemble as the reference. ✅ started in this change.
  2. M2: enlarged eval (P0.1) from public bid tabs + manuals; re-score the reference.
  3. M3: domain-adaptive pretrain + retrain (mf-0.5) on P0.2/P0.3 data; validate with CIs (P0.3).
  4. M4: distillation + hierarchical head (mf-0.6); pick the best single model (P1.1/P1.2).
  5. M5: calibration/abstention + routing in serving; publish with a coverage–precision card (P1.4, P2.1).

7. Explicitly not doing

  • More synthetic line-item augmentation (tested: mf-0.3 — no line-item transfer, manual regression).
  • More TF-IDF feature engineering (already saturated; gains below the CI).
  • Reporting sub-CI improvements as wins on the model card.