Text Classification
Transformers
Joblib
ONNX
Safetensors
English
bert
masterformat
masterformat-classifier
csi-masterformat
construction
construction-technology
sequence-classification
tfidf
ensemble
specs
spec-writing
specifications
takeoff
estimating
cost-code
ufgs
public-domain
Eval Results (legacy)
text-embeddings-inference
Instructions to use constructelligence/masterformat-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use constructelligence/masterformat-classifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="constructelligence/masterformat-classifier")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("constructelligence/masterformat-classifier") model = AutoModelForSequenceClassification.from_pretrained("constructelligence/masterformat-classifier", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Add evidence caveat, calibration/abstention (91% precision at 50% coverage) and improvement roadmap
Browse files- IMPROVEMENT-PLAN.md +126 -0
- README.md +34 -0
IMPROVEMENT-PLAN.md
ADDED
|
@@ -0,0 +1,126 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# MasterFormat classifier — improvement plan
|
| 2 |
+
|
| 3 |
+
**Date:** 2026-10-05 · **Owner:** Constructelligence
|
| 4 |
+
**Published:** [`constructelligence/masterformat-classifier`](https://huggingface.co/constructelligence/masterformat-classifier)
|
| 5 |
+
**Current:** `mf-0.2` transformer + `ensemble/` (word+char TF-IDF blend).
|
| 6 |
+
|
| 7 |
+
| metric | published value | 95 % CI | n |
|
| 8 |
+
|---|---:|---|---:|
|
| 9 |
+
| line-item top-1 (hand-labelled estimates) | 0.702 | **[0.633, 0.762]** | 194 (191 in-taxonomy) |
|
| 10 |
+
| whole-section top-1 | 0.620 | **[0.540, 0.694]** | 153 |
|
| 11 |
+
| whole-section division | 0.804 | ±~6 pts | 153 |
|
| 12 |
+
| UFGS val top-1 | 0.588 | [0.579, 0.596] | 11,598 |
|
| 13 |
+
| manual-chunk top-1 | 0.330 | [0.319, 0.341] | 6,890 |
|
| 14 |
+
|
| 15 |
+
---
|
| 16 |
+
|
| 17 |
+
## 1. The binding constraint is evidence, not modelling
|
| 18 |
+
|
| 19 |
+
The two metrics that matter for the product — **hand-labelled line items** and **whole real spec sections** —
|
| 20 |
+
come from tiny samples. Their confidence intervals are **±6–8 points**, which means:
|
| 21 |
+
|
| 22 |
+
- the recent `0.686 → 0.702` ensemble change is **statistically indistinguishable from noise**;
|
| 23 |
+
- any model change smaller than ~10 points on line items cannot be validated on this eval;
|
| 24 |
+
- the tight, large-sample metric (`val`, ±0.9 pts) is **in-domain UFGS text**, not the target distribution —
|
| 25 |
+
a model can improve `val` a lot (see `mf-0.3`: 0.466 → 0.556) while getting *worse* on real manuals.
|
| 26 |
+
|
| 27 |
+
**Consequence:** the next gains are gated on **a bigger, real-world evaluation set**, not on architecture.
|
| 28 |
+
Chasing sub-CI improvements (more TF-IDF features, more seeds) is a trap we have already fallen into once.
|
| 29 |
+
|
| 30 |
+
## 2. Where the headroom actually is
|
| 31 |
+
|
| 32 |
+
1. **Line items (estimating).** TF-IDF dominates; the transformer adds little. True ceiling unknown because
|
| 33 |
+
the eval is underpowered. Needs *real* labelled items, not synthetic augmentation (`mf-0.3` proved
|
| 34 |
+
synthetic terse augmentation does not transfer: line items stayed 0.597 and manuals regressed).
|
| 35 |
+
2. **Section-level *top-1*** (0.62) is the weak spot, while *division* (0.81) is strong → most residual error
|
| 36 |
+
is **confusion within a division** (e.g. `03 30 00` vs `03 50 00`). A coarse-to-fine or contrastive
|
| 37 |
+
approach targets exactly this.
|
| 38 |
+
3. **Out-of-taxonomy** sections only score at division level; there is no `other`/hierarchical fallback.
|
| 39 |
+
4. **No level-3 / full 8-digit section codes** — real spec books use them.
|
| 40 |
+
|
| 41 |
+
## 3. Definition of done (gates for any new model)
|
| 42 |
+
|
| 43 |
+
| Gate | Target | How to verify |
|
| 44 |
+
|---|---|---|
|
| 45 |
+
| Powered eval | line items ≥ 1,000; sections ≥ 400 | Wilson CI ≤ ±3 pts |
|
| 46 |
+
| Measurable lift | ≥ +3 pts line-item top-1 vs published, 95 % CI non-overlapping | paired bootstrap / McNemar |
|
| 47 |
+
| No regression | manuals + section division within 1 pt | fixed eval sets |
|
| 48 |
+
| Calibrated | ECE ≤ 0.05 after temperature scaling | reliability curve |
|
| 49 |
+
| Useful abstention | ≥ 90 % precision at ≥ 50 % coverage on line items | coverage–precision curve |
|
| 50 |
+
| Reproducible | one `cloud/run_kaggle.py --run <name>` + seed | identical metrics ± noise |
|
| 51 |
+
|
| 52 |
+
## 4. Plan
|
| 53 |
+
|
| 54 |
+
### P0 — Make the measurement trustworthy, then feed it real data
|
| 55 |
+
|
| 56 |
+
**P0.1 — Enlarge the held-out eval (do first).**
|
| 57 |
+
*Why:* everything downstream is unverifiable until this exists. *How:* assemble ≥1,000 real estimate line
|
| 58 |
+
items and ≥400 real spec sections with true MasterFormat labels from public sources (§5); store as
|
| 59 |
+
`eval/line_items_v2.tsv`, `eval/manuals_v2.jsonl`; report Wilson CIs + paired bootstrap for every comparison.
|
| 60 |
+
*Done when:* CIs ≤ ±3 pts and the published vs candidate comparison reports a p-value. *Effort:* M.
|
| 61 |
+
|
| 62 |
+
**P0.2 — Acquire real line-item training labels (not just eval).**
|
| 63 |
+
*Why:* the transformer's line-item ceiling is a data problem, not a capacity problem. *How:* mine public bid
|
| 64 |
+
tabulations / agency item catalogs; map agency item codes to MasterFormat where a crosswalk exists; treat the
|
| 65 |
+
rest as weak/self-supervised signal. Keep a strict train/eval split by project. *Done when:* ≥10k labelled
|
| 66 |
+
items, none from the eval projects. *Effort:* L.
|
| 67 |
+
|
| 68 |
+
**P0.3 — Domain-adaptive pretraining corpus.**
|
| 69 |
+
*Why:* the encoder only ever saw UFGS. *How:* collect public construction prose (specs, bid tabs, RFIs, submittal
|
| 70 |
+
logs) beyond UFGS — VA design manuals, state DOT standard specs, Corps/UFC, NASA. Continued-MLM the encoder,
|
| 71 |
+
then fine-tune. *Done when:* mf-0.5 beats the current model on the *enlarged* eval with non-overlapping CI.
|
| 72 |
+
*Effort:* L (compute: 1–2 GPU-days).
|
| 73 |
+
|
| 74 |
+
**P0.4 — Label space: level-3 sections + explicit out-of-taxonomy handling.**
|
| 75 |
+
*Why:* real specs use 8-digit codes and many sections outside the 171 groups. *How:* extend the taxonomy to
|
| 76 |
+
full sections; add an `other/<division>` fallback and score it honestly; revisit `group_of()`. *Effort:* M.
|
| 77 |
+
|
| 78 |
+
### P1 — Modelling (only once P0.1 is in place)
|
| 79 |
+
|
| 80 |
+
- **P1.1 Distillation** from the TF-IDF/ensemble teacher using **out-of-fold** soft targets (never on-data
|
| 81 |
+
teacher labels) into the transformer. Replaces the current brittle log-prob blend with one model.
|
| 82 |
+
- **P1.2 Coarse-to-fine / hierarchical head** (division → group), directly attacking within-division confusion
|
| 83 |
+
and giving graceful out-of-taxonomy fallback.
|
| 84 |
+
- **P1.3 Retrieval + rerank:** kNN over labelled embeddings for the top-k, plus a cross-encoder to reorder
|
| 85 |
+
them. Often the cheapest way to absorb new labelled data as it arrives.
|
| 86 |
+
- **P1.4 Calibration + abstention:** temperature scaling on `val`; per-class thresholds; a `REVIEW` queue.
|
| 87 |
+
Ship this regardless of accuracy — it converts a probabilistic classifier into a usable one.
|
| 88 |
+
- **P1.5 Multi-task:** shared encoder with group + division + section heads; consistency regularisation between
|
| 89 |
+
them.
|
| 90 |
+
|
| 91 |
+
### P2 — Serving / ops
|
| 92 |
+
|
| 93 |
+
- **P2.1 Routing policy:** transformer for whole sections / long text, ensemble for terse items (already the
|
| 94 |
+
evidence-backed split).
|
| 95 |
+
- **P2.2 Quantization-aware / int8-calibrated ONNX** (the int8 export currently drifts; add QAT or
|
| 96 |
+
per-channel calibration and a browser benchmark).
|
| 97 |
+
- **P2.3 Golden-file CI** on the enlarged eval; nightly regression that fails on a gate breach.
|
| 98 |
+
- **P2.4 Metrics dashboard** with coverage/precision so abstention is tuned in production.
|
| 99 |
+
|
| 100 |
+
## 5. Data sourcing shortlist (public unless noted)
|
| 101 |
+
|
| 102 |
+
| Source | Use | Note |
|
| 103 |
+
|---|---|---|
|
| 104 |
+
| UFGS + UFC | train (in use) | US federal, public domain |
|
| 105 |
+
| State DOT standard specs (Caltrans, TxDOT, WSDOT, MnDOT…) | pretrain / train | mostly public |
|
| 106 |
+
| VA design manuals & specifications | pretrain / train | US federal, public |
|
| 107 |
+
| NASA / DoD / Corps specs | pretrain | public |
|
| 108 |
+
| Public **bid tabulations** (agency open-data portals; city/state capital projects) | line-item text + weak labels | item number + description; codes vary by agency |
|
| 109 |
+
| SAM.gov / FPDS contract line items | weak labels | huge, noisy |
|
| 110 |
+
| Public BOQ / cost datasets (e.g. Kaggle construction cost) | line items | check licence |
|
| 111 |
+
| RSMeans / Gordian cost codes | labels | **proprietary** — licence required |
|
| 112 |
+
|
| 113 |
+
## 6. Sequenced milestones
|
| 114 |
+
|
| 115 |
+
1. **M1 (now):** publish the evaluation caveat + `cloud/stats.py` (CIs, abstention). Freeze `mf-0.2`+ensemble
|
| 116 |
+
as the reference. ✅ started in this change.
|
| 117 |
+
2. **M2:** enlarged eval (P0.1) from public bid tabs + manuals; re-score the reference.
|
| 118 |
+
3. **M3:** domain-adaptive pretrain + retrain (`mf-0.5`) on P0.2/P0.3 data; validate with CIs (P0.3).
|
| 119 |
+
4. **M4:** distillation + hierarchical head (`mf-0.6`); pick the best single model (P1.1/P1.2).
|
| 120 |
+
5. **M5:** calibration/abstention + routing in serving; publish with a coverage–precision card (P1.4, P2.1).
|
| 121 |
+
|
| 122 |
+
## 7. Explicitly not doing
|
| 123 |
+
|
| 124 |
+
- More synthetic line-item augmentation (tested: `mf-0.3` — no line-item transfer, manual regression).
|
| 125 |
+
- More TF-IDF feature engineering (already saturated; gains below the CI).
|
| 126 |
+
- Reporting sub-CI improvements as wins on the model card.
|
README.md
CHANGED
|
@@ -155,6 +155,11 @@ sets are **fixed**, so those columns are directly comparable for every scorer. `
|
|
| 155 |
commercial project manuals; out-of-taxonomy gold labels (e.g. `22 40 00`) count only toward the division
|
| 156 |
score, so top-1 is over in-taxonomy items only.
|
| 157 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 158 |
**Reading it:** `mf-0.2` is the best scorer overall at section-level *division* accuracy (0.810). On short
|
| 159 |
line items a lexical TF-IDF model is stronger (0.696 alone), so the shipped **transformer + TF-IDF ensemble**
|
| 160 |
reaches **0.702** and is the production candidate; see *Ensemble*.
|
|
@@ -194,6 +199,34 @@ the weights on the 194-item test itself and does not generalise. The transformer
|
|
| 194 |
section *division* accuracy (0.810). A reasonable production split is the transformer for whole sections, the
|
| 195 |
ensemble for terse line items.
|
| 196 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 197 |
## Intended use
|
| 198 |
|
| 199 |
- **Use it for:** tagging construction text with a candidate MasterFormat group (top-3 shown), an
|
|
@@ -268,6 +301,7 @@ manual-chunk and whole-section accuracy. Constructelligence's proprietary models
|
|
| 268 |
- `onnx/` — ONNX fp32 and int8 exports.
|
| 269 |
- `ensemble/` — `tfidf.joblib` (word + char n-grams, ~112 MB), `blend.json`, `predict_ensemble.py`: the
|
| 270 |
transformer + TF-IDF line-item ensemble.
|
|
|
|
| 271 |
- `CITATION.cff`, `requirements.txt`.
|
| 272 |
|
| 273 |
## Citation
|
|
|
|
| 155 |
commercial project manuals; out-of-taxonomy gold labels (e.g. `22 40 00`) count only toward the division
|
| 156 |
score, so top-1 is over in-taxonomy items only.
|
| 157 |
|
| 158 |
+
**Evidence caveat.** The line-item and section sets are small. At n = 191 the 95 % Wilson interval on the
|
| 159 |
+
line-item top-1 is **[0.633, 0.762] (±6.4 pts)**, and whole-section is ±~8 pts — so differences of a few
|
| 160 |
+
points between models here are **noise**. Treat the `val` and `manual-chunk` rows (±1 pt) as the measurable
|
| 161 |
+
ones, and the line-item/section rows as indicative. Enlarging these eval sets is the top item on the roadmap.
|
| 162 |
+
|
| 163 |
**Reading it:** `mf-0.2` is the best scorer overall at section-level *division* accuracy (0.810). On short
|
| 164 |
line items a lexical TF-IDF model is stronger (0.696 alone), so the shipped **transformer + TF-IDF ensemble**
|
| 165 |
reaches **0.702** and is the production candidate; see *Ensemble*.
|
|
|
|
| 199 |
section *division* accuracy (0.810). A reasonable production split is the transformer for whole sections, the
|
| 200 |
ensemble for terse line items.
|
| 201 |
|
| 202 |
+
## Calibration and abstention
|
| 203 |
+
|
| 204 |
+
The ensemble is already close to calibrated; a single **temperature** `T = 0.90` (fit on `val`) lowers the
|
| 205 |
+
expected calibration error on `val` from **0.060 to 0.022**. More useful in production: send the
|
| 206 |
+
lowest-confidence predictions to a human instead of filing them.
|
| 207 |
+
|
| 208 |
+
Coverage → precision on the 194 line items (predictions sorted by confidence):
|
| 209 |
+
|
| 210 |
+
| auto-filed (coverage) | 100 % | 90 % | 80 % | 70 % | 50 % | 30 % |
|
| 211 |
+
|---|---:|---:|---:|---:|---:|---:|
|
| 212 |
+
| precision | 0.69 | 0.75 | 0.81 | **0.85** | **0.91** | 0.97 |
|
| 213 |
+
|
| 214 |
+
Routing the least-confident **30 %** to review lifts precision to **91 %**; routing 50 % gives **97 %**. So
|
| 215 |
+
the model is usable today as an **assist that flags its own uncertainty**, even though it is not accurate
|
| 216 |
+
enough to file unattended. Reproduce with `python cloud/stats.py`.
|
| 217 |
+
|
| 218 |
+
## Roadmap
|
| 219 |
+
|
| 220 |
+
Further gains are **evidence-bound**, not architecture-bound (see the repo's `IMPROVEMENT-PLAN.md`):
|
| 221 |
+
|
| 222 |
+
1. **Enlarge the eval** (≥1,000 line items, ≥400 sections) so changes are measurable at ±3 pts — currently
|
| 223 |
+
line-item changes below ~10 pts cannot be validated.
|
| 224 |
+
2. **Real line-item training data** (public bid tabulations, agency item catalogs) — the transformer's
|
| 225 |
+
line-item ceiling is a data problem; synthetic augmentation was tested and did not transfer (`mf-0.3`).
|
| 226 |
+
3. **Domain-adaptive pretraining** on construction prose beyond UFGS, then retrain.
|
| 227 |
+
4. **Distillation + hierarchical (division → group) head** to beat the TF-IDF blend with a single model.
|
| 228 |
+
5. **Level-3 / full-section codes** and an explicit out-of-taxonomy fallback.
|
| 229 |
+
|
| 230 |
## Intended use
|
| 231 |
|
| 232 |
- **Use it for:** tagging construction text with a candidate MasterFormat group (top-3 shown), an
|
|
|
|
| 301 |
- `onnx/` — ONNX fp32 and int8 exports.
|
| 302 |
- `ensemble/` — `tfidf.joblib` (word + char n-grams, ~112 MB), `blend.json`, `predict_ensemble.py`: the
|
| 303 |
transformer + TF-IDF line-item ensemble.
|
| 304 |
+
- `IMPROVEMENT-PLAN.md` — the prioritised roadmap (enlarged eval, data, domain-adaptive pretraining, distillation).
|
| 305 |
- `CITATION.cff`, `requirements.txt`.
|
| 306 |
|
| 307 |
## Citation
|