Text Classification
Transformers
Joblib
ONNX
Safetensors
English
bert
masterformat
masterformat-classifier
csi-masterformat
construction
construction-technology
sequence-classification
tfidf
ensemble
specs
spec-writing
specifications
takeoff
estimating
cost-code
ufgs
public-domain
Eval Results (legacy)
text-embeddings-inference
Instructions to use constructelligence/masterformat-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use constructelligence/masterformat-classifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="constructelligence/masterformat-classifier")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("constructelligence/masterformat-classifier") model = AutoModelForSequenceClassification.from_pretrained("constructelligence/masterformat-classifier", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Add evidence caveat, calibration/abstention (91% precision at 50% coverage) and improvement roadmap
e43c8ca verified |
Download IMPROVEMENT-PLAN.md from constructelligence/masterformat-classifier: direct link, hf CLI and curl.
- Browser
- Download file 7.79 kB
-
https://huggingface.co/constructelligence/masterformat-classifier/resolve/main/IMPROVEMENT-PLAN.md
- Command line
-
hf download hf://constructelligence/masterformat-classifier/IMPROVEMENT-PLAN.md
-
curl -L -o IMPROVEMENT-PLAN.md https://huggingface.co/constructelligence/masterformat-classifier/resolve/main/IMPROVEMENT-PLAN.md
7.79 kB
| # MasterFormat classifier — improvement plan | |
| **Date:** 2026-10-05 · **Owner:** Constructelligence | |
| **Published:** [`constructelligence/masterformat-classifier`](https://huggingface.co/constructelligence/masterformat-classifier) | |
| **Current:** `mf-0.2` transformer + `ensemble/` (word+char TF-IDF blend). | |
| | metric | published value | 95 % CI | n | | |
| |---|---:|---|---:| | |
| | line-item top-1 (hand-labelled estimates) | 0.702 | **[0.633, 0.762]** | 194 (191 in-taxonomy) | | |
| | whole-section top-1 | 0.620 | **[0.540, 0.694]** | 153 | | |
| | whole-section division | 0.804 | ±~6 pts | 153 | | |
| | UFGS val top-1 | 0.588 | [0.579, 0.596] | 11,598 | | |
| | manual-chunk top-1 | 0.330 | [0.319, 0.341] | 6,890 | | |
| --- | |
| ## 1. The binding constraint is evidence, not modelling | |
| The two metrics that matter for the product — **hand-labelled line items** and **whole real spec sections** — | |
| come from tiny samples. Their confidence intervals are **±6–8 points**, which means: | |
| - the recent `0.686 → 0.702` ensemble change is **statistically indistinguishable from noise**; | |
| - any model change smaller than ~10 points on line items cannot be validated on this eval; | |
| - the tight, large-sample metric (`val`, ±0.9 pts) is **in-domain UFGS text**, not the target distribution — | |
| a model can improve `val` a lot (see `mf-0.3`: 0.466 → 0.556) while getting *worse* on real manuals. | |
| **Consequence:** the next gains are gated on **a bigger, real-world evaluation set**, not on architecture. | |
| Chasing sub-CI improvements (more TF-IDF features, more seeds) is a trap we have already fallen into once. | |
| ## 2. Where the headroom actually is | |
| 1. **Line items (estimating).** TF-IDF dominates; the transformer adds little. True ceiling unknown because | |
| the eval is underpowered. Needs *real* labelled items, not synthetic augmentation (`mf-0.3` proved | |
| synthetic terse augmentation does not transfer: line items stayed 0.597 and manuals regressed). | |
| 2. **Section-level *top-1*** (0.62) is the weak spot, while *division* (0.81) is strong → most residual error | |
| is **confusion within a division** (e.g. `03 30 00` vs `03 50 00`). A coarse-to-fine or contrastive | |
| approach targets exactly this. | |
| 3. **Out-of-taxonomy** sections only score at division level; there is no `other`/hierarchical fallback. | |
| 4. **No level-3 / full 8-digit section codes** — real spec books use them. | |
| ## 3. Definition of done (gates for any new model) | |
| | Gate | Target | How to verify | | |
| |---|---|---| | |
| | Powered eval | line items ≥ 1,000; sections ≥ 400 | Wilson CI ≤ ±3 pts | | |
| | Measurable lift | ≥ +3 pts line-item top-1 vs published, 95 % CI non-overlapping | paired bootstrap / McNemar | | |
| | No regression | manuals + section division within 1 pt | fixed eval sets | | |
| | Calibrated | ECE ≤ 0.05 after temperature scaling | reliability curve | | |
| | Useful abstention | ≥ 90 % precision at ≥ 50 % coverage on line items | coverage–precision curve | | |
| | Reproducible | one `cloud/run_kaggle.py --run <name>` + seed | identical metrics ± noise | | |
| ## 4. Plan | |
| ### P0 — Make the measurement trustworthy, then feed it real data | |
| **P0.1 — Enlarge the held-out eval (do first).** | |
| *Why:* everything downstream is unverifiable until this exists. *How:* assemble ≥1,000 real estimate line | |
| items and ≥400 real spec sections with true MasterFormat labels from public sources (§5); store as | |
| `eval/line_items_v2.tsv`, `eval/manuals_v2.jsonl`; report Wilson CIs + paired bootstrap for every comparison. | |
| *Done when:* CIs ≤ ±3 pts and the published vs candidate comparison reports a p-value. *Effort:* M. | |
| **P0.2 — Acquire real line-item training labels (not just eval).** | |
| *Why:* the transformer's line-item ceiling is a data problem, not a capacity problem. *How:* mine public bid | |
| tabulations / agency item catalogs; map agency item codes to MasterFormat where a crosswalk exists; treat the | |
| rest as weak/self-supervised signal. Keep a strict train/eval split by project. *Done when:* ≥10k labelled | |
| items, none from the eval projects. *Effort:* L. | |
| **P0.3 — Domain-adaptive pretraining corpus.** | |
| *Why:* the encoder only ever saw UFGS. *How:* collect public construction prose (specs, bid tabs, RFIs, submittal | |
| logs) beyond UFGS — VA design manuals, state DOT standard specs, Corps/UFC, NASA. Continued-MLM the encoder, | |
| then fine-tune. *Done when:* mf-0.5 beats the current model on the *enlarged* eval with non-overlapping CI. | |
| *Effort:* L (compute: 1–2 GPU-days). | |
| **P0.4 — Label space: level-3 sections + explicit out-of-taxonomy handling.** | |
| *Why:* real specs use 8-digit codes and many sections outside the 171 groups. *How:* extend the taxonomy to | |
| full sections; add an `other/<division>` fallback and score it honestly; revisit `group_of()`. *Effort:* M. | |
| ### P1 — Modelling (only once P0.1 is in place) | |
| - **P1.1 Distillation** from the TF-IDF/ensemble teacher using **out-of-fold** soft targets (never on-data | |
| teacher labels) into the transformer. Replaces the current brittle log-prob blend with one model. | |
| - **P1.2 Coarse-to-fine / hierarchical head** (division → group), directly attacking within-division confusion | |
| and giving graceful out-of-taxonomy fallback. | |
| - **P1.3 Retrieval + rerank:** kNN over labelled embeddings for the top-k, plus a cross-encoder to reorder | |
| them. Often the cheapest way to absorb new labelled data as it arrives. | |
| - **P1.4 Calibration + abstention:** temperature scaling on `val`; per-class thresholds; a `REVIEW` queue. | |
| Ship this regardless of accuracy — it converts a probabilistic classifier into a usable one. | |
| - **P1.5 Multi-task:** shared encoder with group + division + section heads; consistency regularisation between | |
| them. | |
| ### P2 — Serving / ops | |
| - **P2.1 Routing policy:** transformer for whole sections / long text, ensemble for terse items (already the | |
| evidence-backed split). | |
| - **P2.2 Quantization-aware / int8-calibrated ONNX** (the int8 export currently drifts; add QAT or | |
| per-channel calibration and a browser benchmark). | |
| - **P2.3 Golden-file CI** on the enlarged eval; nightly regression that fails on a gate breach. | |
| - **P2.4 Metrics dashboard** with coverage/precision so abstention is tuned in production. | |
| ## 5. Data sourcing shortlist (public unless noted) | |
| | Source | Use | Note | | |
| |---|---|---| | |
| | UFGS + UFC | train (in use) | US federal, public domain | | |
| | State DOT standard specs (Caltrans, TxDOT, WSDOT, MnDOT…) | pretrain / train | mostly public | | |
| | VA design manuals & specifications | pretrain / train | US federal, public | | |
| | NASA / DoD / Corps specs | pretrain | public | | |
| | Public **bid tabulations** (agency open-data portals; city/state capital projects) | line-item text + weak labels | item number + description; codes vary by agency | | |
| | SAM.gov / FPDS contract line items | weak labels | huge, noisy | | |
| | Public BOQ / cost datasets (e.g. Kaggle construction cost) | line items | check licence | | |
| | RSMeans / Gordian cost codes | labels | **proprietary** — licence required | | |
| ## 6. Sequenced milestones | |
| 1. **M1 (now):** publish the evaluation caveat + `cloud/stats.py` (CIs, abstention). Freeze `mf-0.2`+ensemble | |
| as the reference. ✅ started in this change. | |
| 2. **M2:** enlarged eval (P0.1) from public bid tabs + manuals; re-score the reference. | |
| 3. **M3:** domain-adaptive pretrain + retrain (`mf-0.5`) on P0.2/P0.3 data; validate with CIs (P0.3). | |
| 4. **M4:** distillation + hierarchical head (`mf-0.6`); pick the best single model (P1.1/P1.2). | |
| 5. **M5:** calibration/abstention + routing in serving; publish with a coverage–precision card (P1.4, P2.1). | |
| ## 7. Explicitly not doing | |
| - More synthetic line-item augmentation (tested: `mf-0.3` — no line-item transfer, manual regression). | |
| - More TF-IDF feature engineering (already saturated; gains below the CI). | |
| - Reporting sub-CI improvements as wins on the model card. | |