Text Classification
Transformers
Joblib
ONNX
Safetensors
English
bert
masterformat
masterformat-classifier
csi-masterformat
construction
construction-technology
sequence-classification
tfidf
ensemble
specs
spec-writing
specifications
takeoff
estimating
cost-code
ufgs
public-domain
Eval Results (legacy)
text-embeddings-inference
Instructions to use constructelligence/masterformat-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use constructelligence/masterformat-classifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="constructelligence/masterformat-classifier")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("constructelligence/masterformat-classifier") model = AutoModelForSequenceClassification.from_pretrained("constructelligence/masterformat-classifier", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 16,507 Bytes
3ccb233 1d29f6c 3ccb233 1d29f6c 3ccb233 1d29f6c 3ccb233 1d29f6c 3ccb233 1d29f6c 9d4c04e 1d29f6c 3ccb233 1d29f6c 3ccb233 b774c08 3ccb233 b774c08 1d29f6c b774c08 1d29f6c b774c08 3ccb233 b774c08 3ccb233 1d29f6c b774c08 3ccb233 1d29f6c 3ccb233 b774c08 ad1145a 3ccb233 1d29f6c 3ccb233 b774c08 3ccb233 b774c08 1d29f6c b774c08 1d29f6c 3ccb233 1d29f6c 3ccb233 1d29f6c b774c08 3ccb233 1d29f6c 3ccb233 1d29f6c b774c08 3ccb233 1d29f6c 3ccb233 1d29f6c b774c08 1d29f6c b774c08 2a65210 1d29f6c b774c08 1d29f6c b774c08 2a65210 b774c08 2a65210 1d29f6c b774c08 2a65210 b774c08 2a65210 b774c08 e43c8ca ad1145a 1d29f6c 9d4c04e ad1145a 9d4c04e ad1145a 9d4c04e ad1145a 9d4c04e ad1145a 9d4c04e e43c8ca 1d29f6c b774c08 1d29f6c b774c08 3ccb233 1d29f6c b774c08 1d29f6c b774c08 3ccb233 1d29f6c 3ccb233 b774c08 3ccb233 1d29f6c b774c08 1d29f6c b774c08 1d29f6c 9d4c04e ad1145a 9d4c04e 1d29f6c b774c08 1d29f6c ad1145a e43c8ca 1d29f6c b774c08 1d29f6c 3ccb233 1d29f6c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 | ---
language:
- en
license: mit
library_name: transformers
pipeline_tag: text-classification
base_model: BAAI/bge-small-en-v1.5
tags:
- masterformat
- masterformat-classifier
- csi-masterformat
- construction
- construction-technology
- text-classification
- sequence-classification
- bert
- onnx
- safetensors
- tfidf
- ensemble
- specs
- spec-writing
- specifications
- takeoff
- estimating
- cost-code
- ufgs
- public-domain
metrics:
- accuracy
model-index:
- name: masterformat-classifier
results:
- task:
type: text-classification
name: MasterFormat level-2 group classification
dataset:
type: ufgs-val-v2
name: UFGS held-out text units (11,598)
split: validation
metrics:
- type: accuracy
value: 0.466
name: Top-1 accuracy (mf-0.2)
widget:
- text: EPDM membrane roofing
example_title: Membrane roofing
- text: Wet pipe sprinkler system, light hazard
example_title: Fire suppression
- text: 8" CMU wall, grout filled at 32" o.c.
example_title: Unit masonry
- text: Addressable fire alarm system, devices and panel
example_title: Fire alarm
---
# MasterFormat classifier (mf-0.2) β construction spec & line-item classification
A BERT text classifier that maps **construction line items, specification paragraphs, submittals and
section titles** to one of **171 MasterFormat level-2 groups** (e.g. `03 30 00 Cast-in-Place Concrete`,
`23 30 00 HVAC Air Distribution`) across **32 divisions**. It is the machine-learning counterpart to a
cost-code lookup table: hand it "EPDM membrane roofing" and it returns the MasterFormat group, ranked.
English text, one label per input.
Built for estimating, takeoff, spec-writing and RAG pipelines that need to tag free text with CSI
MasterFormat codes without a human choosing from a 171-row list.
> **Completed GPU fine-tune.** `mf-0.2` is the full run that `mf-0.1` (step 700) was an early checkpoint of:
> **6,132 steps / 12 epochs** on a Kaggle T4 over the rebuilt v2 dataset. Top-1 accuracy on the v2 UFGS
> validation split is **46.6 %** β **2.3Γ** `mf-0.1` (20.0 %). Line-item accuracy on 194 hand-labelled
> estimate items is **59.7 %** (was 28.3 %), and whole-section accuracy is **60.7 %** with **81.0 %** at
> division level. Because the transformer is trained on specification prose, short estimate line items are its
> weak spot; the repo therefore also ships a **transformer + TF-IDF ensemble** (`ensemble/`) that lifts
> line-item top-1 to **70.2 %** and improves manual-chunk and section accuracy; see *Ensemble*.
MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or
endorsed by them.
## Quick start
```python
from transformers import pipeline
clf = pipeline("text-classification", model="constructelligence/masterformat-classifier", top_k=3)
for p in clf("EPDM membrane roofing"):
print(p["label"], round(p["score"], 3))
# 07 50 00 Membrane Roofing 0.863
# 07 10 00 Dampproofing and Waterproofing 0.068
# 07 30 00 Steep Slope Roofing 0.047
```
Batched, with the level-1 division roll-up and JSON output:
```bash
pip install -r requirements.txt # transformers + torch
python predict.py --top-k 5 --divisions "Addressable fire alarm system, devices and panel"
echo "12\" RCP storm drain pipe" | python predict.py -
python predict.py --file items.txt --json > out.json
```
ONNX (no torch, ~34 MB int8):
```bash
pip install onnxruntime transformers
python predict.py --onnx onnx/model_quantized.onnx "Wet pipe sprinkler system, light hazard"
```
The encoder is English-only; inputs should be English.
### ONNX exports
`onnx/model.onnx` (fp32, 134 MB) and `onnx/model_quantized.onnx` (dynamic int8, 34 MB) are exported with
dynamic batch and sequence axes, in the layout `transformers.js` / Optimum expect β inputs `input_ids`,
`attention_mask`, `token_type_ids`, output `logits`. The fp32 export matches PyTorch on a 4-text smoke set
(max |Ξ logit| 1e-5, argmax agreement 4/4). Dynamic int8 quantization shifts the logits more than it did for
`mf-0.1` β the better-trained head is more confident β so verify the quantized model on your own inputs;
the fp32 export is the safer default. Reproduce with `python scripts/export_onnx.py runs/mf-0.2`.
## Model
- **Base:** [`BAAI/bge-small-en-v1.5`](https://huggingface.co/BAAI/bge-small-en-v1.5) (BERT, 33M parameters, 384-d, MIT).
- **Head:** 171-way sequence-classification layer; `id2label` / `label2id` are in [`config.json`](config.json).
- **Recipe:** `label smoothing 0.05`, max length 64, batch 128, 12 epochs (6,132 steps), fp16 autocast,
AdamW with 6 % linear warmup; body LR 5e-5, head LR 1e-3; embeddings and the first 4 encoder layers frozen.
- **Checkpoint:** `train_state.json` records `{"step": 6132, "val_acc": 0.4661}`.
- **Trained on:** Kaggle T4, ~16.5 min wall-clock, from `BAAI/bge-small-en-v1.5` (not resumed from `mf-0.1`).
## Results
All numbers are measured in this project on held-out data. The `line_items`, `manuals` and `manual_sections`
sets are **fixed**, so those columns are directly comparable for every scorer. `val` is the **v2** UFGS split
(11,598 units).
**v2 validation split**
| Scorer | val top-1 | val top-3 | val division |
|---|---:|---:|---:|
| **mf-0.2 (step 6132, published)** | 0.466 | 0.641 | 0.593 |
| mf-0.3 (v3 data, unfrozen) | **0.556** | **0.710** | **0.662** |
| mf-0.4 (v2 data, unfrozen) | 0.491 | 0.657 | 0.609 |
| mf-0.1 (step 700) | 0.200 | 0.372 | 0.356 |
| TF-IDF (word + char n-grams) | 0.571 | 0.732 | 0.670 |
**Fixed held-out sets (line items Β· manual chunks Β· whole sections)**
| Scorer | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division |
|---|---:|---:|---:|---:|
| **mf-0.2 (step 6132, published)** | 0.597 | 0.280 | 0.607 | **0.810** |
| mf-0.3 (v3 data, unfrozen) | 0.597 | 0.259 | 0.547 | 0.765 |
| mf-0.4 (v2 data, unfrozen) | **0.618** | 0.270 | 0.567 | 0.817 |
| mf-0.1 (step 700) | 0.283 | 0.176 | 0.287 | 0.588 |
| TF-IDF (word + char n-grams) | 0.696 | 0.287 | 0.587 | 0.752 |
| **Ensemble mf-0.2 + TF-IDF (w = 0.25)** | **0.702** | **0.330** | **0.620** | 0.804 |
`line_items` = 194 hand-labelled estimate line items; `manuals` = 6,890 chunks from 153 sections of two real
commercial project manuals; out-of-taxonomy gold labels (e.g. `22 40 00`) count only toward the division
score, so top-1 is over in-taxonomy items only.
**Evidence caveat.** The line-item and section sets are small. At n = 191 the 95 % Wilson interval on the
line-item top-1 is **[0.633, 0.762] (Β±6.4 pts)**, and whole-section is Β±~8 pts β so differences of a few
points between models here are **noise**. Treat the `val` and `manual-chunk` rows (Β±1 pt) as the measurable
ones, and the line-item/section rows as indicative. Enlarging these eval sets is the top item on the roadmap.
**Reading it:** `mf-0.2` is the best scorer overall at section-level *division* accuracy (0.810). On short
line items a lexical TF-IDF model is stronger (0.696 alone), so the shipped **transformer + TF-IDF ensemble**
reaches **0.702** and is the production candidate; see *Ensemble*.
## Ensemble (closing the line-item gap)
The transformer is trained on specification prose, so on **short, terse estimate line items** a lexical
TF-IDF model is still stronger (0.696 vs 0.597 top-1). To close that gap without giving up the transformer's
long-text accuracy, this repo ships a **log-probability ensemble** of `mf-0.2` and a TF-IDF + SGD scorer:
```
p = softmax( 0.25 Β· log_softmax(mf-0.2) + 0.75 Β· log_softmax(tfidf) )
```
The 0.25 weight is chosen on the UFGS **validation** split β not on the line-item test set. `ensemble/tfidf.joblib`
holds the fitted vectorizer (word 1β2 grams, 100k features, plus `char_wb` 3β5 grams, 150k features, `min_df`
2/3, sublinear tf) and SGD logistic classifier; `ensemble/blend.json` records the weight and all metrics;
`ensemble/predict_ensemble.py` runs the blend.
| Scorer | val top-1 | val division | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division |
|---|---:|---:|---:|---:|---:|---:|
| mf-0.2 (transformer) | 0.466 | 0.593 | 0.597 | 0.280 | 0.607 | **0.810** |
| TF-IDF (word + char n-grams) | 0.571 | 0.670 | 0.696 | 0.287 | 0.587 | 0.752 |
| **Ensemble (w = 0.25)** | **0.588** | **0.689** | **0.702** | **0.330** | **0.620** | 0.804 |
```bash
pip install transformers torch scikit-learn joblib
python ensemble/predict_ensemble.py --model constructelligence/masterformat-classifier \
--tfidf ensemble/tfidf.joblib --top-k 3 --divisions "4000 psi concrete slab on grade"
```
**Honest caveat.** The ensemble beats TF-IDF alone mainly on the transformer's strengths β validation top-1
(0.588 vs 0.571), manual-chunk top-1 (0.330 vs 0.287) and whole-section top-1 (0.620 vs 0.587) β while line
items are nearly saturated by TF-IDF (0.702 vs 0.696). Adding bge-small embeddings does **not** help once the
blend weight is chosen on `val` (its weight goes to ~0); an earlier 0.712 line-item figure came from tuning
the weights on the 194-item test itself and does not generalise. The transformer remains the best scorer at
section *division* accuracy (0.810). A reasonable production split is the transformer for whole sections, the
ensemble for terse line items.
## Calibration and abstention
The ensemble is already close to calibrated; a single **temperature** `T = 0.90` (fit on `val`) lowers the
expected calibration error on `val` from **0.060 to 0.022**. More useful in production: send the
lowest-confidence predictions to a human instead of filing them.
Coverage β precision on the 194 line items (predictions sorted by confidence):
| auto-filed (coverage) | 100 % | 90 % | 80 % | 70 % | 50 % | 30 % |
|---|---:|---:|---:|---:|---:|---:|
| precision | 0.69 | 0.75 | 0.81 | **0.85** | **0.91** | 0.97 |
Routing the least-confident **30 %** to review lifts precision to **91 %**; routing 50 % gives **97 %**. So
the model is usable today as an **assist that flags its own uncertainty**, even though it is not accurate
enough to file unattended. Reproduce with `python cloud/stats.py`.
## Roadmap
Further gains are **evidence-bound**, not architecture-bound (see the repo's `IMPROVEMENT-PLAN.md`):
1. **Enlarge the eval** (β₯1,000 line items, β₯400 sections) so changes are measurable at Β±3 pts β currently
line-item changes below ~10 pts cannot be validated.
2. **Real line-item training data** (public bid tabulations, agency item catalogs) β the transformer's
line-item ceiling is a data problem; synthetic augmentation was tested and did not transfer (`mf-0.3`).
3. **Domain-adaptive pretraining** on construction prose beyond UFGS, then retrain.
4. **Distillation + hierarchical (division β group) head** to beat the TF-IDF blend with a single model.
5. **Level-3 / full-section codes** and an explicit out-of-taxonomy fallback.
## Intended use
- **Use it for:** tagging construction text with a candidate MasterFormat group (top-3 shown), an
auto-classification step for estimating or spec workflows, and as a strong base to fine-tune or distil.
- **Do not use it for:** unattended production takeoff, bid pricing, code compliance, or anything where a
wrong cost code has financial or contractual consequences. Keep a human in the loop.
- **Not a substitute for review.** Classifies text content only β if the input already contains a
MasterFormat number, read the number instead.
## Training data and taxonomy
- **Source:** UFGS (Unified Facilities Guide Specifications) `.SEC` files β US federal works in the
**public domain**. Paragraphs and titles are parsed into labelled text units; boilerplate shared by
multiple sections is dropped, cross-references are stripped so section numbers cannot leak labels, and
long paragraphs are cut on sentence boundaries into 8β60-word windows.
- **Split:** held out by a hash of the normalised source text (10 %), so a paragraph and its augmentations
stay on the same side. **65,496 train / 11,598 val** rows (the v2 set).
- **Taxonomy:** 171 level-2 groups over 32 MasterFormat divisions; group numbers follow the MasterFormat
numbering convention and the short names are this project's own (see `config.json`).
## Limitations
- **Below the ensemble on line items.** 59.7 % vs 70.2 % on 194 hand-labelled estimate items; do not deploy
as the sole classifier.
- **Domain skew.** UFGS over-represents heavy-civil, water/wastewater and process work relative to commercial
building estimates; the training set is class-balanced, so raw predictions do not reflect building-project
priors.
- **Out-of-taxonomy inputs** (sections whose level-2 group is not among the 171) can only be scored at
division level.
- **Short, terse line items are the hardest inputs**; division (2-digit) accuracy is consistently higher than
group (6-digit) accuracy.
- **Quantized ONNX drifts.** The int8 export is smaller but less faithful than `mf-0.1`'s; prefer fp32.
## Bias, risks and safety
- **Estimating bias.** Class-balanced training over a public-domain federal corpus does not represent any
particular firm's cost structure or regional practice. Do not treat output as a standard or an authority.
- **Trademark.** MasterFormat is a registered trademark of CSI / CSC; this model is not endorsed by them and
its group names are the project's own short descriptions, not CSI's official titles.
- **Privacy.** The model runs locally; no input text leaves your machine unless you call a hosted endpoint.
## FAQ
**What is MasterFormat?** The CSI/CSC MasterFormat is the North American standard for organising construction
specifications and cost data into numbered divisions and sections. This model predicts the **level-2 group**
(a 6-digit code such as `03 30 00 Cast-in-Place Concrete`), not the full section number.
**How is this different from `mf-0.1`?** `mf-0.1` was step 700 of an interrupted CPU run; `mf-0.2` is the
completed 12-epoch GPU fine-tune on the rebuilt dataset. Same architecture, ~2.3Γ the validation accuracy.
**Can it classify a whole specification section?** Yes β average the model's log-probabilities over a
section's chunks. On two real project manuals that gives 60.7 % top-1 and 81.0 % at division level.
**Can it read a MasterFormat number out of the text?** No. It classifies the description. If the number is
already present, parse it directly.
**Does it work offline / in the browser?** Yes β the ONNX exports are intended for `onnxruntime` and
`transformers.js`.
**Is a better model available?** Yes β this repo ships a **transformer + TF-IDF ensemble** (`ensemble/`) that
scores **70.2 % top-1 on hand-labelled line items** (vs 59.7 % for the transformer alone) and improves
manual-chunk and whole-section accuracy. Constructelligence's proprietary models are at
[constructelligence.co](https://constructelligence.co).
## Files
- `model.safetensors`, `config.json`, `tokenizer.json`, `tokenizer_config.json`, `vocab.txt`,
`special_tokens_map.json` β standard `transformers` checkpoint.
- `train_state.json` β step and validation accuracy of the saved checkpoint.
- `metrics.json` β full held-out evaluation (`val`, `line_items`, `manuals`, `manual_sections`).
- `predict.py` β CLI example: batching, `--top-k`, `--divisions`, `--json`, stdin/file input, `--onnx`.
- `onnx/` β ONNX fp32 and int8 exports.
- `ensemble/` β `tfidf.joblib` (word + char n-grams, ~112 MB), `blend.json`, `predict_ensemble.py`: the
transformer + TF-IDF line-item ensemble.
- `IMPROVEMENT-PLAN.md` β the prioritised roadmap (enlarged eval, data, domain-adaptive pretraining, distillation).
- `CITATION.cff`, `requirements.txt`.
## Citation
```bibtex
@misc{constructelligence_masterformat_classifier,
title = {MasterFormat Classifier (mf-0.2): construction spec and line-item classification},
author = {Constructelligence},
year = {2026},
howpublished = {\url{https://huggingface.co/constructelligence/masterformat-classifier}},
note = {Fine-tuned from BAAI/bge-small-en-v1.5; 171-way MasterFormat level-2 classifier}
}
```
## Licence and attribution
Released under the **MIT licence**, matching the base model. UFGS source text is public domain. MasterFormat
is a registered trademark of CSI / CSC; this project is not affiliated with or endorsed by them.
Constructelligence's production models are available at [constructelligence.co](https://constructelligence.co).
|