Text Classification
Transformers
Joblib
ONNX
Safetensors
English
bert
masterformat
masterformat-classifier
csi-masterformat
construction
construction-technology
sequence-classification
tfidf
ensemble
specs
spec-writing
specifications
takeoff
estimating
cost-code
ufgs
public-domain
Eval Results (legacy)
text-embeddings-inference
Instructions to use constructelligence/masterformat-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use constructelligence/masterformat-classifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="constructelligence/masterformat-classifier")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("constructelligence/masterformat-classifier") model = AutoModelForSequenceClassification.from_pretrained("constructelligence/masterformat-classifier", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Add evidence caveat, calibration/abstention (91% precision at 50% coverage) and improvement roadmap
e43c8ca verified |
Download README.md from constructelligence/masterformat-classifier: direct link, hf CLI and curl.
- Browser
- Download file 16.5 kB
-
https://huggingface.co/constructelligence/masterformat-classifier/resolve/main/README.md
- Command line
-
hf download hf://constructelligence/masterformat-classifier/README.md
-
curl -L -o README.md https://huggingface.co/constructelligence/masterformat-classifier/resolve/main/README.md
16.5 kB
| language: | |
| - en | |
| license: mit | |
| library_name: transformers | |
| pipeline_tag: text-classification | |
| base_model: BAAI/bge-small-en-v1.5 | |
| tags: | |
| - masterformat | |
| - masterformat-classifier | |
| - csi-masterformat | |
| - construction | |
| - construction-technology | |
| - text-classification | |
| - sequence-classification | |
| - bert | |
| - onnx | |
| - safetensors | |
| - tfidf | |
| - ensemble | |
| - specs | |
| - spec-writing | |
| - specifications | |
| - takeoff | |
| - estimating | |
| - cost-code | |
| - ufgs | |
| - public-domain | |
| metrics: | |
| - accuracy | |
| model-index: | |
| - name: masterformat-classifier | |
| results: | |
| - task: | |
| type: text-classification | |
| name: MasterFormat level-2 group classification | |
| dataset: | |
| type: ufgs-val-v2 | |
| name: UFGS held-out text units (11,598) | |
| split: validation | |
| metrics: | |
| - type: accuracy | |
| value: 0.466 | |
| name: Top-1 accuracy (mf-0.2) | |
| widget: | |
| - text: EPDM membrane roofing | |
| example_title: Membrane roofing | |
| - text: Wet pipe sprinkler system, light hazard | |
| example_title: Fire suppression | |
| - text: 8" CMU wall, grout filled at 32" o.c. | |
| example_title: Unit masonry | |
| - text: Addressable fire alarm system, devices and panel | |
| example_title: Fire alarm | |
| # MasterFormat classifier (mf-0.2) β construction spec & line-item classification | |
| A BERT text classifier that maps **construction line items, specification paragraphs, submittals and | |
| section titles** to one of **171 MasterFormat level-2 groups** (e.g. `03 30 00 Cast-in-Place Concrete`, | |
| `23 30 00 HVAC Air Distribution`) across **32 divisions**. It is the machine-learning counterpart to a | |
| cost-code lookup table: hand it "EPDM membrane roofing" and it returns the MasterFormat group, ranked. | |
| English text, one label per input. | |
| Built for estimating, takeoff, spec-writing and RAG pipelines that need to tag free text with CSI | |
| MasterFormat codes without a human choosing from a 171-row list. | |
| > **Completed GPU fine-tune.** `mf-0.2` is the full run that `mf-0.1` (step 700) was an early checkpoint of: | |
| > **6,132 steps / 12 epochs** on a Kaggle T4 over the rebuilt v2 dataset. Top-1 accuracy on the v2 UFGS | |
| > validation split is **46.6 %** β **2.3Γ** `mf-0.1` (20.0 %). Line-item accuracy on 194 hand-labelled | |
| > estimate items is **59.7 %** (was 28.3 %), and whole-section accuracy is **60.7 %** with **81.0 %** at | |
| > division level. Because the transformer is trained on specification prose, short estimate line items are its | |
| > weak spot; the repo therefore also ships a **transformer + TF-IDF ensemble** (`ensemble/`) that lifts | |
| > line-item top-1 to **70.2 %** and improves manual-chunk and section accuracy; see *Ensemble*. | |
| MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or | |
| endorsed by them. | |
| ## Quick start | |
| ```python | |
| from transformers import pipeline | |
| clf = pipeline("text-classification", model="constructelligence/masterformat-classifier", top_k=3) | |
| for p in clf("EPDM membrane roofing"): | |
| print(p["label"], round(p["score"], 3)) | |
| # 07 50 00 Membrane Roofing 0.863 | |
| # 07 10 00 Dampproofing and Waterproofing 0.068 | |
| # 07 30 00 Steep Slope Roofing 0.047 | |
| ``` | |
| Batched, with the level-1 division roll-up and JSON output: | |
| ```bash | |
| pip install -r requirements.txt # transformers + torch | |
| python predict.py --top-k 5 --divisions "Addressable fire alarm system, devices and panel" | |
| echo "12\" RCP storm drain pipe" | python predict.py - | |
| python predict.py --file items.txt --json > out.json | |
| ``` | |
| ONNX (no torch, ~34 MB int8): | |
| ```bash | |
| pip install onnxruntime transformers | |
| python predict.py --onnx onnx/model_quantized.onnx "Wet pipe sprinkler system, light hazard" | |
| ``` | |
| The encoder is English-only; inputs should be English. | |
| ### ONNX exports | |
| `onnx/model.onnx` (fp32, 134 MB) and `onnx/model_quantized.onnx` (dynamic int8, 34 MB) are exported with | |
| dynamic batch and sequence axes, in the layout `transformers.js` / Optimum expect β inputs `input_ids`, | |
| `attention_mask`, `token_type_ids`, output `logits`. The fp32 export matches PyTorch on a 4-text smoke set | |
| (max |Ξ logit| 1e-5, argmax agreement 4/4). Dynamic int8 quantization shifts the logits more than it did for | |
| `mf-0.1` β the better-trained head is more confident β so verify the quantized model on your own inputs; | |
| the fp32 export is the safer default. Reproduce with `python scripts/export_onnx.py runs/mf-0.2`. | |
| ## Model | |
| - **Base:** [`BAAI/bge-small-en-v1.5`](https://huggingface.co/BAAI/bge-small-en-v1.5) (BERT, 33M parameters, 384-d, MIT). | |
| - **Head:** 171-way sequence-classification layer; `id2label` / `label2id` are in [`config.json`](config.json). | |
| - **Recipe:** `label smoothing 0.05`, max length 64, batch 128, 12 epochs (6,132 steps), fp16 autocast, | |
| AdamW with 6 % linear warmup; body LR 5e-5, head LR 1e-3; embeddings and the first 4 encoder layers frozen. | |
| - **Checkpoint:** `train_state.json` records `{"step": 6132, "val_acc": 0.4661}`. | |
| - **Trained on:** Kaggle T4, ~16.5 min wall-clock, from `BAAI/bge-small-en-v1.5` (not resumed from `mf-0.1`). | |
| ## Results | |
| All numbers are measured in this project on held-out data. The `line_items`, `manuals` and `manual_sections` | |
| sets are **fixed**, so those columns are directly comparable for every scorer. `val` is the **v2** UFGS split | |
| (11,598 units). | |
| **v2 validation split** | |
| | Scorer | val top-1 | val top-3 | val division | | |
| |---|---:|---:|---:| | |
| | **mf-0.2 (step 6132, published)** | 0.466 | 0.641 | 0.593 | | |
| | mf-0.3 (v3 data, unfrozen) | **0.556** | **0.710** | **0.662** | | |
| | mf-0.4 (v2 data, unfrozen) | 0.491 | 0.657 | 0.609 | | |
| | mf-0.1 (step 700) | 0.200 | 0.372 | 0.356 | | |
| | TF-IDF (word + char n-grams) | 0.571 | 0.732 | 0.670 | | |
| **Fixed held-out sets (line items Β· manual chunks Β· whole sections)** | |
| | Scorer | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division | | |
| |---|---:|---:|---:|---:| | |
| | **mf-0.2 (step 6132, published)** | 0.597 | 0.280 | 0.607 | **0.810** | | |
| | mf-0.3 (v3 data, unfrozen) | 0.597 | 0.259 | 0.547 | 0.765 | | |
| | mf-0.4 (v2 data, unfrozen) | **0.618** | 0.270 | 0.567 | 0.817 | | |
| | mf-0.1 (step 700) | 0.283 | 0.176 | 0.287 | 0.588 | | |
| | TF-IDF (word + char n-grams) | 0.696 | 0.287 | 0.587 | 0.752 | | |
| | **Ensemble mf-0.2 + TF-IDF (w = 0.25)** | **0.702** | **0.330** | **0.620** | 0.804 | | |
| `line_items` = 194 hand-labelled estimate line items; `manuals` = 6,890 chunks from 153 sections of two real | |
| commercial project manuals; out-of-taxonomy gold labels (e.g. `22 40 00`) count only toward the division | |
| score, so top-1 is over in-taxonomy items only. | |
| **Evidence caveat.** The line-item and section sets are small. At n = 191 the 95 % Wilson interval on the | |
| line-item top-1 is **[0.633, 0.762] (Β±6.4 pts)**, and whole-section is Β±~8 pts β so differences of a few | |
| points between models here are **noise**. Treat the `val` and `manual-chunk` rows (Β±1 pt) as the measurable | |
| ones, and the line-item/section rows as indicative. Enlarging these eval sets is the top item on the roadmap. | |
| **Reading it:** `mf-0.2` is the best scorer overall at section-level *division* accuracy (0.810). On short | |
| line items a lexical TF-IDF model is stronger (0.696 alone), so the shipped **transformer + TF-IDF ensemble** | |
| reaches **0.702** and is the production candidate; see *Ensemble*. | |
| ## Ensemble (closing the line-item gap) | |
| The transformer is trained on specification prose, so on **short, terse estimate line items** a lexical | |
| TF-IDF model is still stronger (0.696 vs 0.597 top-1). To close that gap without giving up the transformer's | |
| long-text accuracy, this repo ships a **log-probability ensemble** of `mf-0.2` and a TF-IDF + SGD scorer: | |
| ``` | |
| p = softmax( 0.25 Β· log_softmax(mf-0.2) + 0.75 Β· log_softmax(tfidf) ) | |
| ``` | |
| The 0.25 weight is chosen on the UFGS **validation** split β not on the line-item test set. `ensemble/tfidf.joblib` | |
| holds the fitted vectorizer (word 1β2 grams, 100k features, plus `char_wb` 3β5 grams, 150k features, `min_df` | |
| 2/3, sublinear tf) and SGD logistic classifier; `ensemble/blend.json` records the weight and all metrics; | |
| `ensemble/predict_ensemble.py` runs the blend. | |
| | Scorer | val top-1 | val division | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division | | |
| |---|---:|---:|---:|---:|---:|---:| | |
| | mf-0.2 (transformer) | 0.466 | 0.593 | 0.597 | 0.280 | 0.607 | **0.810** | | |
| | TF-IDF (word + char n-grams) | 0.571 | 0.670 | 0.696 | 0.287 | 0.587 | 0.752 | | |
| | **Ensemble (w = 0.25)** | **0.588** | **0.689** | **0.702** | **0.330** | **0.620** | 0.804 | | |
| ```bash | |
| pip install transformers torch scikit-learn joblib | |
| python ensemble/predict_ensemble.py --model constructelligence/masterformat-classifier \ | |
| --tfidf ensemble/tfidf.joblib --top-k 3 --divisions "4000 psi concrete slab on grade" | |
| ``` | |
| **Honest caveat.** The ensemble beats TF-IDF alone mainly on the transformer's strengths β validation top-1 | |
| (0.588 vs 0.571), manual-chunk top-1 (0.330 vs 0.287) and whole-section top-1 (0.620 vs 0.587) β while line | |
| items are nearly saturated by TF-IDF (0.702 vs 0.696). Adding bge-small embeddings does **not** help once the | |
| blend weight is chosen on `val` (its weight goes to ~0); an earlier 0.712 line-item figure came from tuning | |
| the weights on the 194-item test itself and does not generalise. The transformer remains the best scorer at | |
| section *division* accuracy (0.810). A reasonable production split is the transformer for whole sections, the | |
| ensemble for terse line items. | |
| ## Calibration and abstention | |
| The ensemble is already close to calibrated; a single **temperature** `T = 0.90` (fit on `val`) lowers the | |
| expected calibration error on `val` from **0.060 to 0.022**. More useful in production: send the | |
| lowest-confidence predictions to a human instead of filing them. | |
| Coverage β precision on the 194 line items (predictions sorted by confidence): | |
| | auto-filed (coverage) | 100 % | 90 % | 80 % | 70 % | 50 % | 30 % | | |
| |---|---:|---:|---:|---:|---:|---:| | |
| | precision | 0.69 | 0.75 | 0.81 | **0.85** | **0.91** | 0.97 | | |
| Routing the least-confident **30 %** to review lifts precision to **91 %**; routing 50 % gives **97 %**. So | |
| the model is usable today as an **assist that flags its own uncertainty**, even though it is not accurate | |
| enough to file unattended. Reproduce with `python cloud/stats.py`. | |
| ## Roadmap | |
| Further gains are **evidence-bound**, not architecture-bound (see the repo's `IMPROVEMENT-PLAN.md`): | |
| 1. **Enlarge the eval** (β₯1,000 line items, β₯400 sections) so changes are measurable at Β±3 pts β currently | |
| line-item changes below ~10 pts cannot be validated. | |
| 2. **Real line-item training data** (public bid tabulations, agency item catalogs) β the transformer's | |
| line-item ceiling is a data problem; synthetic augmentation was tested and did not transfer (`mf-0.3`). | |
| 3. **Domain-adaptive pretraining** on construction prose beyond UFGS, then retrain. | |
| 4. **Distillation + hierarchical (division β group) head** to beat the TF-IDF blend with a single model. | |
| 5. **Level-3 / full-section codes** and an explicit out-of-taxonomy fallback. | |
| ## Intended use | |
| - **Use it for:** tagging construction text with a candidate MasterFormat group (top-3 shown), an | |
| auto-classification step for estimating or spec workflows, and as a strong base to fine-tune or distil. | |
| - **Do not use it for:** unattended production takeoff, bid pricing, code compliance, or anything where a | |
| wrong cost code has financial or contractual consequences. Keep a human in the loop. | |
| - **Not a substitute for review.** Classifies text content only β if the input already contains a | |
| MasterFormat number, read the number instead. | |
| ## Training data and taxonomy | |
| - **Source:** UFGS (Unified Facilities Guide Specifications) `.SEC` files β US federal works in the | |
| **public domain**. Paragraphs and titles are parsed into labelled text units; boilerplate shared by | |
| multiple sections is dropped, cross-references are stripped so section numbers cannot leak labels, and | |
| long paragraphs are cut on sentence boundaries into 8β60-word windows. | |
| - **Split:** held out by a hash of the normalised source text (10 %), so a paragraph and its augmentations | |
| stay on the same side. **65,496 train / 11,598 val** rows (the v2 set). | |
| - **Taxonomy:** 171 level-2 groups over 32 MasterFormat divisions; group numbers follow the MasterFormat | |
| numbering convention and the short names are this project's own (see `config.json`). | |
| ## Limitations | |
| - **Below the ensemble on line items.** 59.7 % vs 70.2 % on 194 hand-labelled estimate items; do not deploy | |
| as the sole classifier. | |
| - **Domain skew.** UFGS over-represents heavy-civil, water/wastewater and process work relative to commercial | |
| building estimates; the training set is class-balanced, so raw predictions do not reflect building-project | |
| priors. | |
| - **Out-of-taxonomy inputs** (sections whose level-2 group is not among the 171) can only be scored at | |
| division level. | |
| - **Short, terse line items are the hardest inputs**; division (2-digit) accuracy is consistently higher than | |
| group (6-digit) accuracy. | |
| - **Quantized ONNX drifts.** The int8 export is smaller but less faithful than `mf-0.1`'s; prefer fp32. | |
| ## Bias, risks and safety | |
| - **Estimating bias.** Class-balanced training over a public-domain federal corpus does not represent any | |
| particular firm's cost structure or regional practice. Do not treat output as a standard or an authority. | |
| - **Trademark.** MasterFormat is a registered trademark of CSI / CSC; this model is not endorsed by them and | |
| its group names are the project's own short descriptions, not CSI's official titles. | |
| - **Privacy.** The model runs locally; no input text leaves your machine unless you call a hosted endpoint. | |
| ## FAQ | |
| **What is MasterFormat?** The CSI/CSC MasterFormat is the North American standard for organising construction | |
| specifications and cost data into numbered divisions and sections. This model predicts the **level-2 group** | |
| (a 6-digit code such as `03 30 00 Cast-in-Place Concrete`), not the full section number. | |
| **How is this different from `mf-0.1`?** `mf-0.1` was step 700 of an interrupted CPU run; `mf-0.2` is the | |
| completed 12-epoch GPU fine-tune on the rebuilt dataset. Same architecture, ~2.3Γ the validation accuracy. | |
| **Can it classify a whole specification section?** Yes β average the model's log-probabilities over a | |
| section's chunks. On two real project manuals that gives 60.7 % top-1 and 81.0 % at division level. | |
| **Can it read a MasterFormat number out of the text?** No. It classifies the description. If the number is | |
| already present, parse it directly. | |
| **Does it work offline / in the browser?** Yes β the ONNX exports are intended for `onnxruntime` and | |
| `transformers.js`. | |
| **Is a better model available?** Yes β this repo ships a **transformer + TF-IDF ensemble** (`ensemble/`) that | |
| scores **70.2 % top-1 on hand-labelled line items** (vs 59.7 % for the transformer alone) and improves | |
| manual-chunk and whole-section accuracy. Constructelligence's proprietary models are at | |
| [constructelligence.co](https://constructelligence.co). | |
| ## Files | |
| - `model.safetensors`, `config.json`, `tokenizer.json`, `tokenizer_config.json`, `vocab.txt`, | |
| `special_tokens_map.json` β standard `transformers` checkpoint. | |
| - `train_state.json` β step and validation accuracy of the saved checkpoint. | |
| - `metrics.json` β full held-out evaluation (`val`, `line_items`, `manuals`, `manual_sections`). | |
| - `predict.py` β CLI example: batching, `--top-k`, `--divisions`, `--json`, stdin/file input, `--onnx`. | |
| - `onnx/` β ONNX fp32 and int8 exports. | |
| - `ensemble/` β `tfidf.joblib` (word + char n-grams, ~112 MB), `blend.json`, `predict_ensemble.py`: the | |
| transformer + TF-IDF line-item ensemble. | |
| - `IMPROVEMENT-PLAN.md` β the prioritised roadmap (enlarged eval, data, domain-adaptive pretraining, distillation). | |
| - `CITATION.cff`, `requirements.txt`. | |
| ## Citation | |
| ```bibtex | |
| @misc{constructelligence_masterformat_classifier, | |
| title = {MasterFormat Classifier (mf-0.2): construction spec and line-item classification}, | |
| author = {Constructelligence}, | |
| year = {2026}, | |
| howpublished = {\url{https://huggingface.co/constructelligence/masterformat-classifier}}, | |
| note = {Fine-tuned from BAAI/bge-small-en-v1.5; 171-way MasterFormat level-2 classifier} | |
| } | |
| ``` | |
| ## Licence and attribution | |
| Released under the **MIT licence**, matching the base model. UFGS source text is public domain. MasterFormat | |
| is a registered trademark of CSI / CSC; this project is not affiliated with or endorsed by them. | |
| Constructelligence's production models are available at [constructelligence.co](https://constructelligence.co). | |