Text Classification
Transformers
Joblib
ONNX
Safetensors
English
bert
masterformat
masterformat-classifier
csi-masterformat
construction
construction-technology
sequence-classification
tfidf
ensemble
specs
spec-writing
specifications
takeoff
estimating
cost-code
ufgs
public-domain
Eval Results (legacy)
text-embeddings-inference
Instructions to use constructelligence/masterformat-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use constructelligence/masterformat-classifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="constructelligence/masterformat-classifier")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("constructelligence/masterformat-classifier") model = AutoModelForSequenceClassification.from_pretrained("constructelligence/masterformat-classifier", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Publish mf-0.2: completed GPU fine-tune (val 46.6%, 2.3x mf-0.1) + new model card with metrics/ONNX
Browse files- CITATION.cff +2 -2
- README.md +83 -75
- config.json +6 -2
- metrics.json +25 -0
- model.safetensors +1 -1
- onnx/model.onnx +1 -1
- onnx/model_quantized.onnx +1 -1
- predict.py +2 -2
- tokenizer.json +1 -1
- tokenizer_config.json +3 -43
- train_state.json +1 -1
CITATION.cff
CHANGED
|
@@ -1,5 +1,5 @@
|
|
| 1 |
cff-version: 1.2.0
|
| 2 |
-
title: "MasterFormat Classifier (mf-0.
|
| 3 |
message: "If you use this model, please cite it as below."
|
| 4 |
type: software
|
| 5 |
authors:
|
|
@@ -12,7 +12,7 @@ abstract: >-
|
|
| 12 |
A BERT-based text classifier that maps construction line items, specification
|
| 13 |
paragraphs and section titles to one of 171 MasterFormat level-2 groups across
|
| 14 |
32 divisions. Fine-tuned from BAAI/bge-small-en-v1.5 on public-domain UFGS
|
| 15 |
-
specification text.
|
| 16 |
keywords:
|
| 17 |
- MasterFormat
|
| 18 |
- construction
|
|
|
|
| 1 |
cff-version: 1.2.0
|
| 2 |
+
title: "MasterFormat Classifier (mf-0.2)"
|
| 3 |
message: "If you use this model, please cite it as below."
|
| 4 |
type: software
|
| 5 |
authors:
|
|
|
|
| 12 |
A BERT-based text classifier that maps construction line items, specification
|
| 13 |
paragraphs and section titles to one of 171 MasterFormat level-2 groups across
|
| 14 |
32 divisions. Fine-tuned from BAAI/bge-small-en-v1.5 on public-domain UFGS
|
| 15 |
+
specification text. mf-0.2 is the completed 12-epoch GPU fine-tune.
|
| 16 |
keywords:
|
| 17 |
- MasterFormat
|
| 18 |
- construction
|
README.md
CHANGED
|
@@ -33,41 +33,42 @@ model-index:
|
|
| 33 |
type: text-classification
|
| 34 |
name: MasterFormat level-2 group classification
|
| 35 |
dataset:
|
| 36 |
-
type: ufgs-val
|
| 37 |
-
name: UFGS held-out text units (
|
| 38 |
split: validation
|
| 39 |
metrics:
|
| 40 |
- type: accuracy
|
| 41 |
-
value: 0.
|
| 42 |
-
name: Top-1 accuracy (
|
| 43 |
widget:
|
| 44 |
-
- text:
|
| 45 |
-
example_title: Cast-in-place concrete
|
| 46 |
-
- text: TPO roofing, 60 mil, fully adhered
|
| 47 |
example_title: Membrane roofing
|
| 48 |
-
- text: Cat 6 data cabling and jacks
|
| 49 |
-
example_title: Structured cabling
|
| 50 |
- text: Wet pipe sprinkler system, light hazard
|
| 51 |
example_title: Fire suppression
|
|
|
|
|
|
|
|
|
|
|
|
|
| 52 |
---
|
| 53 |
|
| 54 |
-
# MasterFormat classifier (mf-0.
|
| 55 |
|
| 56 |
A BERT text classifier that maps **construction line items, specification paragraphs, submittals and
|
| 57 |
section titles** to one of **171 MasterFormat level-2 groups** (e.g. `03 30 00 Cast-in-Place Concrete`,
|
| 58 |
`23 30 00 HVAC Air Distribution`) across **32 divisions**. It is the machine-learning counterpart to a
|
| 59 |
-
cost-code lookup table: hand it "
|
| 60 |
-
|
| 61 |
|
| 62 |
Built for estimating, takeoff, spec-writing and RAG pipelines that need to tag free text with CSI
|
| 63 |
MasterFormat codes without a human choosing from a 171-row list.
|
| 64 |
|
| 65 |
-
> **
|
| 66 |
-
>
|
| 67 |
-
> **
|
| 68 |
-
>
|
| 69 |
-
>
|
| 70 |
-
>
|
|
|
|
| 71 |
|
| 72 |
MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or
|
| 73 |
endorsed by them.
|
|
@@ -78,19 +79,19 @@ endorsed by them.
|
|
| 78 |
from transformers import pipeline
|
| 79 |
|
| 80 |
clf = pipeline("text-classification", model="constructelligence/masterformat-classifier", top_k=3)
|
| 81 |
-
for p in clf("
|
| 82 |
print(p["label"], round(p["score"], 3))
|
| 83 |
-
#
|
| 84 |
-
#
|
| 85 |
-
#
|
| 86 |
```
|
| 87 |
|
| 88 |
Batched, with the level-1 division roll-up and JSON output:
|
| 89 |
|
| 90 |
```bash
|
| 91 |
pip install -r requirements.txt # transformers + torch
|
| 92 |
-
python predict.py --top-k 5 --divisions "
|
| 93 |
-
echo "
|
| 94 |
python predict.py --file items.txt --json > out.json
|
| 95 |
```
|
| 96 |
|
|
@@ -98,7 +99,7 @@ ONNX (no torch, ~34 MB int8):
|
|
| 98 |
|
| 99 |
```bash
|
| 100 |
pip install onnxruntime transformers
|
| 101 |
-
python predict.py --onnx onnx/model_quantized.onnx "
|
| 102 |
```
|
| 103 |
|
| 104 |
The encoder is English-only; inputs should be English.
|
|
@@ -107,71 +108,78 @@ The encoder is English-only; inputs should be English.
|
|
| 107 |
|
| 108 |
`onnx/model.onnx` (fp32, 134 MB) and `onnx/model_quantized.onnx` (dynamic int8, 34 MB) are exported with
|
| 109 |
dynamic batch and sequence axes, in the layout `transformers.js` / Optimum expect — inputs `input_ids`,
|
| 110 |
-
`attention_mask`, `token_type_ids`, output `logits`. The fp32 export matches PyTorch on a
|
| 111 |
-
(max |Δ logit|
|
| 112 |
-
|
| 113 |
-
|
| 114 |
|
| 115 |
## Model
|
| 116 |
|
| 117 |
- **Base:** [`BAAI/bge-small-en-v1.5`](https://huggingface.co/BAAI/bge-small-en-v1.5) (BERT, 33M parameters, 384-d, MIT).
|
| 118 |
- **Head:** 171-way sequence-classification layer; `id2label` / `label2id` are in [`config.json`](config.json).
|
| 119 |
-
- **Recipe:**
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
- **
|
| 123 |
-
- **A stronger v2 recipe already exists** in the source project (`label smoothing 0.05`, sentence-window
|
| 124 |
-
dataset, max length 64, lower learning rates) but no `mf-0.2` run has been completed yet.
|
| 125 |
|
| 126 |
## Results
|
| 127 |
|
| 128 |
-
All numbers are measured in this project on held-out data. `
|
| 129 |
-
|
| 130 |
-
|
| 131 |
-
|
| 132 |
|
| 133 |
-
|
| 134 |
-
|---|---:|---:|---:|---:|---:|---:|
|
| 135 |
-
| **mf-0.1 (step 700)** | **0.184** | 0.345 | 0.321 | 0.283 | 0.175 | 0.287 |
|
| 136 |
-
| TF-IDF + linear SGD (full run) | 0.531 | 0.690 | 0.634 | 0.665 | 0.290 | 0.607 |
|
| 137 |
-
| bge-small embeddings + logistic regression | 0.349 | 0.517 | 0.458 | 0.576 | 0.256 | 0.593 |
|
| 138 |
-
| Ensemble (0.25 embedding + 0.75 TF-IDF) | **0.532** | **0.700** | **0.637** | **0.702** | **0.318** | **0.640** |
|
| 139 |
|
| 140 |
-
|
| 141 |
-
|
| 142 |
-
|
|
|
|
|
|
|
| 143 |
|
| 144 |
-
|
| 145 |
-
|
| 146 |
-
|
| 147 |
-
|
| 148 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 149 |
|
| 150 |
## Intended use
|
| 151 |
|
| 152 |
-
- **Use it for:** tagging construction text with a candidate MasterFormat group (top-3 shown),
|
| 153 |
-
|
| 154 |
-
further.
|
| 155 |
- **Do not use it for:** unattended production takeoff, bid pricing, code compliance, or anything where a
|
| 156 |
-
wrong cost code has financial or contractual consequences.
|
| 157 |
-
- **Not a substitute for review.**
|
|
|
|
| 158 |
|
| 159 |
## Training data and taxonomy
|
| 160 |
|
| 161 |
- **Source:** UFGS (Unified Facilities Guide Specifications) `.SEC` files — US federal works in the
|
| 162 |
**public domain**. Paragraphs and titles are parsed into labelled text units; boilerplate shared by
|
| 163 |
multiple sections is dropped, cross-references are stripped so section numbers cannot leak labels, and
|
| 164 |
-
long paragraphs are
|
| 165 |
- **Split:** held out by a hash of the normalised source text (10 %), so a paragraph and its augmentations
|
| 166 |
-
stay on the same side.
|
| 167 |
-
65,496 / 11,598.
|
| 168 |
- **Taxonomy:** 171 level-2 groups over 32 MasterFormat divisions; group numbers follow the MasterFormat
|
| 169 |
numbering convention and the short names are this project's own (see `config.json`).
|
| 170 |
|
| 171 |
## Limitations
|
| 172 |
|
| 173 |
-
- **
|
| 174 |
-
the
|
| 175 |
- **Domain skew.** UFGS over-represents heavy-civil, water/wastewater and process work relative to commercial
|
| 176 |
building estimates; the training set is class-balanced, so raw predictions do not reflect building-project
|
| 177 |
priors.
|
|
@@ -179,8 +187,7 @@ original v1 split (13,518 units).
|
|
| 179 |
division level.
|
| 180 |
- **Short, terse line items are the hardest inputs**; division (2-digit) accuracy is consistently higher than
|
| 181 |
group (6-digit) accuracy.
|
| 182 |
-
- **
|
| 183 |
-
if your input already contains a MasterFormat number, don't ask the model — read the number.
|
| 184 |
|
| 185 |
## Bias, risks and safety
|
| 186 |
|
|
@@ -196,9 +203,11 @@ original v1 split (13,518 units).
|
|
| 196 |
specifications and cost data into numbered divisions and sections. This model predicts the **level-2 group**
|
| 197 |
(a 6-digit code such as `03 30 00 Cast-in-Place Concrete`), not the full section number.
|
| 198 |
|
| 199 |
-
**
|
| 200 |
-
|
| 201 |
-
|
|
|
|
|
|
|
| 202 |
|
| 203 |
**Can it read a MasterFormat number out of the text?** No. It classifies the description. If the number is
|
| 204 |
already present, parse it directly.
|
|
@@ -206,17 +215,16 @@ already present, parse it directly.
|
|
| 206 |
**Does it work offline / in the browser?** Yes — the ONNX exports are intended for `onnxruntime` and
|
| 207 |
`transformers.js`.
|
| 208 |
|
| 209 |
-
**Is a better model available?**
|
| 210 |
-
|
| 211 |
-
|
| 212 |
-
are at [constructelligence.co](https://constructelligence.co).
|
| 213 |
|
| 214 |
## Files
|
| 215 |
|
| 216 |
- `model.safetensors`, `config.json`, `tokenizer.json`, `tokenizer_config.json`, `vocab.txt`,
|
| 217 |
`special_tokens_map.json` — standard `transformers` checkpoint.
|
| 218 |
-
- `train_state.json` — step and validation accuracy of the saved checkpoint
|
| 219 |
-
|
| 220 |
- `predict.py` — CLI example: batching, `--top-k`, `--divisions`, `--json`, stdin/file input, `--onnx`.
|
| 221 |
- `onnx/` — ONNX fp32 and int8 exports.
|
| 222 |
- `CITATION.cff`, `requirements.txt`.
|
|
@@ -225,7 +233,7 @@ are at [constructelligence.co](https://constructelligence.co).
|
|
| 225 |
|
| 226 |
```bibtex
|
| 227 |
@misc{constructelligence_masterformat_classifier,
|
| 228 |
-
title = {MasterFormat Classifier (mf-0.
|
| 229 |
author = {Constructelligence},
|
| 230 |
year = {2026},
|
| 231 |
howpublished = {\url{https://huggingface.co/constructelligence/masterformat-classifier}},
|
|
|
|
| 33 |
type: text-classification
|
| 34 |
name: MasterFormat level-2 group classification
|
| 35 |
dataset:
|
| 36 |
+
type: ufgs-val-v2
|
| 37 |
+
name: UFGS held-out text units (11,598)
|
| 38 |
split: validation
|
| 39 |
metrics:
|
| 40 |
- type: accuracy
|
| 41 |
+
value: 0.466
|
| 42 |
+
name: Top-1 accuracy (mf-0.2)
|
| 43 |
widget:
|
| 44 |
+
- text: EPDM membrane roofing
|
|
|
|
|
|
|
| 45 |
example_title: Membrane roofing
|
|
|
|
|
|
|
| 46 |
- text: Wet pipe sprinkler system, light hazard
|
| 47 |
example_title: Fire suppression
|
| 48 |
+
- text: 8" CMU wall, grout filled at 32" o.c.
|
| 49 |
+
example_title: Unit masonry
|
| 50 |
+
- text: Addressable fire alarm system, devices and panel
|
| 51 |
+
example_title: Fire alarm
|
| 52 |
---
|
| 53 |
|
| 54 |
+
# MasterFormat classifier (mf-0.2) — construction spec & line-item classification
|
| 55 |
|
| 56 |
A BERT text classifier that maps **construction line items, specification paragraphs, submittals and
|
| 57 |
section titles** to one of **171 MasterFormat level-2 groups** (e.g. `03 30 00 Cast-in-Place Concrete`,
|
| 58 |
`23 30 00 HVAC Air Distribution`) across **32 divisions**. It is the machine-learning counterpart to a
|
| 59 |
+
cost-code lookup table: hand it "EPDM membrane roofing" and it returns the MasterFormat group, ranked.
|
| 60 |
+
English text, one label per input.
|
| 61 |
|
| 62 |
Built for estimating, takeoff, spec-writing and RAG pipelines that need to tag free text with CSI
|
| 63 |
MasterFormat codes without a human choosing from a 171-row list.
|
| 64 |
|
| 65 |
+
> **Completed GPU fine-tune.** `mf-0.2` is the full run that `mf-0.1` (step 700) was an early checkpoint of:
|
| 66 |
+
> **6,132 steps / 12 epochs** on a Kaggle T4 over the rebuilt v2 dataset. Top-1 accuracy on the v2 UFGS
|
| 67 |
+
> validation split is **46.6 %** — **2.3×** `mf-0.1` (20.0 %). Line-item accuracy on 194 hand-labelled
|
| 68 |
+
> estimate items is **59.7 %** (was 28.3 %), and whole-section accuracy is **60.7 %** with **81.0 %** at
|
| 69 |
+
> division level. It is still **below this project's TF-IDF + embedding ensemble** on line items (70.2 %) —
|
| 70 |
+
> see *Results* — but it now matches TF-IDF on section-level top-1 and beats every baseline on section-level
|
| 71 |
+
> division accuracy.
|
| 72 |
|
| 73 |
MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or
|
| 74 |
endorsed by them.
|
|
|
|
| 79 |
from transformers import pipeline
|
| 80 |
|
| 81 |
clf = pipeline("text-classification", model="constructelligence/masterformat-classifier", top_k=3)
|
| 82 |
+
for p in clf("EPDM membrane roofing"):
|
| 83 |
print(p["label"], round(p["score"], 3))
|
| 84 |
+
# 07 50 00 Membrane Roofing 0.863
|
| 85 |
+
# 07 10 00 Dampproofing and Waterproofing 0.068
|
| 86 |
+
# 07 30 00 Steep Slope Roofing 0.047
|
| 87 |
```
|
| 88 |
|
| 89 |
Batched, with the level-1 division roll-up and JSON output:
|
| 90 |
|
| 91 |
```bash
|
| 92 |
pip install -r requirements.txt # transformers + torch
|
| 93 |
+
python predict.py --top-k 5 --divisions "Addressable fire alarm system, devices and panel"
|
| 94 |
+
echo "12\" RCP storm drain pipe" | python predict.py -
|
| 95 |
python predict.py --file items.txt --json > out.json
|
| 96 |
```
|
| 97 |
|
|
|
|
| 99 |
|
| 100 |
```bash
|
| 101 |
pip install onnxruntime transformers
|
| 102 |
+
python predict.py --onnx onnx/model_quantized.onnx "Wet pipe sprinkler system, light hazard"
|
| 103 |
```
|
| 104 |
|
| 105 |
The encoder is English-only; inputs should be English.
|
|
|
|
| 108 |
|
| 109 |
`onnx/model.onnx` (fp32, 134 MB) and `onnx/model_quantized.onnx` (dynamic int8, 34 MB) are exported with
|
| 110 |
dynamic batch and sequence axes, in the layout `transformers.js` / Optimum expect — inputs `input_ids`,
|
| 111 |
+
`attention_mask`, `token_type_ids`, output `logits`. The fp32 export matches PyTorch on a 4-text smoke set
|
| 112 |
+
(max |Δ logit| 1e-5, argmax agreement 4/4). Dynamic int8 quantization shifts the logits more than it did for
|
| 113 |
+
`mf-0.1` — the better-trained head is more confident — so verify the quantized model on your own inputs;
|
| 114 |
+
the fp32 export is the safer default. Reproduce with `python scripts/export_onnx.py runs/mf-0.2`.
|
| 115 |
|
| 116 |
## Model
|
| 117 |
|
| 118 |
- **Base:** [`BAAI/bge-small-en-v1.5`](https://huggingface.co/BAAI/bge-small-en-v1.5) (BERT, 33M parameters, 384-d, MIT).
|
| 119 |
- **Head:** 171-way sequence-classification layer; `id2label` / `label2id` are in [`config.json`](config.json).
|
| 120 |
+
- **Recipe:** `label smoothing 0.05`, max length 64, batch 128, 12 epochs (6,132 steps), fp16 autocast,
|
| 121 |
+
AdamW with 6 % linear warmup; body LR 5e-5, head LR 1e-3; embeddings and the first 4 encoder layers frozen.
|
| 122 |
+
- **Checkpoint:** `train_state.json` records `{"step": 6132, "val_acc": 0.4661}`.
|
| 123 |
+
- **Trained on:** Kaggle T4, ~16.5 min wall-clock, from `BAAI/bge-small-en-v1.5` (not resumed from `mf-0.1`).
|
|
|
|
|
|
|
| 124 |
|
| 125 |
## Results
|
| 126 |
|
| 127 |
+
All numbers are measured in this project on held-out data. The `line_items`, `manuals` and `manual_sections`
|
| 128 |
+
sets are **fixed**, so those columns are directly comparable for every scorer. `val` depends on the dataset
|
| 129 |
+
version: `mf-0.2` and `mf-0.1` below are on the **v2** split (11,598 units); the embeddings/ensemble rows
|
| 130 |
+
were tuned on the v1 split (13,518 units) and are marked *v1*.
|
| 131 |
|
| 132 |
+
**v2 validation split**
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 133 |
|
| 134 |
+
| Scorer | val top-1 | val top-3 | val division |
|
| 135 |
+
|---|---:|---:|---:|
|
| 136 |
+
| **mf-0.2 (step 6132)** | **0.466** | **0.641** | **0.593** |
|
| 137 |
+
| mf-0.1 (step 700) | 0.200 | 0.372 | 0.356 |
|
| 138 |
+
| TF-IDF + linear SGD | 0.575 | 0.732 | 0.675 |
|
| 139 |
|
| 140 |
+
**Fixed held-out sets (line items · manual chunks · whole sections)**
|
| 141 |
+
|
| 142 |
+
| Scorer | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division |
|
| 143 |
+
|---|---:|---:|---:|---:|
|
| 144 |
+
| **mf-0.2 (step 6132)** | **0.597** | 0.280 | **0.607** | **0.810** |
|
| 145 |
+
| mf-0.1 (step 700) | 0.283 | 0.176 | 0.287 | 0.588 |
|
| 146 |
+
| TF-IDF + linear SGD | 0.665 | **0.290** | **0.607** | 0.732 |
|
| 147 |
+
| bge-small embeddings + logistic regression *(v1)* | 0.576 | 0.256 | 0.593 | 0.725 |
|
| 148 |
+
| Ensemble 0.25 embedding + 0.75 TF-IDF *(v1)* | **0.702** | **0.318** | 0.640 | 0.771 |
|
| 149 |
+
|
| 150 |
+
`line_items` = 194 hand-labelled estimate line items; `manuals` = 6,890 chunks from 153 sections of two real
|
| 151 |
+
commercial project manuals; out-of-taxonomy gold labels (e.g. `22 40 00`) count only toward the division
|
| 152 |
+
score, so top-1 is over in-taxonomy items only.
|
| 153 |
+
|
| 154 |
+
**Reading it:** `mf-0.2` is the best *single transformer* here and the best scorer overall at section-level
|
| 155 |
+
division accuracy (0.810). The TF-IDF + embedding **ensemble still leads on short line items** (0.702 vs
|
| 156 |
+
0.597), which is the real-world target — so the ensemble remains the production candidate, with `mf-0.2` a
|
| 157 |
+
much stronger transformer baseline than `mf-0.1`.
|
| 158 |
|
| 159 |
## Intended use
|
| 160 |
|
| 161 |
+
- **Use it for:** tagging construction text with a candidate MasterFormat group (top-3 shown), an
|
| 162 |
+
auto-classification step for estimating or spec workflows, and as a strong base to fine-tune or distil.
|
|
|
|
| 163 |
- **Do not use it for:** unattended production takeoff, bid pricing, code compliance, or anything where a
|
| 164 |
+
wrong cost code has financial or contractual consequences. Keep a human in the loop.
|
| 165 |
+
- **Not a substitute for review.** Classifies text content only — if the input already contains a
|
| 166 |
+
MasterFormat number, read the number instead.
|
| 167 |
|
| 168 |
## Training data and taxonomy
|
| 169 |
|
| 170 |
- **Source:** UFGS (Unified Facilities Guide Specifications) `.SEC` files — US federal works in the
|
| 171 |
**public domain**. Paragraphs and titles are parsed into labelled text units; boilerplate shared by
|
| 172 |
multiple sections is dropped, cross-references are stripped so section numbers cannot leak labels, and
|
| 173 |
+
long paragraphs are cut on sentence boundaries into 8–60-word windows.
|
| 174 |
- **Split:** held out by a hash of the normalised source text (10 %), so a paragraph and its augmentations
|
| 175 |
+
stay on the same side. **65,496 train / 11,598 val** rows (the v2 set).
|
|
|
|
| 176 |
- **Taxonomy:** 171 level-2 groups over 32 MasterFormat divisions; group numbers follow the MasterFormat
|
| 177 |
numbering convention and the short names are this project's own (see `config.json`).
|
| 178 |
|
| 179 |
## Limitations
|
| 180 |
|
| 181 |
+
- **Below the ensemble on line items.** 59.7 % vs 70.2 % on 194 hand-labelled estimate items; do not deploy
|
| 182 |
+
as the sole classifier.
|
| 183 |
- **Domain skew.** UFGS over-represents heavy-civil, water/wastewater and process work relative to commercial
|
| 184 |
building estimates; the training set is class-balanced, so raw predictions do not reflect building-project
|
| 185 |
priors.
|
|
|
|
| 187 |
division level.
|
| 188 |
- **Short, terse line items are the hardest inputs**; division (2-digit) accuracy is consistently higher than
|
| 189 |
group (6-digit) accuracy.
|
| 190 |
+
- **Quantized ONNX drifts.** The int8 export is smaller but less faithful than `mf-0.1`'s; prefer fp32.
|
|
|
|
| 191 |
|
| 192 |
## Bias, risks and safety
|
| 193 |
|
|
|
|
| 203 |
specifications and cost data into numbered divisions and sections. This model predicts the **level-2 group**
|
| 204 |
(a 6-digit code such as `03 30 00 Cast-in-Place Concrete`), not the full section number.
|
| 205 |
|
| 206 |
+
**How is this different from `mf-0.1`?** `mf-0.1` was step 700 of an interrupted CPU run; `mf-0.2` is the
|
| 207 |
+
completed 12-epoch GPU fine-tune on the rebuilt dataset. Same architecture, ~2.3× the validation accuracy.
|
| 208 |
+
|
| 209 |
+
**Can it classify a whole specification section?** Yes — average the model's log-probabilities over a
|
| 210 |
+
section's chunks. On two real project manuals that gives 60.7 % top-1 and 81.0 % at division level.
|
| 211 |
|
| 212 |
**Can it read a MasterFormat number out of the text?** No. It classifies the description. If the number is
|
| 213 |
already present, parse it directly.
|
|
|
|
| 215 |
**Does it work offline / in the browser?** Yes — the ONNX exports are intended for `onnxruntime` and
|
| 216 |
`transformers.js`.
|
| 217 |
|
| 218 |
+
**Is a better model available?** The project's TF-IDF + embedding **ensemble scores 70.2 % top-1 on
|
| 219 |
+
hand-labelled line items** and remains the production candidate. Constructelligence's proprietary models are
|
| 220 |
+
at [constructelligence.co](https://constructelligence.co).
|
|
|
|
| 221 |
|
| 222 |
## Files
|
| 223 |
|
| 224 |
- `model.safetensors`, `config.json`, `tokenizer.json`, `tokenizer_config.json`, `vocab.txt`,
|
| 225 |
`special_tokens_map.json` — standard `transformers` checkpoint.
|
| 226 |
+
- `train_state.json` — step and validation accuracy of the saved checkpoint.
|
| 227 |
+
- `metrics.json` — full held-out evaluation (`val`, `line_items`, `manuals`, `manual_sections`).
|
| 228 |
- `predict.py` — CLI example: batching, `--top-k`, `--divisions`, `--json`, stdin/file input, `--onnx`.
|
| 229 |
- `onnx/` — ONNX fp32 and int8 exports.
|
| 230 |
- `CITATION.cff`, `requirements.txt`.
|
|
|
|
| 233 |
|
| 234 |
```bibtex
|
| 235 |
@misc{constructelligence_masterformat_classifier,
|
| 236 |
+
title = {MasterFormat Classifier (mf-0.2): construction spec and line-item classification},
|
| 237 |
author = {Constructelligence},
|
| 238 |
year = {2026},
|
| 239 |
howpublished = {\url{https://huggingface.co/constructelligence/masterformat-classifier}},
|
config.json
CHANGED
|
@@ -1,10 +1,13 @@
|
|
| 1 |
{
|
|
|
|
| 2 |
"architectures": [
|
| 3 |
"BertForSequenceClassification"
|
| 4 |
],
|
| 5 |
"attention_probs_dropout_prob": 0.1,
|
|
|
|
| 6 |
"classifier_dropout": null,
|
| 7 |
"dtype": "float32",
|
|
|
|
| 8 |
"hidden_act": "gelu",
|
| 9 |
"hidden_dropout_prob": 0.1,
|
| 10 |
"hidden_size": 384,
|
|
@@ -183,6 +186,7 @@
|
|
| 183 |
},
|
| 184 |
"initializer_range": 0.02,
|
| 185 |
"intermediate_size": 1536,
|
|
|
|
| 186 |
"label2id": {
|
| 187 |
"01 10 00 Summary of Work": 0,
|
| 188 |
"01 20 00 Price and Payment Procedures": 1,
|
|
@@ -363,8 +367,8 @@
|
|
| 363 |
"num_hidden_layers": 12,
|
| 364 |
"pad_token_id": 0,
|
| 365 |
"position_embedding_type": "absolute",
|
| 366 |
-
"
|
| 367 |
-
"transformers_version": "
|
| 368 |
"type_vocab_size": 2,
|
| 369 |
"use_cache": true,
|
| 370 |
"vocab_size": 30522
|
|
|
|
| 1 |
{
|
| 2 |
+
"add_cross_attention": false,
|
| 3 |
"architectures": [
|
| 4 |
"BertForSequenceClassification"
|
| 5 |
],
|
| 6 |
"attention_probs_dropout_prob": 0.1,
|
| 7 |
+
"bos_token_id": null,
|
| 8 |
"classifier_dropout": null,
|
| 9 |
"dtype": "float32",
|
| 10 |
+
"eos_token_id": null,
|
| 11 |
"hidden_act": "gelu",
|
| 12 |
"hidden_dropout_prob": 0.1,
|
| 13 |
"hidden_size": 384,
|
|
|
|
| 186 |
},
|
| 187 |
"initializer_range": 0.02,
|
| 188 |
"intermediate_size": 1536,
|
| 189 |
+
"is_decoder": false,
|
| 190 |
"label2id": {
|
| 191 |
"01 10 00 Summary of Work": 0,
|
| 192 |
"01 20 00 Price and Payment Procedures": 1,
|
|
|
|
| 367 |
"num_hidden_layers": 12,
|
| 368 |
"pad_token_id": 0,
|
| 369 |
"position_embedding_type": "absolute",
|
| 370 |
+
"tie_word_embeddings": true,
|
| 371 |
+
"transformers_version": "5.16.1",
|
| 372 |
"type_vocab_size": 2,
|
| 373 |
"use_cache": true,
|
| 374 |
"vocab_size": 30522
|
metrics.json
ADDED
|
@@ -0,0 +1,25 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"val": {
|
| 3 |
+
"n": 11598,
|
| 4 |
+
"top1": 0.4661148473874806,
|
| 5 |
+
"top3": 0.6408001379548198,
|
| 6 |
+
"division": 0.5929470598379031
|
| 7 |
+
},
|
| 8 |
+
"line_items": {
|
| 9 |
+
"n": 194,
|
| 10 |
+
"top1": 0.5968586387434555,
|
| 11 |
+
"top3": 0.7591623036649214,
|
| 12 |
+
"division": 0.7371134020618557
|
| 13 |
+
},
|
| 14 |
+
"manuals": {
|
| 15 |
+
"n": 6890,
|
| 16 |
+
"top1": 0.2797942689199118,
|
| 17 |
+
"top3": 0.44584864070536373,
|
| 18 |
+
"division": 0.46357039187227866
|
| 19 |
+
},
|
| 20 |
+
"manual_sections": {
|
| 21 |
+
"sections": 153,
|
| 22 |
+
"top1": 0.6066666666666667,
|
| 23 |
+
"division": 0.8104575163398693
|
| 24 |
+
}
|
| 25 |
+
}
|
model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 133726636
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:47d0e681abcfa7f5c588035bd451cc963e335d46bc1de91ef505cccc025a0af9
|
| 3 |
size 133726636
|
onnx/model.onnx
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 133960852
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:10628d545f4b23125421a30ebd11678d7f20911d52c0a7d2e4c893f52c916827
|
| 3 |
size 133960852
|
onnx/model_quantized.onnx
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 34067631
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:18a2d9bc236170e7ee6a06f27f7e1a73073d651d6e85dc2f0a8eb998dcd30055
|
| 3 |
size 34067631
|
predict.py
CHANGED
|
@@ -8,7 +8,7 @@ Examples
|
|
| 8 |
echo "Cat 6 data cabling and jacks" | python predict.py -
|
| 9 |
python predict.py --file items.txt --json > out.json
|
| 10 |
|
| 11 |
-
ONNX (no torch, ~
|
| 12 |
pip install onnxruntime transformers
|
| 13 |
python predict.py --onnx onnx/model_quantized.onnx "Panelboards 208/120V 42 circuit"
|
| 14 |
|
|
@@ -116,7 +116,7 @@ def predict(texts, score, top_k, want_div):
|
|
| 116 |
|
| 117 |
|
| 118 |
def main(argv=None):
|
| 119 |
-
ap = argparse.ArgumentParser(description="MasterFormat level-2 classifier (
|
| 120 |
ap.add_argument("text", nargs="*", help="text to classify; use '-' to read lines from stdin")
|
| 121 |
ap.add_argument("--model", default=MODEL, help="HF repo id or local checkpoint dir")
|
| 122 |
ap.add_argument("--onnx", metavar="PATH", help="classify with an ONNX model instead of PyTorch")
|
|
|
|
| 8 |
echo "Cat 6 data cabling and jacks" | python predict.py -
|
| 9 |
python predict.py --file items.txt --json > out.json
|
| 10 |
|
| 11 |
+
ONNX (no torch, ~34 MB int8 weights):
|
| 12 |
pip install onnxruntime transformers
|
| 13 |
python predict.py --onnx onnx/model_quantized.onnx "Panelboards 208/120V 42 circuit"
|
| 14 |
|
|
|
|
| 116 |
|
| 117 |
|
| 118 |
def main(argv=None):
|
| 119 |
+
ap = argparse.ArgumentParser(description="MasterFormat level-2 classifier (171 groups, 32 divisions).")
|
| 120 |
ap.add_argument("text", nargs="*", help="text to classify; use '-' to read lines from stdin")
|
| 121 |
ap.add_argument("--model", default=MODEL, help="HF repo id or local checkpoint dir")
|
| 122 |
ap.add_argument("--onnx", metavar="PATH", help="classify with an ONNX model instead of PyTorch")
|
tokenizer.json
CHANGED
|
@@ -2,7 +2,7 @@
|
|
| 2 |
"version": "1.0",
|
| 3 |
"truncation": {
|
| 4 |
"direction": "Right",
|
| 5 |
-
"max_length":
|
| 6 |
"strategy": "LongestFirst",
|
| 7 |
"stride": 0
|
| 8 |
},
|
|
|
|
| 2 |
"version": "1.0",
|
| 3 |
"truncation": {
|
| 4 |
"direction": "Right",
|
| 5 |
+
"max_length": 128,
|
| 6 |
"strategy": "LongestFirst",
|
| 7 |
"stride": 0
|
| 8 |
},
|
tokenizer_config.json
CHANGED
|
@@ -1,51 +1,11 @@
|
|
| 1 |
{
|
| 2 |
-
"
|
| 3 |
-
"0": {
|
| 4 |
-
"content": "[PAD]",
|
| 5 |
-
"lstrip": false,
|
| 6 |
-
"normalized": false,
|
| 7 |
-
"rstrip": false,
|
| 8 |
-
"single_word": false,
|
| 9 |
-
"special": true
|
| 10 |
-
},
|
| 11 |
-
"100": {
|
| 12 |
-
"content": "[UNK]",
|
| 13 |
-
"lstrip": false,
|
| 14 |
-
"normalized": false,
|
| 15 |
-
"rstrip": false,
|
| 16 |
-
"single_word": false,
|
| 17 |
-
"special": true
|
| 18 |
-
},
|
| 19 |
-
"101": {
|
| 20 |
-
"content": "[CLS]",
|
| 21 |
-
"lstrip": false,
|
| 22 |
-
"normalized": false,
|
| 23 |
-
"rstrip": false,
|
| 24 |
-
"single_word": false,
|
| 25 |
-
"special": true
|
| 26 |
-
},
|
| 27 |
-
"102": {
|
| 28 |
-
"content": "[SEP]",
|
| 29 |
-
"lstrip": false,
|
| 30 |
-
"normalized": false,
|
| 31 |
-
"rstrip": false,
|
| 32 |
-
"single_word": false,
|
| 33 |
-
"special": true
|
| 34 |
-
},
|
| 35 |
-
"103": {
|
| 36 |
-
"content": "[MASK]",
|
| 37 |
-
"lstrip": false,
|
| 38 |
-
"normalized": false,
|
| 39 |
-
"rstrip": false,
|
| 40 |
-
"single_word": false,
|
| 41 |
-
"special": true
|
| 42 |
-
}
|
| 43 |
-
},
|
| 44 |
"clean_up_tokenization_spaces": true,
|
| 45 |
"cls_token": "[CLS]",
|
| 46 |
"do_basic_tokenize": true,
|
| 47 |
"do_lower_case": true,
|
| 48 |
-
"
|
|
|
|
| 49 |
"mask_token": "[MASK]",
|
| 50 |
"model_max_length": 512,
|
| 51 |
"never_split": null,
|
|
|
|
| 1 |
{
|
| 2 |
+
"backend": "tokenizers",
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
"clean_up_tokenization_spaces": true,
|
| 4 |
"cls_token": "[CLS]",
|
| 5 |
"do_basic_tokenize": true,
|
| 6 |
"do_lower_case": true,
|
| 7 |
+
"is_local": false,
|
| 8 |
+
"local_files_only": false,
|
| 9 |
"mask_token": "[MASK]",
|
| 10 |
"model_max_length": 512,
|
| 11 |
"never_split": null,
|
train_state.json
CHANGED
|
@@ -1 +1 @@
|
|
| 1 |
-
{"step":
|
|
|
|
| 1 |
+
{"step": 6132, "val_acc": 0.4661148473874806}
|