Text Classification
Transformers
Joblib
ONNX
Safetensors
English
bert
masterformat
masterformat-classifier
csi-masterformat
construction
construction-technology
sequence-classification
tfidf
ensemble
specs
spec-writing
specifications
takeoff
estimating
cost-code
ufgs
public-domain
Eval Results (legacy)
text-embeddings-inference
Instructions to use constructelligence/masterformat-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use constructelligence/masterformat-classifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="constructelligence/masterformat-classifier")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("constructelligence/masterformat-classifier") model = AutoModelForSequenceClassification.from_pretrained("constructelligence/masterformat-classifier", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Improve model card (SEO title, FAQ, citation, widgets) and add predict.py CLI, CITATION.cff, requirements.txt
Browse files- CITATION.cff +24 -0
- README.md +175 -67
- predict.py +143 -10
- requirements.txt +7 -0
CITATION.cff
ADDED
|
@@ -0,0 +1,24 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
cff-version: 1.2.0
|
| 2 |
+
title: "MasterFormat Classifier (mf-0.1)"
|
| 3 |
+
message: "If you use this model, please cite it as below."
|
| 4 |
+
type: software
|
| 5 |
+
authors:
|
| 6 |
+
- name: "Constructelligence"
|
| 7 |
+
website: "https://constructelligence.co"
|
| 8 |
+
license: MIT
|
| 9 |
+
url: "https://huggingface.co/constructelligence/masterformat-classifier"
|
| 10 |
+
repository-code: "https://huggingface.co/constructelligence/masterformat-classifier"
|
| 11 |
+
abstract: >-
|
| 12 |
+
A BERT-based text classifier that maps construction line items, specification
|
| 13 |
+
paragraphs and section titles to one of 171 MasterFormat level-2 groups across
|
| 14 |
+
32 divisions. Fine-tuned from BAAI/bge-small-en-v1.5 on public-domain UFGS
|
| 15 |
+
specification text. Published as an early (step-700) checkpoint.
|
| 16 |
+
keywords:
|
| 17 |
+
- MasterFormat
|
| 18 |
+
- construction
|
| 19 |
+
- text-classification
|
| 20 |
+
- specifications
|
| 21 |
+
- estimating
|
| 22 |
+
- takeoff
|
| 23 |
+
- UFGS
|
| 24 |
+
- BERT
|
README.md
CHANGED
|
@@ -1,17 +1,29 @@
|
|
| 1 |
---
|
|
|
|
|
|
|
| 2 |
license: mit
|
| 3 |
-
base_model: BAAI/bge-small-en-v1.5
|
| 4 |
library_name: transformers
|
| 5 |
pipeline_tag: text-classification
|
| 6 |
-
|
| 7 |
tags:
|
| 8 |
-
- construction
|
| 9 |
- masterformat
|
|
|
|
|
|
|
|
|
|
|
|
|
| 10 |
- text-classification
|
| 11 |
-
-
|
| 12 |
-
- takeoff
|
| 13 |
- bert
|
| 14 |
- onnx
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 15 |
metrics:
|
| 16 |
- accuracy
|
| 17 |
model-index:
|
|
@@ -19,7 +31,7 @@ model-index:
|
|
| 19 |
results:
|
| 20 |
- task:
|
| 21 |
type: text-classification
|
| 22 |
-
name: MasterFormat level-2 classification
|
| 23 |
dataset:
|
| 24 |
type: ufgs-val
|
| 25 |
name: UFGS held-out text units (13,518)
|
|
@@ -28,43 +40,37 @@ model-index:
|
|
| 28 |
- type: accuracy
|
| 29 |
value: 0.184
|
| 30 |
name: Top-1 accuracy (step-700 checkpoint)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 31 |
---
|
| 32 |
|
| 33 |
-
# MasterFormat
|
| 34 |
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
|
| 41 |
-
|
| 42 |
-
MasterFormat
|
| 43 |
-
across 32 divisions. English text, one label per input.
|
| 44 |
|
| 45 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 46 |
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
- Base: [`BAAI/bge-small-en-v1.5`](https://huggingface.co/BAAI/bge-small-en-v1.5) (BERT, 33M parameters, 384-d), MIT.
|
| 50 |
-
- Fine-tuned with a 171-way sequence-classification head; `id2label` / `label2id` are in [`config.json`](config.json).
|
| 51 |
-
- Training: 2 epochs planned over 66,396 rows (per-class balance capped at 450 / floored at 300, estimate-style
|
| 52 |
-
augmentation), batch 32, max length 96, AdamW, 6% linear warmup; body LR 6e-5, head LR 2e-3; embeddings and
|
| 53 |
-
the first 4 encoder layers frozen.
|
| 54 |
-
- Checkpoint stopped at step 700; `train_state.json` records `{"step": 700, "val_acc": 0.18725}`.
|
| 55 |
-
|
| 56 |
-
## Results
|
| 57 |
-
|
| 58 |
-
All numbers are measured on held-out data from the same project. `val` = UFGS text units,
|
| 59 |
-
`line_items` = 194 hand-labelled estimate line items, `manuals` = 6,890 chunks from 153 sections of two real
|
| 60 |
-
commercial project manuals (out-of-taxonomy gold labels count only toward the division score).
|
| 61 |
-
|
| 62 |
-
| Scorer | val top-1 | val top-3 | val division | line-item top-1 | manual-chunk top-1 | manual-section top-1 |
|
| 63 |
-
|---|---:|---:|---:|---:|---:|---:|
|
| 64 |
-
| **mf-0.1 (step 700)** | 0.184 | 0.345 | 0.321 | 0.283 | 0.175 | 0.287 |
|
| 65 |
-
| TF-IDF + linear SGD (full run) | 0.531 | 0.690 | 0.634 | 0.665 | 0.290 | 0.607 |
|
| 66 |
-
| bge-small embeddings + logistic regression | 0.349 | 0.517 | 0.458 | 0.576 | 0.256 | 0.593 |
|
| 67 |
-
| Ensemble (0.25 embedding + 0.75 TF-IDF) | 0.532 | 0.700 | 0.637 | 0.702 | 0.318 | 0.640 |
|
| 68 |
|
| 69 |
## Quick start
|
| 70 |
|
|
@@ -74,59 +80,161 @@ from transformers import pipeline
|
|
| 74 |
clf = pipeline("text-classification", model="constructelligence/masterformat-classifier", top_k=3)
|
| 75 |
for p in clf("4000 psi concrete slab on grade"):
|
| 76 |
print(p["label"], round(p["score"], 3))
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 77 |
```
|
| 78 |
|
| 79 |
-
|
| 80 |
|
| 81 |
```bash
|
| 82 |
-
pip install
|
| 83 |
-
python predict.py
|
| 84 |
```
|
| 85 |
|
| 86 |
The encoder is English-only; inputs should be English.
|
| 87 |
|
| 88 |
-
### ONNX
|
| 89 |
|
| 90 |
`onnx/model.onnx` (fp32, 134 MB) and `onnx/model_quantized.onnx` (dynamic int8, 34 MB) are exported with
|
| 91 |
-
dynamic batch and sequence axes, in the layout `transformers.js` / Optimum expect
|
| 92 |
-
`attention_mask`, `token_type_ids`, output `logits`. The fp32 export matches PyTorch on a
|
| 93 |
-
(max |
|
| 94 |
-
(max |
|
| 95 |
-
inputs before
|
| 96 |
|
| 97 |
-
##
|
| 98 |
|
| 99 |
-
-
|
| 100 |
-
|
| 101 |
-
-
|
| 102 |
-
|
| 103 |
-
|
| 104 |
-
- `
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 105 |
|
| 106 |
## Training data and taxonomy
|
| 107 |
|
| 108 |
-
- **
|
| 109 |
-
Paragraphs and titles are parsed into labelled text units; boilerplate shared by
|
| 110 |
-
cross-references are stripped so section numbers cannot leak labels, and
|
| 111 |
-
20–60-word units.
|
| 112 |
-
- **Split:** held out by a hash of the normalised source text (10%), so a paragraph and its augmentations
|
| 113 |
-
on the same side. 66,396 train / 13,518 val rows
|
|
|
|
| 114 |
- **Taxonomy:** 171 level-2 groups over 32 MasterFormat divisions; group numbers follow the MasterFormat
|
| 115 |
-
numbering convention and the short names are
|
| 116 |
|
| 117 |
## Limitations
|
| 118 |
|
| 119 |
-
- **Early checkpoint.** The 18.4% top-1 is a snapshot of an interrupted run; it is not competitive with
|
| 120 |
-
TF-IDF baseline yet. Do not use it for production takeoff without further training and re-evaluation.
|
| 121 |
- **Domain skew.** UFGS over-represents heavy-civil, water/wastewater and process work relative to commercial
|
| 122 |
building estimates; the training set is class-balanced, so raw predictions do not reflect building-project
|
| 123 |
priors.
|
| 124 |
-
- **Out-of-taxonomy inputs** (sections whose level-2 group is not
|
| 125 |
-
level.
|
| 126 |
-
- Short, terse line items are the hardest inputs; division (2-digit) accuracy is higher than
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 127 |
|
| 128 |
## Licence and attribution
|
| 129 |
|
| 130 |
Released under the **MIT licence**, matching the base model. UFGS source text is public domain. MasterFormat
|
| 131 |
-
is a registered trademark of CSI / CSC; this project is not affiliated with or endorsed by them.
|
| 132 |
-
production models are available at [constructelligence.co](https://constructelligence.co).
|
|
|
|
| 1 |
---
|
| 2 |
+
language:
|
| 3 |
+
- en
|
| 4 |
license: mit
|
|
|
|
| 5 |
library_name: transformers
|
| 6 |
pipeline_tag: text-classification
|
| 7 |
+
base_model: BAAI/bge-small-en-v1.5
|
| 8 |
tags:
|
|
|
|
| 9 |
- masterformat
|
| 10 |
+
- masterformat-classifier
|
| 11 |
+
- csi-masterformat
|
| 12 |
+
- construction
|
| 13 |
+
- construction-technology
|
| 14 |
- text-classification
|
| 15 |
+
- sequence-classification
|
|
|
|
| 16 |
- bert
|
| 17 |
- onnx
|
| 18 |
+
- safetensors
|
| 19 |
+
- specs
|
| 20 |
+
- spec-writing
|
| 21 |
+
- specifications
|
| 22 |
+
- takeoff
|
| 23 |
+
- estimating
|
| 24 |
+
- cost-code
|
| 25 |
+
- ufgs
|
| 26 |
+
- public-domain
|
| 27 |
metrics:
|
| 28 |
- accuracy
|
| 29 |
model-index:
|
|
|
|
| 31 |
results:
|
| 32 |
- task:
|
| 33 |
type: text-classification
|
| 34 |
+
name: MasterFormat level-2 group classification
|
| 35 |
dataset:
|
| 36 |
type: ufgs-val
|
| 37 |
name: UFGS held-out text units (13,518)
|
|
|
|
| 40 |
- type: accuracy
|
| 41 |
value: 0.184
|
| 42 |
name: Top-1 accuracy (step-700 checkpoint)
|
| 43 |
+
widget:
|
| 44 |
+
- text: 4000 psi concrete slab on grade
|
| 45 |
+
example_title: Cast-in-place concrete
|
| 46 |
+
- text: TPO roofing, 60 mil, fully adhered
|
| 47 |
+
example_title: Membrane roofing
|
| 48 |
+
- text: Cat 6 data cabling and jacks
|
| 49 |
+
example_title: Structured cabling
|
| 50 |
+
- text: Wet pipe sprinkler system, light hazard
|
| 51 |
+
example_title: Fire suppression
|
| 52 |
---
|
| 53 |
|
| 54 |
+
# MasterFormat classifier (mf-0.1) — construction spec & line-item classification
|
| 55 |
|
| 56 |
+
A BERT text classifier that maps **construction line items, specification paragraphs, submittals and
|
| 57 |
+
section titles** to one of **171 MasterFormat level-2 groups** (e.g. `03 30 00 Cast-in-Place Concrete`,
|
| 58 |
+
`23 30 00 HVAC Air Distribution`) across **32 divisions**. It is the machine-learning counterpart to a
|
| 59 |
+
cost-code lookup table: hand it "TPO roofing, 60 mil, fully adhered" and it returns the MasterFormat group,
|
| 60 |
+
ranked. English text, one label per input.
|
| 61 |
|
| 62 |
+
Built for estimating, takeoff, spec-writing and RAG pipelines that need to tag free text with CSI
|
| 63 |
+
MasterFormat codes without a human choosing from a 171-row list.
|
|
|
|
| 64 |
|
| 65 |
+
> **Early checkpoint, published as-is.** This is **step 700 of a planned 4,150-step fine-tune**
|
| 66 |
+
> (`mf-0.1`); training was paused on an 8 GB CPU-only Mac. Top-1 accuracy on the UFGS validation split is
|
| 67 |
+
> **18.4 %** (v1 split; **20.0 %** on the rebuilt v2 split) — well below this project's TF-IDF baseline
|
| 68 |
+
> (**53.1 %** / **57.5 %**) on the same splits. Treat this release
|
| 69 |
+
> as a reproducible starting point and a working deployment format, **not as a finished classifier**. A v2
|
| 70 |
+
> dataset and trainer are in the repository; see *Model* and *Limitations*.
|
| 71 |
|
| 72 |
+
MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or
|
| 73 |
+
endorsed by them.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 74 |
|
| 75 |
## Quick start
|
| 76 |
|
|
|
|
| 80 |
clf = pipeline("text-classification", model="constructelligence/masterformat-classifier", top_k=3)
|
| 81 |
for p in clf("4000 psi concrete slab on grade"):
|
| 82 |
print(p["label"], round(p["score"], 3))
|
| 83 |
+
# 04 20 00 Unit Masonry 0.082
|
| 84 |
+
# 03 60 00 Grouting 0.076
|
| 85 |
+
# 09 30 00 Tiling 0.063 <- not competitive yet; this is the step-700 checkpoint
|
| 86 |
+
```
|
| 87 |
+
|
| 88 |
+
Batched, with the level-1 division roll-up and JSON output:
|
| 89 |
+
|
| 90 |
+
```bash
|
| 91 |
+
pip install -r requirements.txt # transformers + torch
|
| 92 |
+
python predict.py --top-k 5 --divisions "8\" CMU wall, grout filled at 32\" o.c."
|
| 93 |
+
echo "Cat 6 data cabling and jacks" | python predict.py -
|
| 94 |
+
python predict.py --file items.txt --json > out.json
|
| 95 |
```
|
| 96 |
|
| 97 |
+
ONNX (no torch, ~34 MB int8):
|
| 98 |
|
| 99 |
```bash
|
| 100 |
+
pip install onnxruntime transformers
|
| 101 |
+
python predict.py --onnx onnx/model_quantized.onnx "Panelboards 208/120V 42 circuit"
|
| 102 |
```
|
| 103 |
|
| 104 |
The encoder is English-only; inputs should be English.
|
| 105 |
|
| 106 |
+
### ONNX exports
|
| 107 |
|
| 108 |
`onnx/model.onnx` (fp32, 134 MB) and `onnx/model_quantized.onnx` (dynamic int8, 34 MB) are exported with
|
| 109 |
+
dynamic batch and sequence axes, in the layout `transformers.js` / Optimum expect — inputs `input_ids`,
|
| 110 |
+
`attention_mask`, `token_type_ids`, output `logits`. The fp32 export matches PyTorch on a 5-text smoke set
|
| 111 |
+
(max |Δ logit| 5e-6, argmax agreement 5/5). Dynamic int8 quantization shifts low-confidence logits
|
| 112 |
+
(max |Δ| ≈ 1.5 on the same smoke set, argmax agreement 3/5), so verify the quantized model on your own
|
| 113 |
+
inputs before relying on it. Reproduce with `python scripts/export_onnx.py runs/mf-0.1`.
|
| 114 |
|
| 115 |
+
## Model
|
| 116 |
|
| 117 |
+
- **Base:** [`BAAI/bge-small-en-v1.5`](https://huggingface.co/BAAI/bge-small-en-v1.5) (BERT, 33M parameters, 384-d, MIT).
|
| 118 |
+
- **Head:** 171-way sequence-classification layer; `id2label` / `label2id` are in [`config.json`](config.json).
|
| 119 |
+
- **Recipe:** 2 planned epochs over the training set (per-class balance capped at 450 / floored at 300,
|
| 120 |
+
estimate-style augmentation), batch 32, max length 96, AdamW, 6 % linear warmup; body LR 6e-5, head LR 2e-3;
|
| 121 |
+
embeddings and the first 4 encoder layers frozen.
|
| 122 |
+
- **Checkpoint:** stopped at step 700; `train_state.json` records `{"step": 700, "val_acc": 0.18725}`.
|
| 123 |
+
- **A stronger v2 recipe already exists** in the source project (`label smoothing 0.05`, sentence-window
|
| 124 |
+
dataset, max length 64, lower learning rates) but no `mf-0.2` run has been completed yet.
|
| 125 |
+
|
| 126 |
+
## Results
|
| 127 |
+
|
| 128 |
+
All numbers are measured in this project on held-out data. `val` = UFGS text units, held out by a hash of
|
| 129 |
+
the normalised source text; `line_items` = 194 hand-labelled estimate line items; `manuals` = 6,890 chunks
|
| 130 |
+
from 153 sections of two real commercial project manuals. Out-of-taxonomy gold labels (e.g. `22 40 00`)
|
| 131 |
+
count only toward the division score, so a row's top-1 is computed over in-taxonomy items only.
|
| 132 |
+
|
| 133 |
+
| Scorer | val top-1 | val top-3 | val division | line-item top-1 | manual-chunk top-1 | manual-section top-1 |
|
| 134 |
+
|---|---:|---:|---:|---:|---:|---:|
|
| 135 |
+
| **mf-0.1 (step 700)** | **0.184** | 0.345 | 0.321 | 0.283 | 0.175 | 0.287 |
|
| 136 |
+
| TF-IDF + linear SGD (full run) | 0.531 | 0.690 | 0.634 | 0.665 | 0.290 | 0.607 |
|
| 137 |
+
| bge-small embeddings + logistic regression | 0.349 | 0.517 | 0.458 | 0.576 | 0.256 | 0.593 |
|
| 138 |
+
| Ensemble (0.25 embedding + 0.75 TF-IDF) | **0.532** | **0.700** | **0.637** | **0.702** | **0.318** | **0.640** |
|
| 139 |
+
|
| 140 |
+
The published checkpoint is the weakest scorer in this table. The best result on the real-world target
|
| 141 |
+
(hand-labelled estimate line items) is the **ensemble at 70.2 % top-1**, which is why the project's next
|
| 142 |
+
milestone is to finish the transformer run rather than ship `mf-0.1` as the answer.
|
| 143 |
+
|
| 144 |
+
`val` is the only column that depends on the dataset version: `line_items`, `manuals` and `manual_sections`
|
| 145 |
+
are fixed held-out sets, so those columns are directly comparable for every scorer. On the rebuilt **v2**
|
| 146 |
+
validation split (11,598 units) the same `mf-0.1` checkpoint scores **0.200 top-1 / 0.372 top-3 / 0.356
|
| 147 |
+
division**, against **0.575 / 0.732 / 0.675** for TF-IDF on that split. The `val` columns above are on the
|
| 148 |
+
original v1 split (13,518 units).
|
| 149 |
+
|
| 150 |
+
## Intended use
|
| 151 |
+
|
| 152 |
+
- **Use it for:** tagging construction text with a candidate MasterFormat group (top-3 shown), prototyping
|
| 153 |
+
an auto-classification step for estimating or spec workflows, and as a reproducible baseline to fine-tune
|
| 154 |
+
further.
|
| 155 |
+
- **Do not use it for:** unattended production takeoff, bid pricing, code compliance, or anything where a
|
| 156 |
+
wrong cost code has financial or contractual consequences. At 18.4 % top-1 it is not accurate enough.
|
| 157 |
+
- **Not a substitute for review.** Always keep a human in the loop for classification decisions.
|
| 158 |
|
| 159 |
## Training data and taxonomy
|
| 160 |
|
| 161 |
+
- **Source:** UFGS (Unified Facilities Guide Specifications) `.SEC` files — US federal works in the
|
| 162 |
+
**public domain**. Paragraphs and titles are parsed into labelled text units; boilerplate shared by
|
| 163 |
+
multiple sections is dropped, cross-references are stripped so section numbers cannot leak labels, and
|
| 164 |
+
long paragraphs are windowed into 20–60-word units.
|
| 165 |
+
- **Split:** held out by a hash of the normalised source text (10 %), so a paragraph and its augmentations
|
| 166 |
+
stay on the same side. The `mf-0.1` run used 66,396 train / 13,518 val rows; the rebuilt v2 dataset has
|
| 167 |
+
65,496 / 11,598.
|
| 168 |
- **Taxonomy:** 171 level-2 groups over 32 MasterFormat divisions; group numbers follow the MasterFormat
|
| 169 |
+
numbering convention and the short names are this project's own (see `config.json`).
|
| 170 |
|
| 171 |
## Limitations
|
| 172 |
|
| 173 |
+
- **Early checkpoint.** The 18.4 % top-1 is a snapshot of an interrupted run; it is not competitive with
|
| 174 |
+
the TF-IDF baseline yet. Do not use it for production takeoff without further training and re-evaluation.
|
| 175 |
- **Domain skew.** UFGS over-represents heavy-civil, water/wastewater and process work relative to commercial
|
| 176 |
building estimates; the training set is class-balanced, so raw predictions do not reflect building-project
|
| 177 |
priors.
|
| 178 |
+
- **Out-of-taxonomy inputs** (sections whose level-2 group is not among the 171) can only be scored at
|
| 179 |
+
division level.
|
| 180 |
+
- **Short, terse line items are the hardest inputs**; division (2-digit) accuracy is consistently higher than
|
| 181 |
+
group (6-digit) accuracy.
|
| 182 |
+
- **No section-number leakage, but no section-number help either.** The model classifies text content only;
|
| 183 |
+
if your input already contains a MasterFormat number, don't ask the model — read the number.
|
| 184 |
+
|
| 185 |
+
## Bias, risks and safety
|
| 186 |
+
|
| 187 |
+
- **Estimating bias.** Class-balanced training over a public-domain federal corpus does not represent any
|
| 188 |
+
particular firm's cost structure or regional practice. Do not treat output as a standard or an authority.
|
| 189 |
+
- **Trademark.** MasterFormat is a registered trademark of CSI / CSC; this model is not endorsed by them and
|
| 190 |
+
its group names are the project's own short descriptions, not CSI's official titles.
|
| 191 |
+
- **Privacy.** The model runs locally; no input text leaves your machine unless you call a hosted endpoint.
|
| 192 |
+
|
| 193 |
+
## FAQ
|
| 194 |
+
|
| 195 |
+
**What is MasterFormat?** The CSI/CSC MasterFormat is the North American standard for organising construction
|
| 196 |
+
specifications and cost data into numbered divisions and sections. This model predicts the **level-2 group**
|
| 197 |
+
(a 6-digit code such as `03 30 00 Cast-in-Place Concrete`), not the full section number.
|
| 198 |
+
|
| 199 |
+
**Can it classify a whole specification section?** Partly. Averaging the model's log-probabilities over a
|
| 200 |
+
section's chunks is the intended way to score a section; at step 700 that gets 28.7 % top-1 on two real
|
| 201 |
+
project manuals (53 % at division level). The TF-IDF baseline reaches 60.7 %.
|
| 202 |
+
|
| 203 |
+
**Can it read a MasterFormat number out of the text?** No. It classifies the description. If the number is
|
| 204 |
+
already present, parse it directly.
|
| 205 |
+
|
| 206 |
+
**Does it work offline / in the browser?** Yes — the ONNX exports are intended for `onnxruntime` and
|
| 207 |
+
`transformers.js`.
|
| 208 |
+
|
| 209 |
+
**Is a better model available?** Not yet as a published checkpoint. The project's TF-IDF + embedding
|
| 210 |
+
**ensemble scores 70.2 % top-1 on hand-labelled line items** and is the current production-candidate; the
|
| 211 |
+
next transformer run (`mf-0.2`) is expected to close most of the gap. Constructelligence's proprietary models
|
| 212 |
+
are at [constructelligence.co](https://constructelligence.co).
|
| 213 |
+
|
| 214 |
+
## Files
|
| 215 |
+
|
| 216 |
+
- `model.safetensors`, `config.json`, `tokenizer.json`, `tokenizer_config.json`, `vocab.txt`,
|
| 217 |
+
`special_tokens_map.json` — standard `transformers` checkpoint.
|
| 218 |
+
- `train_state.json` — step and validation accuracy of the saved checkpoint (text logs and optimizer state
|
| 219 |
+
are not published).
|
| 220 |
+
- `predict.py` — CLI example: batching, `--top-k`, `--divisions`, `--json`, stdin/file input, `--onnx`.
|
| 221 |
+
- `onnx/` — ONNX fp32 and int8 exports.
|
| 222 |
+
- `CITATION.cff`, `requirements.txt`.
|
| 223 |
+
|
| 224 |
+
## Citation
|
| 225 |
+
|
| 226 |
+
```bibtex
|
| 227 |
+
@misc{constructelligence_masterformat_classifier,
|
| 228 |
+
title = {MasterFormat Classifier (mf-0.1): construction spec and line-item classification},
|
| 229 |
+
author = {Constructelligence},
|
| 230 |
+
year = {2026},
|
| 231 |
+
howpublished = {\url{https://huggingface.co/constructelligence/masterformat-classifier}},
|
| 232 |
+
note = {Fine-tuned from BAAI/bge-small-en-v1.5; 171-way MasterFormat level-2 classifier}
|
| 233 |
+
}
|
| 234 |
+
```
|
| 235 |
|
| 236 |
## Licence and attribution
|
| 237 |
|
| 238 |
Released under the **MIT licence**, matching the base model. UFGS source text is public domain. MasterFormat
|
| 239 |
+
is a registered trademark of CSI / CSC; this project is not affiliated with or endorsed by them.
|
| 240 |
+
Constructelligence's production models are available at [constructelligence.co](https://constructelligence.co).
|
predict.py
CHANGED
|
@@ -1,24 +1,157 @@
|
|
| 1 |
-
"""Classify construction text into MasterFormat level-2 groups.
|
| 2 |
|
|
|
|
|
|
|
| 3 |
pip install transformers torch
|
| 4 |
python predict.py "4000 psi concrete slab on grade" "TPO roofing, 60 mil, fully adhered"
|
|
|
|
|
|
|
|
|
|
| 5 |
|
| 6 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 7 |
"""
|
|
|
|
|
|
|
| 8 |
import sys
|
| 9 |
|
| 10 |
-
from transformers import pipeline
|
| 11 |
-
|
| 12 |
MODEL = "constructelligence/masterformat-classifier"
|
| 13 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 14 |
|
| 15 |
-
def
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 20 |
for p in preds:
|
| 21 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 22 |
|
| 23 |
|
| 24 |
if __name__ == "__main__":
|
|
|
|
| 1 |
+
"""Classify construction text into MasterFormat level-2 groups (171 classes, 32 divisions).
|
| 2 |
|
| 3 |
+
Examples
|
| 4 |
+
--------
|
| 5 |
pip install transformers torch
|
| 6 |
python predict.py "4000 psi concrete slab on grade" "TPO roofing, 60 mil, fully adhered"
|
| 7 |
+
python predict.py --top-k 5 --divisions "8\" CMU wall, grout filled at 32\" o.c."
|
| 8 |
+
echo "Cat 6 data cabling and jacks" | python predict.py -
|
| 9 |
+
python predict.py --file items.txt --json > out.json
|
| 10 |
|
| 11 |
+
ONNX (no torch, ~4 MB quantized weights):
|
| 12 |
+
pip install onnxruntime transformers
|
| 13 |
+
python predict.py --onnx onnx/model_quantized.onnx "Panelboards 208/120V 42 circuit"
|
| 14 |
+
|
| 15 |
+
Without arguments it classifies one built-in example.
|
| 16 |
"""
|
| 17 |
+
import argparse
|
| 18 |
+
import json
|
| 19 |
import sys
|
| 20 |
|
|
|
|
|
|
|
| 21 |
MODEL = "constructelligence/masterformat-classifier"
|
| 22 |
|
| 23 |
+
# Division code -> name. Kept local so --divisions needs no extra package.
|
| 24 |
+
DIVISIONS = {
|
| 25 |
+
"01": "General Requirements", "02": "Existing Conditions", "03": "Concrete",
|
| 26 |
+
"04": "Masonry", "05": "Metals", "06": "Wood, Plastics, and Composites",
|
| 27 |
+
"07": "Thermal and Moisture Protection", "08": "Openings", "09": "Finishes",
|
| 28 |
+
"10": "Specialties", "11": "Equipment", "12": "Furnishings",
|
| 29 |
+
"13": "Special Construction", "14": "Conveying Equipment", "21": "Fire Suppression",
|
| 30 |
+
"22": "Plumbing", "23": "HVAC", "25": "Integrated Automation", "26": "Electrical",
|
| 31 |
+
"27": "Communications", "28": "Electronic Safety and Security", "31": "Earthwork",
|
| 32 |
+
"32": "Exterior Improvements", "33": "Utilities", "34": "Transportation",
|
| 33 |
+
"35": "Waterway and Marine Construction", "40": "Process Interconnections",
|
| 34 |
+
"41": "Material Processing and Handling Equipment",
|
| 35 |
+
"43": "Process Gas and Liquid Handling, Purification, and Storage Equipment",
|
| 36 |
+
"44": "Pollution and Waste Control Equipment", "46": "Water and Wastewater Equipment",
|
| 37 |
+
"48": "Electrical Power Generation",
|
| 38 |
+
}
|
| 39 |
+
|
| 40 |
+
EXAMPLE = "4000 psi concrete slab on grade"
|
| 41 |
+
|
| 42 |
+
|
| 43 |
+
def split_label(label):
|
| 44 |
+
"""'03 30 00 Cast-in-Place Concrete' -> ('03 30 00', 'Cast-in-Place Concrete')."""
|
| 45 |
+
code, name = label[:8].strip(), label[8:].strip()
|
| 46 |
+
return code, (name or label)
|
| 47 |
+
|
| 48 |
+
|
| 49 |
+
def division_rollup(preds):
|
| 50 |
+
"""Sum level-2 probabilities by the first two digits (division)."""
|
| 51 |
+
totals = {}
|
| 52 |
+
for p in preds:
|
| 53 |
+
div = p["code"][:2]
|
| 54 |
+
totals[div] = totals.get(div, 0.0) + p["score"]
|
| 55 |
+
return [
|
| 56 |
+
{"code": d, "name": DIVISIONS.get(d, d), "score": s}
|
| 57 |
+
for d, s in sorted(totals.items(), key=lambda kv: -kv[1])
|
| 58 |
+
]
|
| 59 |
+
|
| 60 |
+
|
| 61 |
+
def hf_scorer(model, top_k):
|
| 62 |
+
from transformers import pipeline
|
| 63 |
+
|
| 64 |
+
clf = pipeline("text-classification", model=model, top_k=top_k)
|
| 65 |
+
return lambda texts: [[d for d in out] for out in clf(texts)]
|
| 66 |
+
|
| 67 |
|
| 68 |
+
def onnx_scorer(path, top_k):
|
| 69 |
+
"""Run onnx/model*.onnx directly. Mirrors scripts/export_onnx.py: inputs
|
| 70 |
+
input_ids / attention_mask / token_type_ids, output logits [batch, num_labels]."""
|
| 71 |
+
import numpy as np
|
| 72 |
+
import onnxruntime as ort
|
| 73 |
+
from pathlib import Path
|
| 74 |
+
from transformers import AutoTokenizer
|
| 75 |
+
|
| 76 |
+
# The labels and tokenizer live beside onnx/ in the repo; fall back to the Hub.
|
| 77 |
+
root = Path(path).resolve().parent.parent
|
| 78 |
+
src = str(root) if (root / "tokenizer.json").exists() or (root / "vocab.txt").exists() else MODEL
|
| 79 |
+
tok = AutoTokenizer.from_pretrained(src)
|
| 80 |
+
id2label = None
|
| 81 |
+
try:
|
| 82 |
+
c = json.loads((root / "config.json").read_text())
|
| 83 |
+
id2label = {int(k): v for k, v in c.get("id2label", {}).items()}
|
| 84 |
+
except Exception:
|
| 85 |
+
id2label = None
|
| 86 |
+
sess = ort.InferenceSession(path, providers=["CPUExecutionProvider"])
|
| 87 |
+
names = {i.name for i in sess.get_inputs()}
|
| 88 |
+
|
| 89 |
+
def score(texts):
|
| 90 |
+
enc = tok(texts, truncation=True, max_length=128, padding=True, return_tensors="np")
|
| 91 |
+
feed = {n: enc[n].astype(np.int64) for n in ("input_ids", "attention_mask", "token_type_ids") if n in names}
|
| 92 |
+
logits = sess.run(None, feed)[0]
|
| 93 |
+
idx = np.argsort(-logits, 1)[:, :top_k]
|
| 94 |
+
shifted = logits - logits.max(1, keepdims=True)
|
| 95 |
+
probs = np.exp(shifted) / np.exp(shifted).sum(1, keepdims=True)
|
| 96 |
+
out = []
|
| 97 |
+
for row, cols in zip(probs, idx):
|
| 98 |
+
out.append([{"label": id2label.get(int(c), str(int(c))), "score": float(row[c])} for c in cols])
|
| 99 |
+
return out
|
| 100 |
+
|
| 101 |
+
return score
|
| 102 |
+
|
| 103 |
+
|
| 104 |
+
def predict(texts, score, top_k, want_div):
|
| 105 |
+
results = []
|
| 106 |
+
for text, preds in zip(texts, score(texts)):
|
| 107 |
+
pl = []
|
| 108 |
for p in preds:
|
| 109 |
+
code, name = split_label(p["label"])
|
| 110 |
+
pl.append({"code": code, "label": p["label"], "name": name, "score": round(float(p["score"]), 4)})
|
| 111 |
+
item = {"text": text, "predictions": pl}
|
| 112 |
+
if want_div:
|
| 113 |
+
item["divisions"] = [dict(d, score=round(d["score"], 4)) for d in division_rollup(pl)]
|
| 114 |
+
results.append(item)
|
| 115 |
+
return results
|
| 116 |
+
|
| 117 |
+
|
| 118 |
+
def main(argv=None):
|
| 119 |
+
ap = argparse.ArgumentParser(description="MasterFormat level-2 classifier (mf-0.1).")
|
| 120 |
+
ap.add_argument("text", nargs="*", help="text to classify; use '-' to read lines from stdin")
|
| 121 |
+
ap.add_argument("--model", default=MODEL, help="HF repo id or local checkpoint dir")
|
| 122 |
+
ap.add_argument("--onnx", metavar="PATH", help="classify with an ONNX model instead of PyTorch")
|
| 123 |
+
ap.add_argument("--top-k", type=int, default=3, help="number of level-2 predictions to show")
|
| 124 |
+
ap.add_argument("--divisions", action="store_true", help="also show the level-1 (division) roll-up")
|
| 125 |
+
ap.add_argument("--file", help="read newline-separated inputs from a file")
|
| 126 |
+
ap.add_argument("--json", action="store_true", help="emit JSON instead of a table")
|
| 127 |
+
a = ap.parse_args(argv)
|
| 128 |
+
|
| 129 |
+
texts = list(a.text)
|
| 130 |
+
if a.file:
|
| 131 |
+
texts += [l.rstrip("\n") for l in open(a.file, encoding="utf-8") if l.strip()]
|
| 132 |
+
if "-" in texts:
|
| 133 |
+
texts = [t for t in texts if t != "-"] + [l.rstrip("\n") for l in sys.stdin if l.strip()]
|
| 134 |
+
if not texts:
|
| 135 |
+
texts = [EXAMPLE]
|
| 136 |
+
texts = [t for t in texts if t.strip()]
|
| 137 |
+
if not texts:
|
| 138 |
+
ap.error("no input text")
|
| 139 |
+
|
| 140 |
+
score = onnx_scorer(a.onnx, a.top_k) if a.onnx else hf_scorer(a.model, a.top_k)
|
| 141 |
+
results = predict(texts, score, a.top_k, a.divisions)
|
| 142 |
+
|
| 143 |
+
if a.json:
|
| 144 |
+
json.dump(results, sys.stdout, indent=2, ensure_ascii=False)
|
| 145 |
+
print()
|
| 146 |
+
return
|
| 147 |
+
for r in results:
|
| 148 |
+
print(f"\n{r['text']}")
|
| 149 |
+
for p in r["predictions"]:
|
| 150 |
+
print(f" {p['code']} {p['name']:<48.48} {p['score']:.3f}")
|
| 151 |
+
if a.divisions:
|
| 152 |
+
print(" -- divisions --")
|
| 153 |
+
for d in r.get("divisions", []):
|
| 154 |
+
print(f" {d['code']} {d['name']:<48.48} {d['score']:.3f}")
|
| 155 |
|
| 156 |
|
| 157 |
if __name__ == "__main__":
|
requirements.txt
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Core inference (PyTorch)
|
| 2 |
+
transformers>=4.40
|
| 3 |
+
torch>=2.0
|
| 4 |
+
|
| 5 |
+
# ONNX path: pip install -r requirements.txt onnxruntime
|
| 6 |
+
# onnxruntime>=1.17
|
| 7 |
+
# numpy>=1.24
|