constructelligence commited on
Commit
1d29f6c
·
verified ·
1 Parent(s): 3ccb233

Improve model card (SEO title, FAQ, citation, widgets) and add predict.py CLI, CITATION.cff, requirements.txt

Browse files
Files changed (4) hide show
  1. CITATION.cff +24 -0
  2. README.md +175 -67
  3. predict.py +143 -10
  4. requirements.txt +7 -0
CITATION.cff ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ cff-version: 1.2.0
2
+ title: "MasterFormat Classifier (mf-0.1)"
3
+ message: "If you use this model, please cite it as below."
4
+ type: software
5
+ authors:
6
+ - name: "Constructelligence"
7
+ website: "https://constructelligence.co"
8
+ license: MIT
9
+ url: "https://huggingface.co/constructelligence/masterformat-classifier"
10
+ repository-code: "https://huggingface.co/constructelligence/masterformat-classifier"
11
+ abstract: >-
12
+ A BERT-based text classifier that maps construction line items, specification
13
+ paragraphs and section titles to one of 171 MasterFormat level-2 groups across
14
+ 32 divisions. Fine-tuned from BAAI/bge-small-en-v1.5 on public-domain UFGS
15
+ specification text. Published as an early (step-700) checkpoint.
16
+ keywords:
17
+ - MasterFormat
18
+ - construction
19
+ - text-classification
20
+ - specifications
21
+ - estimating
22
+ - takeoff
23
+ - UFGS
24
+ - BERT
README.md CHANGED
@@ -1,17 +1,29 @@
1
  ---
 
 
2
  license: mit
3
- base_model: BAAI/bge-small-en-v1.5
4
  library_name: transformers
5
  pipeline_tag: text-classification
6
- language: en
7
  tags:
8
- - construction
9
  - masterformat
 
 
 
 
10
  - text-classification
11
- - spec-writing
12
- - takeoff
13
  - bert
14
  - onnx
 
 
 
 
 
 
 
 
 
15
  metrics:
16
  - accuracy
17
  model-index:
@@ -19,7 +31,7 @@ model-index:
19
  results:
20
  - task:
21
  type: text-classification
22
- name: MasterFormat level-2 classification
23
  dataset:
24
  type: ufgs-val
25
  name: UFGS held-out text units (13,518)
@@ -28,43 +40,37 @@ model-index:
28
  - type: accuracy
29
  value: 0.184
30
  name: Top-1 accuracy (step-700 checkpoint)
 
 
 
 
 
 
 
 
 
31
  ---
32
 
33
- # MasterFormat level-2 classifier (mf-0.1)
34
 
35
- > **Early checkpoint, published as-is.** This is step 700 of a planned 4,150-step fine-tune; training was
36
- > paused on an 8 GB CPU-only Mac. Top-1 accuracy on the full UFGS validation split is **18.4%** — well below the
37
- > TF-IDF baseline (53.1%) on the same split. Treat this release as a reproducible starting point and a
38
- > working deployment format, not as a finished classifier. A resume script and the full recipe are in the
39
- > source project.
40
 
41
- Classifies a construction line item, specification paragraph, submittal or section title into one of **171
42
- MasterFormat level-2 groups** (e.g. `03 30 00 Cast-in-Place Concrete`, `23 30 00 HVAC Air Distribution`)
43
- across 32 divisions. English text, one label per input.
44
 
45
- MasterFormat is a registered trademark of CSI / CSC. This project is not affiliated with or endorsed by them.
 
 
 
 
 
46
 
47
- ## Model
48
-
49
- - Base: [`BAAI/bge-small-en-v1.5`](https://huggingface.co/BAAI/bge-small-en-v1.5) (BERT, 33M parameters, 384-d), MIT.
50
- - Fine-tuned with a 171-way sequence-classification head; `id2label` / `label2id` are in [`config.json`](config.json).
51
- - Training: 2 epochs planned over 66,396 rows (per-class balance capped at 450 / floored at 300, estimate-style
52
- augmentation), batch 32, max length 96, AdamW, 6% linear warmup; body LR 6e-5, head LR 2e-3; embeddings and
53
- the first 4 encoder layers frozen.
54
- - Checkpoint stopped at step 700; `train_state.json` records `{"step": 700, "val_acc": 0.18725}`.
55
-
56
- ## Results
57
-
58
- All numbers are measured on held-out data from the same project. `val` = UFGS text units,
59
- `line_items` = 194 hand-labelled estimate line items, `manuals` = 6,890 chunks from 153 sections of two real
60
- commercial project manuals (out-of-taxonomy gold labels count only toward the division score).
61
-
62
- | Scorer | val top-1 | val top-3 | val division | line-item top-1 | manual-chunk top-1 | manual-section top-1 |
63
- |---|---:|---:|---:|---:|---:|---:|
64
- | **mf-0.1 (step 700)** | 0.184 | 0.345 | 0.321 | 0.283 | 0.175 | 0.287 |
65
- | TF-IDF + linear SGD (full run) | 0.531 | 0.690 | 0.634 | 0.665 | 0.290 | 0.607 |
66
- | bge-small embeddings + logistic regression | 0.349 | 0.517 | 0.458 | 0.576 | 0.256 | 0.593 |
67
- | Ensemble (0.25 embedding + 0.75 TF-IDF) | 0.532 | 0.700 | 0.637 | 0.702 | 0.318 | 0.640 |
68
 
69
  ## Quick start
70
 
@@ -74,59 +80,161 @@ from transformers import pipeline
74
  clf = pipeline("text-classification", model="constructelligence/masterformat-classifier", top_k=3)
75
  for p in clf("4000 psi concrete slab on grade"):
76
  print(p["label"], round(p["score"], 3))
 
 
 
 
 
 
 
 
 
 
 
 
77
  ```
78
 
79
- Or with the included example:
80
 
81
  ```bash
82
- pip install transformers torch
83
- python predict.py "TPO roofing, 60 mil, fully adhered"
84
  ```
85
 
86
  The encoder is English-only; inputs should be English.
87
 
88
- ### ONNX
89
 
90
  `onnx/model.onnx` (fp32, 134 MB) and `onnx/model_quantized.onnx` (dynamic int8, 34 MB) are exported with
91
- dynamic batch and sequence axes, in the layout `transformers.js` / Optimum expect, inputs `input_ids`,
92
- `attention_mask`, `token_type_ids`, output `logits`. The fp32 export matches PyTorch on a 4-text smoke set
93
- (max |delta logit| 5e-6, argmax agreement 4/4). Dynamic int8 quantization shifts low-confidence logits
94
- (max |delta| ~1.4 on the same smoke set, argmax agreement 3/4), so verify the quantized model on your own
95
- inputs before using it; see the project's `scripts/export_onnx.py` for the exact command.
96
 
97
- ## Files
98
 
99
- - `model.safetensors`, `config.json`, `tokenizer.json`, `tokenizer_config.json`, `vocab.txt`,
100
- `special_tokens_map.json` — standard `transformers` checkpoint.
101
- - `train_state.json` — step and validation accuracy of the saved checkpoint (text logs and optimizer state
102
- are not published).
103
- - `predict.py` — minimal pipeline example.
104
- - `onnx/` — ONNX fp32 and int8 exports.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
105
 
106
  ## Training data and taxonomy
107
 
108
- - **UFGS** (Unified Facilities Guide Specifications) `.SEC` files, US federal works in the **public domain**.
109
- Paragraphs and titles are parsed into labelled text units; boilerplate shared by 3+ sections is dropped,
110
- cross-references are stripped so section numbers cannot leak labels, and long paragraphs are windowed into
111
- 20–60-word units.
112
- - **Split:** held out by a hash of the normalised source text (10%), so a paragraph and its augmentations stay
113
- on the same side. 66,396 train / 13,518 val rows.
 
114
  - **Taxonomy:** 171 level-2 groups over 32 MasterFormat divisions; group numbers follow the MasterFormat
115
- numbering convention and the short names are the project's own (see `config.json`).
116
 
117
  ## Limitations
118
 
119
- - **Early checkpoint.** The 18.4% top-1 is a snapshot of an interrupted run; it is not competitive with the
120
- TF-IDF baseline yet. Do not use it for production takeoff without further training and re-evaluation.
121
  - **Domain skew.** UFGS over-represents heavy-civil, water/wastewater and process work relative to commercial
122
  building estimates; the training set is class-balanced, so raw predictions do not reflect building-project
123
  priors.
124
- - **Out-of-taxonomy inputs** (sections whose level-2 group is not in the 171) can only be scored at division
125
- level.
126
- - Short, terse line items are the hardest inputs; division (2-digit) accuracy is higher than group accuracy.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
127
 
128
  ## Licence and attribution
129
 
130
  Released under the **MIT licence**, matching the base model. UFGS source text is public domain. MasterFormat
131
- is a registered trademark of CSI / CSC; this project is not affiliated with or endorsed by them. Constructelligence's
132
- production models are available at [constructelligence.co](https://constructelligence.co).
 
1
  ---
2
+ language:
3
+ - en
4
  license: mit
 
5
  library_name: transformers
6
  pipeline_tag: text-classification
7
+ base_model: BAAI/bge-small-en-v1.5
8
  tags:
 
9
  - masterformat
10
+ - masterformat-classifier
11
+ - csi-masterformat
12
+ - construction
13
+ - construction-technology
14
  - text-classification
15
+ - sequence-classification
 
16
  - bert
17
  - onnx
18
+ - safetensors
19
+ - specs
20
+ - spec-writing
21
+ - specifications
22
+ - takeoff
23
+ - estimating
24
+ - cost-code
25
+ - ufgs
26
+ - public-domain
27
  metrics:
28
  - accuracy
29
  model-index:
 
31
  results:
32
  - task:
33
  type: text-classification
34
+ name: MasterFormat level-2 group classification
35
  dataset:
36
  type: ufgs-val
37
  name: UFGS held-out text units (13,518)
 
40
  - type: accuracy
41
  value: 0.184
42
  name: Top-1 accuracy (step-700 checkpoint)
43
+ widget:
44
+ - text: 4000 psi concrete slab on grade
45
+ example_title: Cast-in-place concrete
46
+ - text: TPO roofing, 60 mil, fully adhered
47
+ example_title: Membrane roofing
48
+ - text: Cat 6 data cabling and jacks
49
+ example_title: Structured cabling
50
+ - text: Wet pipe sprinkler system, light hazard
51
+ example_title: Fire suppression
52
  ---
53
 
54
+ # MasterFormat classifier (mf-0.1) — construction spec & line-item classification
55
 
56
+ A BERT text classifier that maps **construction line items, specification paragraphs, submittals and
57
+ section titles** to one of **171 MasterFormat level-2 groups** (e.g. `03 30 00 Cast-in-Place Concrete`,
58
+ `23 30 00 HVAC Air Distribution`) across **32 divisions**. It is the machine-learning counterpart to a
59
+ cost-code lookup table: hand it "TPO roofing, 60 mil, fully adhered" and it returns the MasterFormat group,
60
+ ranked. English text, one label per input.
61
 
62
+ Built for estimating, takeoff, spec-writing and RAG pipelines that need to tag free text with CSI
63
+ MasterFormat codes without a human choosing from a 171-row list.
 
64
 
65
+ > **Early checkpoint, published as-is.** This is **step 700 of a planned 4,150-step fine-tune**
66
+ > (`mf-0.1`); training was paused on an 8 GB CPU-only Mac. Top-1 accuracy on the UFGS validation split is
67
+ > **18.4 %** (v1 split; **20.0 %** on the rebuilt v2 split) — well below this project's TF-IDF baseline
68
+ > (**53.1 %** / **57.5 %**) on the same splits. Treat this release
69
+ > as a reproducible starting point and a working deployment format, **not as a finished classifier**. A v2
70
+ > dataset and trainer are in the repository; see *Model* and *Limitations*.
71
 
72
+ MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or
73
+ endorsed by them.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
74
 
75
  ## Quick start
76
 
 
80
  clf = pipeline("text-classification", model="constructelligence/masterformat-classifier", top_k=3)
81
  for p in clf("4000 psi concrete slab on grade"):
82
  print(p["label"], round(p["score"], 3))
83
+ # 04 20 00 Unit Masonry 0.082
84
+ # 03 60 00 Grouting 0.076
85
+ # 09 30 00 Tiling 0.063 <- not competitive yet; this is the step-700 checkpoint
86
+ ```
87
+
88
+ Batched, with the level-1 division roll-up and JSON output:
89
+
90
+ ```bash
91
+ pip install -r requirements.txt # transformers + torch
92
+ python predict.py --top-k 5 --divisions "8\" CMU wall, grout filled at 32\" o.c."
93
+ echo "Cat 6 data cabling and jacks" | python predict.py -
94
+ python predict.py --file items.txt --json > out.json
95
  ```
96
 
97
+ ONNX (no torch, ~34 MB int8):
98
 
99
  ```bash
100
+ pip install onnxruntime transformers
101
+ python predict.py --onnx onnx/model_quantized.onnx "Panelboards 208/120V 42 circuit"
102
  ```
103
 
104
  The encoder is English-only; inputs should be English.
105
 
106
+ ### ONNX exports
107
 
108
  `onnx/model.onnx` (fp32, 134 MB) and `onnx/model_quantized.onnx` (dynamic int8, 34 MB) are exported with
109
+ dynamic batch and sequence axes, in the layout `transformers.js` / Optimum expect — inputs `input_ids`,
110
+ `attention_mask`, `token_type_ids`, output `logits`. The fp32 export matches PyTorch on a 5-text smoke set
111
+ (max |Δ logit| 5e-6, argmax agreement 5/5). Dynamic int8 quantization shifts low-confidence logits
112
+ (max |Δ| ≈ 1.5 on the same smoke set, argmax agreement 3/5), so verify the quantized model on your own
113
+ inputs before relying on it. Reproduce with `python scripts/export_onnx.py runs/mf-0.1`.
114
 
115
+ ## Model
116
 
117
+ - **Base:** [`BAAI/bge-small-en-v1.5`](https://huggingface.co/BAAI/bge-small-en-v1.5) (BERT, 33M parameters, 384-d, MIT).
118
+ - **Head:** 171-way sequence-classification layer; `id2label` / `label2id` are in [`config.json`](config.json).
119
+ - **Recipe:** 2 planned epochs over the training set (per-class balance capped at 450 / floored at 300,
120
+ estimate-style augmentation), batch 32, max length 96, AdamW, 6 % linear warmup; body LR 6e-5, head LR 2e-3;
121
+ embeddings and the first 4 encoder layers frozen.
122
+ - **Checkpoint:** stopped at step 700; `train_state.json` records `{"step": 700, "val_acc": 0.18725}`.
123
+ - **A stronger v2 recipe already exists** in the source project (`label smoothing 0.05`, sentence-window
124
+ dataset, max length 64, lower learning rates) but no `mf-0.2` run has been completed yet.
125
+
126
+ ## Results
127
+
128
+ All numbers are measured in this project on held-out data. `val` = UFGS text units, held out by a hash of
129
+ the normalised source text; `line_items` = 194 hand-labelled estimate line items; `manuals` = 6,890 chunks
130
+ from 153 sections of two real commercial project manuals. Out-of-taxonomy gold labels (e.g. `22 40 00`)
131
+ count only toward the division score, so a row's top-1 is computed over in-taxonomy items only.
132
+
133
+ | Scorer | val top-1 | val top-3 | val division | line-item top-1 | manual-chunk top-1 | manual-section top-1 |
134
+ |---|---:|---:|---:|---:|---:|---:|
135
+ | **mf-0.1 (step 700)** | **0.184** | 0.345 | 0.321 | 0.283 | 0.175 | 0.287 |
136
+ | TF-IDF + linear SGD (full run) | 0.531 | 0.690 | 0.634 | 0.665 | 0.290 | 0.607 |
137
+ | bge-small embeddings + logistic regression | 0.349 | 0.517 | 0.458 | 0.576 | 0.256 | 0.593 |
138
+ | Ensemble (0.25 embedding + 0.75 TF-IDF) | **0.532** | **0.700** | **0.637** | **0.702** | **0.318** | **0.640** |
139
+
140
+ The published checkpoint is the weakest scorer in this table. The best result on the real-world target
141
+ (hand-labelled estimate line items) is the **ensemble at 70.2 % top-1**, which is why the project's next
142
+ milestone is to finish the transformer run rather than ship `mf-0.1` as the answer.
143
+
144
+ `val` is the only column that depends on the dataset version: `line_items`, `manuals` and `manual_sections`
145
+ are fixed held-out sets, so those columns are directly comparable for every scorer. On the rebuilt **v2**
146
+ validation split (11,598 units) the same `mf-0.1` checkpoint scores **0.200 top-1 / 0.372 top-3 / 0.356
147
+ division**, against **0.575 / 0.732 / 0.675** for TF-IDF on that split. The `val` columns above are on the
148
+ original v1 split (13,518 units).
149
+
150
+ ## Intended use
151
+
152
+ - **Use it for:** tagging construction text with a candidate MasterFormat group (top-3 shown), prototyping
153
+ an auto-classification step for estimating or spec workflows, and as a reproducible baseline to fine-tune
154
+ further.
155
+ - **Do not use it for:** unattended production takeoff, bid pricing, code compliance, or anything where a
156
+ wrong cost code has financial or contractual consequences. At 18.4 % top-1 it is not accurate enough.
157
+ - **Not a substitute for review.** Always keep a human in the loop for classification decisions.
158
 
159
  ## Training data and taxonomy
160
 
161
+ - **Source:** UFGS (Unified Facilities Guide Specifications) `.SEC` files — US federal works in the
162
+ **public domain**. Paragraphs and titles are parsed into labelled text units; boilerplate shared by
163
+ multiple sections is dropped, cross-references are stripped so section numbers cannot leak labels, and
164
+ long paragraphs are windowed into 20–60-word units.
165
+ - **Split:** held out by a hash of the normalised source text (10 %), so a paragraph and its augmentations
166
+ stay on the same side. The `mf-0.1` run used 66,396 train / 13,518 val rows; the rebuilt v2 dataset has
167
+ 65,496 / 11,598.
168
  - **Taxonomy:** 171 level-2 groups over 32 MasterFormat divisions; group numbers follow the MasterFormat
169
+ numbering convention and the short names are this project's own (see `config.json`).
170
 
171
  ## Limitations
172
 
173
+ - **Early checkpoint.** The 18.4 % top-1 is a snapshot of an interrupted run; it is not competitive with
174
+ the TF-IDF baseline yet. Do not use it for production takeoff without further training and re-evaluation.
175
  - **Domain skew.** UFGS over-represents heavy-civil, water/wastewater and process work relative to commercial
176
  building estimates; the training set is class-balanced, so raw predictions do not reflect building-project
177
  priors.
178
+ - **Out-of-taxonomy inputs** (sections whose level-2 group is not among the 171) can only be scored at
179
+ division level.
180
+ - **Short, terse line items are the hardest inputs**; division (2-digit) accuracy is consistently higher than
181
+ group (6-digit) accuracy.
182
+ - **No section-number leakage, but no section-number help either.** The model classifies text content only;
183
+ if your input already contains a MasterFormat number, don't ask the model — read the number.
184
+
185
+ ## Bias, risks and safety
186
+
187
+ - **Estimating bias.** Class-balanced training over a public-domain federal corpus does not represent any
188
+ particular firm's cost structure or regional practice. Do not treat output as a standard or an authority.
189
+ - **Trademark.** MasterFormat is a registered trademark of CSI / CSC; this model is not endorsed by them and
190
+ its group names are the project's own short descriptions, not CSI's official titles.
191
+ - **Privacy.** The model runs locally; no input text leaves your machine unless you call a hosted endpoint.
192
+
193
+ ## FAQ
194
+
195
+ **What is MasterFormat?** The CSI/CSC MasterFormat is the North American standard for organising construction
196
+ specifications and cost data into numbered divisions and sections. This model predicts the **level-2 group**
197
+ (a 6-digit code such as `03 30 00 Cast-in-Place Concrete`), not the full section number.
198
+
199
+ **Can it classify a whole specification section?** Partly. Averaging the model's log-probabilities over a
200
+ section's chunks is the intended way to score a section; at step 700 that gets 28.7 % top-1 on two real
201
+ project manuals (53 % at division level). The TF-IDF baseline reaches 60.7 %.
202
+
203
+ **Can it read a MasterFormat number out of the text?** No. It classifies the description. If the number is
204
+ already present, parse it directly.
205
+
206
+ **Does it work offline / in the browser?** Yes — the ONNX exports are intended for `onnxruntime` and
207
+ `transformers.js`.
208
+
209
+ **Is a better model available?** Not yet as a published checkpoint. The project's TF-IDF + embedding
210
+ **ensemble scores 70.2 % top-1 on hand-labelled line items** and is the current production-candidate; the
211
+ next transformer run (`mf-0.2`) is expected to close most of the gap. Constructelligence's proprietary models
212
+ are at [constructelligence.co](https://constructelligence.co).
213
+
214
+ ## Files
215
+
216
+ - `model.safetensors`, `config.json`, `tokenizer.json`, `tokenizer_config.json`, `vocab.txt`,
217
+ `special_tokens_map.json` — standard `transformers` checkpoint.
218
+ - `train_state.json` — step and validation accuracy of the saved checkpoint (text logs and optimizer state
219
+ are not published).
220
+ - `predict.py` — CLI example: batching, `--top-k`, `--divisions`, `--json`, stdin/file input, `--onnx`.
221
+ - `onnx/` — ONNX fp32 and int8 exports.
222
+ - `CITATION.cff`, `requirements.txt`.
223
+
224
+ ## Citation
225
+
226
+ ```bibtex
227
+ @misc{constructelligence_masterformat_classifier,
228
+ title = {MasterFormat Classifier (mf-0.1): construction spec and line-item classification},
229
+ author = {Constructelligence},
230
+ year = {2026},
231
+ howpublished = {\url{https://huggingface.co/constructelligence/masterformat-classifier}},
232
+ note = {Fine-tuned from BAAI/bge-small-en-v1.5; 171-way MasterFormat level-2 classifier}
233
+ }
234
+ ```
235
 
236
  ## Licence and attribution
237
 
238
  Released under the **MIT licence**, matching the base model. UFGS source text is public domain. MasterFormat
239
+ is a registered trademark of CSI / CSC; this project is not affiliated with or endorsed by them.
240
+ Constructelligence's production models are available at [constructelligence.co](https://constructelligence.co).
predict.py CHANGED
@@ -1,24 +1,157 @@
1
- """Classify construction text into MasterFormat level-2 groups.
2
 
 
 
3
  pip install transformers torch
4
  python predict.py "4000 psi concrete slab on grade" "TPO roofing, 60 mil, fully adhered"
 
 
 
5
 
6
- Without arguments it classifies one example line item.
 
 
 
 
7
  """
 
 
8
  import sys
9
 
10
- from transformers import pipeline
11
-
12
  MODEL = "constructelligence/masterformat-classifier"
13
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14
 
15
- def main():
16
- texts = sys.argv[1:] or ["4000 psi concrete slab on grade"]
17
- clf = pipeline("text-classification", model=MODEL, top_k=3)
18
- for text, preds in zip(texts, clf(texts)):
19
- print(f"\n{text}")
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
20
  for p in preds:
21
- print(f" {p['label']:64s} {p['score']:.3f}")
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
22
 
23
 
24
  if __name__ == "__main__":
 
1
+ """Classify construction text into MasterFormat level-2 groups (171 classes, 32 divisions).
2
 
3
+ Examples
4
+ --------
5
  pip install transformers torch
6
  python predict.py "4000 psi concrete slab on grade" "TPO roofing, 60 mil, fully adhered"
7
+ python predict.py --top-k 5 --divisions "8\" CMU wall, grout filled at 32\" o.c."
8
+ echo "Cat 6 data cabling and jacks" | python predict.py -
9
+ python predict.py --file items.txt --json > out.json
10
 
11
+ ONNX (no torch, ~4 MB quantized weights):
12
+ pip install onnxruntime transformers
13
+ python predict.py --onnx onnx/model_quantized.onnx "Panelboards 208/120V 42 circuit"
14
+
15
+ Without arguments it classifies one built-in example.
16
  """
17
+ import argparse
18
+ import json
19
  import sys
20
 
 
 
21
  MODEL = "constructelligence/masterformat-classifier"
22
 
23
+ # Division code -> name. Kept local so --divisions needs no extra package.
24
+ DIVISIONS = {
25
+ "01": "General Requirements", "02": "Existing Conditions", "03": "Concrete",
26
+ "04": "Masonry", "05": "Metals", "06": "Wood, Plastics, and Composites",
27
+ "07": "Thermal and Moisture Protection", "08": "Openings", "09": "Finishes",
28
+ "10": "Specialties", "11": "Equipment", "12": "Furnishings",
29
+ "13": "Special Construction", "14": "Conveying Equipment", "21": "Fire Suppression",
30
+ "22": "Plumbing", "23": "HVAC", "25": "Integrated Automation", "26": "Electrical",
31
+ "27": "Communications", "28": "Electronic Safety and Security", "31": "Earthwork",
32
+ "32": "Exterior Improvements", "33": "Utilities", "34": "Transportation",
33
+ "35": "Waterway and Marine Construction", "40": "Process Interconnections",
34
+ "41": "Material Processing and Handling Equipment",
35
+ "43": "Process Gas and Liquid Handling, Purification, and Storage Equipment",
36
+ "44": "Pollution and Waste Control Equipment", "46": "Water and Wastewater Equipment",
37
+ "48": "Electrical Power Generation",
38
+ }
39
+
40
+ EXAMPLE = "4000 psi concrete slab on grade"
41
+
42
+
43
+ def split_label(label):
44
+ """'03 30 00 Cast-in-Place Concrete' -> ('03 30 00', 'Cast-in-Place Concrete')."""
45
+ code, name = label[:8].strip(), label[8:].strip()
46
+ return code, (name or label)
47
+
48
+
49
+ def division_rollup(preds):
50
+ """Sum level-2 probabilities by the first two digits (division)."""
51
+ totals = {}
52
+ for p in preds:
53
+ div = p["code"][:2]
54
+ totals[div] = totals.get(div, 0.0) + p["score"]
55
+ return [
56
+ {"code": d, "name": DIVISIONS.get(d, d), "score": s}
57
+ for d, s in sorted(totals.items(), key=lambda kv: -kv[1])
58
+ ]
59
+
60
+
61
+ def hf_scorer(model, top_k):
62
+ from transformers import pipeline
63
+
64
+ clf = pipeline("text-classification", model=model, top_k=top_k)
65
+ return lambda texts: [[d for d in out] for out in clf(texts)]
66
+
67
 
68
+ def onnx_scorer(path, top_k):
69
+ """Run onnx/model*.onnx directly. Mirrors scripts/export_onnx.py: inputs
70
+ input_ids / attention_mask / token_type_ids, output logits [batch, num_labels]."""
71
+ import numpy as np
72
+ import onnxruntime as ort
73
+ from pathlib import Path
74
+ from transformers import AutoTokenizer
75
+
76
+ # The labels and tokenizer live beside onnx/ in the repo; fall back to the Hub.
77
+ root = Path(path).resolve().parent.parent
78
+ src = str(root) if (root / "tokenizer.json").exists() or (root / "vocab.txt").exists() else MODEL
79
+ tok = AutoTokenizer.from_pretrained(src)
80
+ id2label = None
81
+ try:
82
+ c = json.loads((root / "config.json").read_text())
83
+ id2label = {int(k): v for k, v in c.get("id2label", {}).items()}
84
+ except Exception:
85
+ id2label = None
86
+ sess = ort.InferenceSession(path, providers=["CPUExecutionProvider"])
87
+ names = {i.name for i in sess.get_inputs()}
88
+
89
+ def score(texts):
90
+ enc = tok(texts, truncation=True, max_length=128, padding=True, return_tensors="np")
91
+ feed = {n: enc[n].astype(np.int64) for n in ("input_ids", "attention_mask", "token_type_ids") if n in names}
92
+ logits = sess.run(None, feed)[0]
93
+ idx = np.argsort(-logits, 1)[:, :top_k]
94
+ shifted = logits - logits.max(1, keepdims=True)
95
+ probs = np.exp(shifted) / np.exp(shifted).sum(1, keepdims=True)
96
+ out = []
97
+ for row, cols in zip(probs, idx):
98
+ out.append([{"label": id2label.get(int(c), str(int(c))), "score": float(row[c])} for c in cols])
99
+ return out
100
+
101
+ return score
102
+
103
+
104
+ def predict(texts, score, top_k, want_div):
105
+ results = []
106
+ for text, preds in zip(texts, score(texts)):
107
+ pl = []
108
  for p in preds:
109
+ code, name = split_label(p["label"])
110
+ pl.append({"code": code, "label": p["label"], "name": name, "score": round(float(p["score"]), 4)})
111
+ item = {"text": text, "predictions": pl}
112
+ if want_div:
113
+ item["divisions"] = [dict(d, score=round(d["score"], 4)) for d in division_rollup(pl)]
114
+ results.append(item)
115
+ return results
116
+
117
+
118
+ def main(argv=None):
119
+ ap = argparse.ArgumentParser(description="MasterFormat level-2 classifier (mf-0.1).")
120
+ ap.add_argument("text", nargs="*", help="text to classify; use '-' to read lines from stdin")
121
+ ap.add_argument("--model", default=MODEL, help="HF repo id or local checkpoint dir")
122
+ ap.add_argument("--onnx", metavar="PATH", help="classify with an ONNX model instead of PyTorch")
123
+ ap.add_argument("--top-k", type=int, default=3, help="number of level-2 predictions to show")
124
+ ap.add_argument("--divisions", action="store_true", help="also show the level-1 (division) roll-up")
125
+ ap.add_argument("--file", help="read newline-separated inputs from a file")
126
+ ap.add_argument("--json", action="store_true", help="emit JSON instead of a table")
127
+ a = ap.parse_args(argv)
128
+
129
+ texts = list(a.text)
130
+ if a.file:
131
+ texts += [l.rstrip("\n") for l in open(a.file, encoding="utf-8") if l.strip()]
132
+ if "-" in texts:
133
+ texts = [t for t in texts if t != "-"] + [l.rstrip("\n") for l in sys.stdin if l.strip()]
134
+ if not texts:
135
+ texts = [EXAMPLE]
136
+ texts = [t for t in texts if t.strip()]
137
+ if not texts:
138
+ ap.error("no input text")
139
+
140
+ score = onnx_scorer(a.onnx, a.top_k) if a.onnx else hf_scorer(a.model, a.top_k)
141
+ results = predict(texts, score, a.top_k, a.divisions)
142
+
143
+ if a.json:
144
+ json.dump(results, sys.stdout, indent=2, ensure_ascii=False)
145
+ print()
146
+ return
147
+ for r in results:
148
+ print(f"\n{r['text']}")
149
+ for p in r["predictions"]:
150
+ print(f" {p['code']} {p['name']:<48.48} {p['score']:.3f}")
151
+ if a.divisions:
152
+ print(" -- divisions --")
153
+ for d in r.get("divisions", []):
154
+ print(f" {d['code']} {d['name']:<48.48} {d['score']:.3f}")
155
 
156
 
157
  if __name__ == "__main__":
requirements.txt ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ # Core inference (PyTorch)
2
+ transformers>=4.40
3
+ torch>=2.0
4
+
5
+ # ONNX path: pip install -r requirements.txt onnxruntime
6
+ # onnxruntime>=1.17
7
+ # numpy>=1.24