constructelligence commited on
Commit
b774c08
·
verified ·
1 Parent(s): 1d29f6c

Publish mf-0.2: completed GPU fine-tune (val 46.6%, 2.3x mf-0.1) + new model card with metrics/ONNX

Browse files
CITATION.cff CHANGED
@@ -1,5 +1,5 @@
1
  cff-version: 1.2.0
2
- title: "MasterFormat Classifier (mf-0.1)"
3
  message: "If you use this model, please cite it as below."
4
  type: software
5
  authors:
@@ -12,7 +12,7 @@ abstract: >-
12
  A BERT-based text classifier that maps construction line items, specification
13
  paragraphs and section titles to one of 171 MasterFormat level-2 groups across
14
  32 divisions. Fine-tuned from BAAI/bge-small-en-v1.5 on public-domain UFGS
15
- specification text. Published as an early (step-700) checkpoint.
16
  keywords:
17
  - MasterFormat
18
  - construction
 
1
  cff-version: 1.2.0
2
+ title: "MasterFormat Classifier (mf-0.2)"
3
  message: "If you use this model, please cite it as below."
4
  type: software
5
  authors:
 
12
  A BERT-based text classifier that maps construction line items, specification
13
  paragraphs and section titles to one of 171 MasterFormat level-2 groups across
14
  32 divisions. Fine-tuned from BAAI/bge-small-en-v1.5 on public-domain UFGS
15
+ specification text. mf-0.2 is the completed 12-epoch GPU fine-tune.
16
  keywords:
17
  - MasterFormat
18
  - construction
README.md CHANGED
@@ -33,41 +33,42 @@ model-index:
33
  type: text-classification
34
  name: MasterFormat level-2 group classification
35
  dataset:
36
- type: ufgs-val
37
- name: UFGS held-out text units (13,518)
38
  split: validation
39
  metrics:
40
  - type: accuracy
41
- value: 0.184
42
- name: Top-1 accuracy (step-700 checkpoint)
43
  widget:
44
- - text: 4000 psi concrete slab on grade
45
- example_title: Cast-in-place concrete
46
- - text: TPO roofing, 60 mil, fully adhered
47
  example_title: Membrane roofing
48
- - text: Cat 6 data cabling and jacks
49
- example_title: Structured cabling
50
  - text: Wet pipe sprinkler system, light hazard
51
  example_title: Fire suppression
 
 
 
 
52
  ---
53
 
54
- # MasterFormat classifier (mf-0.1) — construction spec & line-item classification
55
 
56
  A BERT text classifier that maps **construction line items, specification paragraphs, submittals and
57
  section titles** to one of **171 MasterFormat level-2 groups** (e.g. `03 30 00 Cast-in-Place Concrete`,
58
  `23 30 00 HVAC Air Distribution`) across **32 divisions**. It is the machine-learning counterpart to a
59
- cost-code lookup table: hand it "TPO roofing, 60 mil, fully adhered" and it returns the MasterFormat group,
60
- ranked. English text, one label per input.
61
 
62
  Built for estimating, takeoff, spec-writing and RAG pipelines that need to tag free text with CSI
63
  MasterFormat codes without a human choosing from a 171-row list.
64
 
65
- > **Early checkpoint, published as-is.** This is **step 700 of a planned 4,150-step fine-tune**
66
- > (`mf-0.1`); training was paused on an 8 GB CPU-only Mac. Top-1 accuracy on the UFGS validation split is
67
- > **18.4 %** (v1 split; **20.0 %** on the rebuilt v2 split) — well below this project's TF-IDF baseline
68
- > (**53.1 %** / **57.5 %**) on the same splits. Treat this release
69
- > as a reproducible starting point and a working deployment format, **not as a finished classifier**. A v2
70
- > dataset and trainer are in the repository; see *Model* and *Limitations*.
 
71
 
72
  MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or
73
  endorsed by them.
@@ -78,19 +79,19 @@ endorsed by them.
78
  from transformers import pipeline
79
 
80
  clf = pipeline("text-classification", model="constructelligence/masterformat-classifier", top_k=3)
81
- for p in clf("4000 psi concrete slab on grade"):
82
  print(p["label"], round(p["score"], 3))
83
- # 04 20 00 Unit Masonry 0.082
84
- # 03 60 00 Grouting 0.076
85
- # 09 30 00 Tiling 0.063 <- not competitive yet; this is the step-700 checkpoint
86
  ```
87
 
88
  Batched, with the level-1 division roll-up and JSON output:
89
 
90
  ```bash
91
  pip install -r requirements.txt # transformers + torch
92
- python predict.py --top-k 5 --divisions "8\" CMU wall, grout filled at 32\" o.c."
93
- echo "Cat 6 data cabling and jacks" | python predict.py -
94
  python predict.py --file items.txt --json > out.json
95
  ```
96
 
@@ -98,7 +99,7 @@ ONNX (no torch, ~34 MB int8):
98
 
99
  ```bash
100
  pip install onnxruntime transformers
101
- python predict.py --onnx onnx/model_quantized.onnx "Panelboards 208/120V 42 circuit"
102
  ```
103
 
104
  The encoder is English-only; inputs should be English.
@@ -107,71 +108,78 @@ The encoder is English-only; inputs should be English.
107
 
108
  `onnx/model.onnx` (fp32, 134 MB) and `onnx/model_quantized.onnx` (dynamic int8, 34 MB) are exported with
109
  dynamic batch and sequence axes, in the layout `transformers.js` / Optimum expect — inputs `input_ids`,
110
- `attention_mask`, `token_type_ids`, output `logits`. The fp32 export matches PyTorch on a 5-text smoke set
111
- (max |Δ logit| 5e-6, argmax agreement 5/5). Dynamic int8 quantization shifts low-confidence logits
112
- (max |Δ| ≈ 1.5 on the same smoke set, argmax agreement 3/5), so verify the quantized model on your own
113
- inputs before relying on it. Reproduce with `python scripts/export_onnx.py runs/mf-0.1`.
114
 
115
  ## Model
116
 
117
  - **Base:** [`BAAI/bge-small-en-v1.5`](https://huggingface.co/BAAI/bge-small-en-v1.5) (BERT, 33M parameters, 384-d, MIT).
118
  - **Head:** 171-way sequence-classification layer; `id2label` / `label2id` are in [`config.json`](config.json).
119
- - **Recipe:** 2 planned epochs over the training set (per-class balance capped at 450 / floored at 300,
120
- estimate-style augmentation), batch 32, max length 96, AdamW, 6 % linear warmup; body LR 6e-5, head LR 2e-3;
121
- embeddings and the first 4 encoder layers frozen.
122
- - **Checkpoint:** stopped at step 700; `train_state.json` records `{"step": 700, "val_acc": 0.18725}`.
123
- - **A stronger v2 recipe already exists** in the source project (`label smoothing 0.05`, sentence-window
124
- dataset, max length 64, lower learning rates) but no `mf-0.2` run has been completed yet.
125
 
126
  ## Results
127
 
128
- All numbers are measured in this project on held-out data. `val` = UFGS text units, held out by a hash of
129
- the normalised source text; `line_items` = 194 hand-labelled estimate line items; `manuals` = 6,890 chunks
130
- from 153 sections of two real commercial project manuals. Out-of-taxonomy gold labels (e.g. `22 40 00`)
131
- count only toward the division score, so a row's top-1 is computed over in-taxonomy items only.
132
 
133
- | Scorer | val top-1 | val top-3 | val division | line-item top-1 | manual-chunk top-1 | manual-section top-1 |
134
- |---|---:|---:|---:|---:|---:|---:|
135
- | **mf-0.1 (step 700)** | **0.184** | 0.345 | 0.321 | 0.283 | 0.175 | 0.287 |
136
- | TF-IDF + linear SGD (full run) | 0.531 | 0.690 | 0.634 | 0.665 | 0.290 | 0.607 |
137
- | bge-small embeddings + logistic regression | 0.349 | 0.517 | 0.458 | 0.576 | 0.256 | 0.593 |
138
- | Ensemble (0.25 embedding + 0.75 TF-IDF) | **0.532** | **0.700** | **0.637** | **0.702** | **0.318** | **0.640** |
139
 
140
- The published checkpoint is the weakest scorer in this table. The best result on the real-world target
141
- (hand-labelled estimate line items) is the **ensemble at 70.2 % top-1**, which is why the project's next
142
- milestone is to finish the transformer run rather than ship `mf-0.1` as the answer.
 
 
143
 
144
- `val` is the only column that depends on the dataset version: `line_items`, `manuals` and `manual_sections`
145
- are fixed held-out sets, so those columns are directly comparable for every scorer. On the rebuilt **v2**
146
- validation split (11,598 units) the same `mf-0.1` checkpoint scores **0.200 top-1 / 0.372 top-3 / 0.356
147
- division**, against **0.575 / 0.732 / 0.675** for TF-IDF on that split. The `val` columns above are on the
148
- original v1 split (13,518 units).
 
 
 
 
 
 
 
 
 
 
 
 
 
149
 
150
  ## Intended use
151
 
152
- - **Use it for:** tagging construction text with a candidate MasterFormat group (top-3 shown), prototyping
153
- an auto-classification step for estimating or spec workflows, and as a reproducible baseline to fine-tune
154
- further.
155
  - **Do not use it for:** unattended production takeoff, bid pricing, code compliance, or anything where a
156
- wrong cost code has financial or contractual consequences. At 18.4 % top-1 it is not accurate enough.
157
- - **Not a substitute for review.** Always keep a human in the loop for classification decisions.
 
158
 
159
  ## Training data and taxonomy
160
 
161
  - **Source:** UFGS (Unified Facilities Guide Specifications) `.SEC` files — US federal works in the
162
  **public domain**. Paragraphs and titles are parsed into labelled text units; boilerplate shared by
163
  multiple sections is dropped, cross-references are stripped so section numbers cannot leak labels, and
164
- long paragraphs are windowed into 20–60-word units.
165
  - **Split:** held out by a hash of the normalised source text (10 %), so a paragraph and its augmentations
166
- stay on the same side. The `mf-0.1` run used 66,396 train / 13,518 val rows; the rebuilt v2 dataset has
167
- 65,496 / 11,598.
168
  - **Taxonomy:** 171 level-2 groups over 32 MasterFormat divisions; group numbers follow the MasterFormat
169
  numbering convention and the short names are this project's own (see `config.json`).
170
 
171
  ## Limitations
172
 
173
- - **Early checkpoint.** The 18.4 % top-1 is a snapshot of an interrupted run; it is not competitive with
174
- the TF-IDF baseline yet. Do not use it for production takeoff without further training and re-evaluation.
175
  - **Domain skew.** UFGS over-represents heavy-civil, water/wastewater and process work relative to commercial
176
  building estimates; the training set is class-balanced, so raw predictions do not reflect building-project
177
  priors.
@@ -179,8 +187,7 @@ original v1 split (13,518 units).
179
  division level.
180
  - **Short, terse line items are the hardest inputs**; division (2-digit) accuracy is consistently higher than
181
  group (6-digit) accuracy.
182
- - **No section-number leakage, but no section-number help either.** The model classifies text content only;
183
- if your input already contains a MasterFormat number, don't ask the model — read the number.
184
 
185
  ## Bias, risks and safety
186
 
@@ -196,9 +203,11 @@ original v1 split (13,518 units).
196
  specifications and cost data into numbered divisions and sections. This model predicts the **level-2 group**
197
  (a 6-digit code such as `03 30 00 Cast-in-Place Concrete`), not the full section number.
198
 
199
- **Can it classify a whole specification section?** Partly. Averaging the model's log-probabilities over a
200
- section's chunks is the intended way to score a section; at step 700 that gets 28.7 % top-1 on two real
201
- project manuals (53 % at division level). The TF-IDF baseline reaches 60.7 %.
 
 
202
 
203
  **Can it read a MasterFormat number out of the text?** No. It classifies the description. If the number is
204
  already present, parse it directly.
@@ -206,17 +215,16 @@ already present, parse it directly.
206
  **Does it work offline / in the browser?** Yes — the ONNX exports are intended for `onnxruntime` and
207
  `transformers.js`.
208
 
209
- **Is a better model available?** Not yet as a published checkpoint. The project's TF-IDF + embedding
210
- **ensemble scores 70.2 % top-1 on hand-labelled line items** and is the current production-candidate; the
211
- next transformer run (`mf-0.2`) is expected to close most of the gap. Constructelligence's proprietary models
212
- are at [constructelligence.co](https://constructelligence.co).
213
 
214
  ## Files
215
 
216
  - `model.safetensors`, `config.json`, `tokenizer.json`, `tokenizer_config.json`, `vocab.txt`,
217
  `special_tokens_map.json` — standard `transformers` checkpoint.
218
- - `train_state.json` — step and validation accuracy of the saved checkpoint (text logs and optimizer state
219
- are not published).
220
  - `predict.py` — CLI example: batching, `--top-k`, `--divisions`, `--json`, stdin/file input, `--onnx`.
221
  - `onnx/` — ONNX fp32 and int8 exports.
222
  - `CITATION.cff`, `requirements.txt`.
@@ -225,7 +233,7 @@ are at [constructelligence.co](https://constructelligence.co).
225
 
226
  ```bibtex
227
  @misc{constructelligence_masterformat_classifier,
228
- title = {MasterFormat Classifier (mf-0.1): construction spec and line-item classification},
229
  author = {Constructelligence},
230
  year = {2026},
231
  howpublished = {\url{https://huggingface.co/constructelligence/masterformat-classifier}},
 
33
  type: text-classification
34
  name: MasterFormat level-2 group classification
35
  dataset:
36
+ type: ufgs-val-v2
37
+ name: UFGS held-out text units (11,598)
38
  split: validation
39
  metrics:
40
  - type: accuracy
41
+ value: 0.466
42
+ name: Top-1 accuracy (mf-0.2)
43
  widget:
44
+ - text: EPDM membrane roofing
 
 
45
  example_title: Membrane roofing
 
 
46
  - text: Wet pipe sprinkler system, light hazard
47
  example_title: Fire suppression
48
+ - text: 8" CMU wall, grout filled at 32" o.c.
49
+ example_title: Unit masonry
50
+ - text: Addressable fire alarm system, devices and panel
51
+ example_title: Fire alarm
52
  ---
53
 
54
+ # MasterFormat classifier (mf-0.2) — construction spec & line-item classification
55
 
56
  A BERT text classifier that maps **construction line items, specification paragraphs, submittals and
57
  section titles** to one of **171 MasterFormat level-2 groups** (e.g. `03 30 00 Cast-in-Place Concrete`,
58
  `23 30 00 HVAC Air Distribution`) across **32 divisions**. It is the machine-learning counterpart to a
59
+ cost-code lookup table: hand it "EPDM membrane roofing" and it returns the MasterFormat group, ranked.
60
+ English text, one label per input.
61
 
62
  Built for estimating, takeoff, spec-writing and RAG pipelines that need to tag free text with CSI
63
  MasterFormat codes without a human choosing from a 171-row list.
64
 
65
+ > **Completed GPU fine-tune.** `mf-0.2` is the full run that `mf-0.1` (step 700) was an early checkpoint of:
66
+ > **6,132 steps / 12 epochs** on a Kaggle T4 over the rebuilt v2 dataset. Top-1 accuracy on the v2 UFGS
67
+ > validation split is **46.6 %** — **2.3×** `mf-0.1` (20.0 %). Line-item accuracy on 194 hand-labelled
68
+ > estimate items is **59.7 %** (was 28.3 %), and whole-section accuracy is **60.7 %** with **81.0 %** at
69
+ > division level. It is still **below this project's TF-IDF + embedding ensemble** on line items (70.2 %) —
70
+ > see *Results* — but it now matches TF-IDF on section-level top-1 and beats every baseline on section-level
71
+ > division accuracy.
72
 
73
  MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or
74
  endorsed by them.
 
79
  from transformers import pipeline
80
 
81
  clf = pipeline("text-classification", model="constructelligence/masterformat-classifier", top_k=3)
82
+ for p in clf("EPDM membrane roofing"):
83
  print(p["label"], round(p["score"], 3))
84
+ # 07 50 00 Membrane Roofing 0.863
85
+ # 07 10 00 Dampproofing and Waterproofing 0.068
86
+ # 07 30 00 Steep Slope Roofing 0.047
87
  ```
88
 
89
  Batched, with the level-1 division roll-up and JSON output:
90
 
91
  ```bash
92
  pip install -r requirements.txt # transformers + torch
93
+ python predict.py --top-k 5 --divisions "Addressable fire alarm system, devices and panel"
94
+ echo "12\" RCP storm drain pipe" | python predict.py -
95
  python predict.py --file items.txt --json > out.json
96
  ```
97
 
 
99
 
100
  ```bash
101
  pip install onnxruntime transformers
102
+ python predict.py --onnx onnx/model_quantized.onnx "Wet pipe sprinkler system, light hazard"
103
  ```
104
 
105
  The encoder is English-only; inputs should be English.
 
108
 
109
  `onnx/model.onnx` (fp32, 134 MB) and `onnx/model_quantized.onnx` (dynamic int8, 34 MB) are exported with
110
  dynamic batch and sequence axes, in the layout `transformers.js` / Optimum expect — inputs `input_ids`,
111
+ `attention_mask`, `token_type_ids`, output `logits`. The fp32 export matches PyTorch on a 4-text smoke set
112
+ (max |Δ logit| 1e-5, argmax agreement 4/4). Dynamic int8 quantization shifts the logits more than it did for
113
+ `mf-0.1` — the better-trained head is more confident — so verify the quantized model on your own inputs;
114
+ the fp32 export is the safer default. Reproduce with `python scripts/export_onnx.py runs/mf-0.2`.
115
 
116
  ## Model
117
 
118
  - **Base:** [`BAAI/bge-small-en-v1.5`](https://huggingface.co/BAAI/bge-small-en-v1.5) (BERT, 33M parameters, 384-d, MIT).
119
  - **Head:** 171-way sequence-classification layer; `id2label` / `label2id` are in [`config.json`](config.json).
120
+ - **Recipe:** `label smoothing 0.05`, max length 64, batch 128, 12 epochs (6,132 steps), fp16 autocast,
121
+ AdamW with 6 % linear warmup; body LR 5e-5, head LR 1e-3; embeddings and the first 4 encoder layers frozen.
122
+ - **Checkpoint:** `train_state.json` records `{"step": 6132, "val_acc": 0.4661}`.
123
+ - **Trained on:** Kaggle T4, ~16.5 min wall-clock, from `BAAI/bge-small-en-v1.5` (not resumed from `mf-0.1`).
 
 
124
 
125
  ## Results
126
 
127
+ All numbers are measured in this project on held-out data. The `line_items`, `manuals` and `manual_sections`
128
+ sets are **fixed**, so those columns are directly comparable for every scorer. `val` depends on the dataset
129
+ version: `mf-0.2` and `mf-0.1` below are on the **v2** split (11,598 units); the embeddings/ensemble rows
130
+ were tuned on the v1 split (13,518 units) and are marked *v1*.
131
 
132
+ **v2 validation split**
 
 
 
 
 
133
 
134
+ | Scorer | val top-1 | val top-3 | val division |
135
+ |---|---:|---:|---:|
136
+ | **mf-0.2 (step 6132)** | **0.466** | **0.641** | **0.593** |
137
+ | mf-0.1 (step 700) | 0.200 | 0.372 | 0.356 |
138
+ | TF-IDF + linear SGD | 0.575 | 0.732 | 0.675 |
139
 
140
+ **Fixed held-out sets (line items · manual chunks · whole sections)**
141
+
142
+ | Scorer | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division |
143
+ |---|---:|---:|---:|---:|
144
+ | **mf-0.2 (step 6132)** | **0.597** | 0.280 | **0.607** | **0.810** |
145
+ | mf-0.1 (step 700) | 0.283 | 0.176 | 0.287 | 0.588 |
146
+ | TF-IDF + linear SGD | 0.665 | **0.290** | **0.607** | 0.732 |
147
+ | bge-small embeddings + logistic regression *(v1)* | 0.576 | 0.256 | 0.593 | 0.725 |
148
+ | Ensemble 0.25 embedding + 0.75 TF-IDF *(v1)* | **0.702** | **0.318** | 0.640 | 0.771 |
149
+
150
+ `line_items` = 194 hand-labelled estimate line items; `manuals` = 6,890 chunks from 153 sections of two real
151
+ commercial project manuals; out-of-taxonomy gold labels (e.g. `22 40 00`) count only toward the division
152
+ score, so top-1 is over in-taxonomy items only.
153
+
154
+ **Reading it:** `mf-0.2` is the best *single transformer* here and the best scorer overall at section-level
155
+ division accuracy (0.810). The TF-IDF + embedding **ensemble still leads on short line items** (0.702 vs
156
+ 0.597), which is the real-world target — so the ensemble remains the production candidate, with `mf-0.2` a
157
+ much stronger transformer baseline than `mf-0.1`.
158
 
159
  ## Intended use
160
 
161
+ - **Use it for:** tagging construction text with a candidate MasterFormat group (top-3 shown), an
162
+ auto-classification step for estimating or spec workflows, and as a strong base to fine-tune or distil.
 
163
  - **Do not use it for:** unattended production takeoff, bid pricing, code compliance, or anything where a
164
+ wrong cost code has financial or contractual consequences. Keep a human in the loop.
165
+ - **Not a substitute for review.** Classifies text content only — if the input already contains a
166
+ MasterFormat number, read the number instead.
167
 
168
  ## Training data and taxonomy
169
 
170
  - **Source:** UFGS (Unified Facilities Guide Specifications) `.SEC` files — US federal works in the
171
  **public domain**. Paragraphs and titles are parsed into labelled text units; boilerplate shared by
172
  multiple sections is dropped, cross-references are stripped so section numbers cannot leak labels, and
173
+ long paragraphs are cut on sentence boundaries into 8–60-word windows.
174
  - **Split:** held out by a hash of the normalised source text (10 %), so a paragraph and its augmentations
175
+ stay on the same side. **65,496 train / 11,598 val** rows (the v2 set).
 
176
  - **Taxonomy:** 171 level-2 groups over 32 MasterFormat divisions; group numbers follow the MasterFormat
177
  numbering convention and the short names are this project's own (see `config.json`).
178
 
179
  ## Limitations
180
 
181
+ - **Below the ensemble on line items.** 59.7 % vs 70.2 % on 194 hand-labelled estimate items; do not deploy
182
+ as the sole classifier.
183
  - **Domain skew.** UFGS over-represents heavy-civil, water/wastewater and process work relative to commercial
184
  building estimates; the training set is class-balanced, so raw predictions do not reflect building-project
185
  priors.
 
187
  division level.
188
  - **Short, terse line items are the hardest inputs**; division (2-digit) accuracy is consistently higher than
189
  group (6-digit) accuracy.
190
+ - **Quantized ONNX drifts.** The int8 export is smaller but less faithful than `mf-0.1`'s; prefer fp32.
 
191
 
192
  ## Bias, risks and safety
193
 
 
203
  specifications and cost data into numbered divisions and sections. This model predicts the **level-2 group**
204
  (a 6-digit code such as `03 30 00 Cast-in-Place Concrete`), not the full section number.
205
 
206
+ **How is this different from `mf-0.1`?** `mf-0.1` was step 700 of an interrupted CPU run; `mf-0.2` is the
207
+ completed 12-epoch GPU fine-tune on the rebuilt dataset. Same architecture, ~2.3× the validation accuracy.
208
+
209
+ **Can it classify a whole specification section?** Yes — average the model's log-probabilities over a
210
+ section's chunks. On two real project manuals that gives 60.7 % top-1 and 81.0 % at division level.
211
 
212
  **Can it read a MasterFormat number out of the text?** No. It classifies the description. If the number is
213
  already present, parse it directly.
 
215
  **Does it work offline / in the browser?** Yes — the ONNX exports are intended for `onnxruntime` and
216
  `transformers.js`.
217
 
218
+ **Is a better model available?** The project's TF-IDF + embedding **ensemble scores 70.2 % top-1 on
219
+ hand-labelled line items** and remains the production candidate. Constructelligence's proprietary models are
220
+ at [constructelligence.co](https://constructelligence.co).
 
221
 
222
  ## Files
223
 
224
  - `model.safetensors`, `config.json`, `tokenizer.json`, `tokenizer_config.json`, `vocab.txt`,
225
  `special_tokens_map.json` — standard `transformers` checkpoint.
226
+ - `train_state.json` — step and validation accuracy of the saved checkpoint.
227
+ - `metrics.json` — full held-out evaluation (`val`, `line_items`, `manuals`, `manual_sections`).
228
  - `predict.py` — CLI example: batching, `--top-k`, `--divisions`, `--json`, stdin/file input, `--onnx`.
229
  - `onnx/` — ONNX fp32 and int8 exports.
230
  - `CITATION.cff`, `requirements.txt`.
 
233
 
234
  ```bibtex
235
  @misc{constructelligence_masterformat_classifier,
236
+ title = {MasterFormat Classifier (mf-0.2): construction spec and line-item classification},
237
  author = {Constructelligence},
238
  year = {2026},
239
  howpublished = {\url{https://huggingface.co/constructelligence/masterformat-classifier}},
config.json CHANGED
@@ -1,10 +1,13 @@
1
  {
 
2
  "architectures": [
3
  "BertForSequenceClassification"
4
  ],
5
  "attention_probs_dropout_prob": 0.1,
 
6
  "classifier_dropout": null,
7
  "dtype": "float32",
 
8
  "hidden_act": "gelu",
9
  "hidden_dropout_prob": 0.1,
10
  "hidden_size": 384,
@@ -183,6 +186,7 @@
183
  },
184
  "initializer_range": 0.02,
185
  "intermediate_size": 1536,
 
186
  "label2id": {
187
  "01 10 00 Summary of Work": 0,
188
  "01 20 00 Price and Payment Procedures": 1,
@@ -363,8 +367,8 @@
363
  "num_hidden_layers": 12,
364
  "pad_token_id": 0,
365
  "position_embedding_type": "absolute",
366
- "problem_type": "single_label_classification",
367
- "transformers_version": "4.57.6",
368
  "type_vocab_size": 2,
369
  "use_cache": true,
370
  "vocab_size": 30522
 
1
  {
2
+ "add_cross_attention": false,
3
  "architectures": [
4
  "BertForSequenceClassification"
5
  ],
6
  "attention_probs_dropout_prob": 0.1,
7
+ "bos_token_id": null,
8
  "classifier_dropout": null,
9
  "dtype": "float32",
10
+ "eos_token_id": null,
11
  "hidden_act": "gelu",
12
  "hidden_dropout_prob": 0.1,
13
  "hidden_size": 384,
 
186
  },
187
  "initializer_range": 0.02,
188
  "intermediate_size": 1536,
189
+ "is_decoder": false,
190
  "label2id": {
191
  "01 10 00 Summary of Work": 0,
192
  "01 20 00 Price and Payment Procedures": 1,
 
367
  "num_hidden_layers": 12,
368
  "pad_token_id": 0,
369
  "position_embedding_type": "absolute",
370
+ "tie_word_embeddings": true,
371
+ "transformers_version": "5.16.1",
372
  "type_vocab_size": 2,
373
  "use_cache": true,
374
  "vocab_size": 30522
metrics.json ADDED
@@ -0,0 +1,25 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "val": {
3
+ "n": 11598,
4
+ "top1": 0.4661148473874806,
5
+ "top3": 0.6408001379548198,
6
+ "division": 0.5929470598379031
7
+ },
8
+ "line_items": {
9
+ "n": 194,
10
+ "top1": 0.5968586387434555,
11
+ "top3": 0.7591623036649214,
12
+ "division": 0.7371134020618557
13
+ },
14
+ "manuals": {
15
+ "n": 6890,
16
+ "top1": 0.2797942689199118,
17
+ "top3": 0.44584864070536373,
18
+ "division": 0.46357039187227866
19
+ },
20
+ "manual_sections": {
21
+ "sections": 153,
22
+ "top1": 0.6066666666666667,
23
+ "division": 0.8104575163398693
24
+ }
25
+ }
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:0770e36bd9b8e8f054ac5750ab0da3078c1e9b4d00ce907a90e9b9f59fa048ca
3
  size 133726636
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:47d0e681abcfa7f5c588035bd451cc963e335d46bc1de91ef505cccc025a0af9
3
  size 133726636
onnx/model.onnx CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:c359f7e6387049e671c1c2039c9e65003f80205183a5702134ebf0b4a39e3fe8
3
  size 133960852
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:10628d545f4b23125421a30ebd11678d7f20911d52c0a7d2e4c893f52c916827
3
  size 133960852
onnx/model_quantized.onnx CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:3096cf891670cf74be258a21c860825dd3e4bd3aa3f4b48bb5e6d716ff4fca5a
3
  size 34067631
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:18a2d9bc236170e7ee6a06f27f7e1a73073d651d6e85dc2f0a8eb998dcd30055
3
  size 34067631
predict.py CHANGED
@@ -8,7 +8,7 @@ Examples
8
  echo "Cat 6 data cabling and jacks" | python predict.py -
9
  python predict.py --file items.txt --json > out.json
10
 
11
- ONNX (no torch, ~4 MB quantized weights):
12
  pip install onnxruntime transformers
13
  python predict.py --onnx onnx/model_quantized.onnx "Panelboards 208/120V 42 circuit"
14
 
@@ -116,7 +116,7 @@ def predict(texts, score, top_k, want_div):
116
 
117
 
118
  def main(argv=None):
119
- ap = argparse.ArgumentParser(description="MasterFormat level-2 classifier (mf-0.1).")
120
  ap.add_argument("text", nargs="*", help="text to classify; use '-' to read lines from stdin")
121
  ap.add_argument("--model", default=MODEL, help="HF repo id or local checkpoint dir")
122
  ap.add_argument("--onnx", metavar="PATH", help="classify with an ONNX model instead of PyTorch")
 
8
  echo "Cat 6 data cabling and jacks" | python predict.py -
9
  python predict.py --file items.txt --json > out.json
10
 
11
+ ONNX (no torch, ~34 MB int8 weights):
12
  pip install onnxruntime transformers
13
  python predict.py --onnx onnx/model_quantized.onnx "Panelboards 208/120V 42 circuit"
14
 
 
116
 
117
 
118
  def main(argv=None):
119
+ ap = argparse.ArgumentParser(description="MasterFormat level-2 classifier (171 groups, 32 divisions).")
120
  ap.add_argument("text", nargs="*", help="text to classify; use '-' to read lines from stdin")
121
  ap.add_argument("--model", default=MODEL, help="HF repo id or local checkpoint dir")
122
  ap.add_argument("--onnx", metavar="PATH", help="classify with an ONNX model instead of PyTorch")
tokenizer.json CHANGED
@@ -2,7 +2,7 @@
2
  "version": "1.0",
3
  "truncation": {
4
  "direction": "Right",
5
- "max_length": 96,
6
  "strategy": "LongestFirst",
7
  "stride": 0
8
  },
 
2
  "version": "1.0",
3
  "truncation": {
4
  "direction": "Right",
5
+ "max_length": 128,
6
  "strategy": "LongestFirst",
7
  "stride": 0
8
  },
tokenizer_config.json CHANGED
@@ -1,51 +1,11 @@
1
  {
2
- "added_tokens_decoder": {
3
- "0": {
4
- "content": "[PAD]",
5
- "lstrip": false,
6
- "normalized": false,
7
- "rstrip": false,
8
- "single_word": false,
9
- "special": true
10
- },
11
- "100": {
12
- "content": "[UNK]",
13
- "lstrip": false,
14
- "normalized": false,
15
- "rstrip": false,
16
- "single_word": false,
17
- "special": true
18
- },
19
- "101": {
20
- "content": "[CLS]",
21
- "lstrip": false,
22
- "normalized": false,
23
- "rstrip": false,
24
- "single_word": false,
25
- "special": true
26
- },
27
- "102": {
28
- "content": "[SEP]",
29
- "lstrip": false,
30
- "normalized": false,
31
- "rstrip": false,
32
- "single_word": false,
33
- "special": true
34
- },
35
- "103": {
36
- "content": "[MASK]",
37
- "lstrip": false,
38
- "normalized": false,
39
- "rstrip": false,
40
- "single_word": false,
41
- "special": true
42
- }
43
- },
44
  "clean_up_tokenization_spaces": true,
45
  "cls_token": "[CLS]",
46
  "do_basic_tokenize": true,
47
  "do_lower_case": true,
48
- "extra_special_tokens": {},
 
49
  "mask_token": "[MASK]",
50
  "model_max_length": 512,
51
  "never_split": null,
 
1
  {
2
+ "backend": "tokenizers",
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  "clean_up_tokenization_spaces": true,
4
  "cls_token": "[CLS]",
5
  "do_basic_tokenize": true,
6
  "do_lower_case": true,
7
+ "is_local": false,
8
+ "local_files_only": false,
9
  "mask_token": "[MASK]",
10
  "model_max_length": 512,
11
  "never_split": null,
train_state.json CHANGED
@@ -1 +1 @@
1
- {"step": 700, "val_acc": 0.18725}
 
1
+ {"step": 6132, "val_acc": 0.4661148473874806}