Text Classification
Transformers
Joblib
ONNX
Safetensors
English
bert
masterformat
masterformat-classifier
csi-masterformat
construction
construction-technology
sequence-classification
tfidf
ensemble
specs
spec-writing
specifications
takeoff
estimating
cost-code
ufgs
public-domain
Eval Results (legacy)
text-embeddings-inference
Instructions to use constructelligence/masterformat-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use constructelligence/masterformat-classifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="constructelligence/masterformat-classifier")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("constructelligence/masterformat-classifier") model = AutoModelForSequenceClassification.from_pretrained("constructelligence/masterformat-classifier", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Upgrade ensemble to word+char TF-IDF: line-item 0.686->0.702, section 0.613->0.620 (weights chosen on val)
Browse files- README.md +23 -23
- ensemble/blend.json +12 -11
- ensemble/predict_ensemble.py +9 -5
- ensemble/tfidf.joblib +2 -2
README.md
CHANGED
|
@@ -68,11 +68,9 @@ MasterFormat codes without a human choosing from a 171-row list.
|
|
| 68 |
> **6,132 steps / 12 epochs** on a Kaggle T4 over the rebuilt v2 dataset. Top-1 accuracy on the v2 UFGS
|
| 69 |
> validation split is **46.6 %** β **2.3Γ** `mf-0.1` (20.0 %). Line-item accuracy on 194 hand-labelled
|
| 70 |
> estimate items is **59.7 %** (was 28.3 %), and whole-section accuracy is **60.7 %** with **81.0 %** at
|
| 71 |
-
> division level.
|
| 72 |
-
>
|
| 73 |
-
>
|
| 74 |
-
> ensemble** (`ensemble/`) that lifts line-item top-1 to **68.6 %** and gives the best manual-chunk and
|
| 75 |
-
> section-accuracy results; see *Ensemble*.
|
| 76 |
|
| 77 |
MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or
|
| 78 |
endorsed by them.
|
|
@@ -155,15 +153,14 @@ were tuned on the v1 split (13,518 units) and are marked *v1*.
|
|
| 155 |
commercial project manuals; out-of-taxonomy gold labels (e.g. `22 40 00`) count only toward the division
|
| 156 |
score, so top-1 is over in-taxonomy items only.
|
| 157 |
|
| 158 |
-
**Reading it:** `mf-0.2` is the best
|
| 159 |
-
|
| 160 |
-
0.
|
| 161 |
-
much stronger transformer baseline than `mf-0.1`.
|
| 162 |
|
| 163 |
## Ensemble (closing the line-item gap)
|
| 164 |
|
| 165 |
The transformer is trained on specification prose, so on **short, terse estimate line items** a lexical
|
| 166 |
-
TF-IDF model is still stronger (0.
|
| 167 |
long-text accuracy, this repo ships a **log-probability ensemble** of `mf-0.2` and a TF-IDF + SGD scorer:
|
| 168 |
|
| 169 |
```
|
|
@@ -171,14 +168,15 @@ p = softmax( 0.25 Β· log_softmax(mf-0.2) + 0.75 Β· log_softmax(tfidf) )
|
|
| 171 |
```
|
| 172 |
|
| 173 |
The 0.25 weight is chosen on the UFGS **validation** split β not on the line-item test set. `ensemble/tfidf.joblib`
|
| 174 |
-
holds the fitted vectorizer (word 1β2 grams, 100k features, `
|
| 175 |
-
classifier; `ensemble/blend.json` records the weight and all metrics;
|
|
|
|
| 176 |
|
| 177 |
| Scorer | val top-1 | val division | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division |
|
| 178 |
|---|---:|---:|---:|---:|---:|---:|
|
| 179 |
| mf-0.2 (transformer) | 0.466 | 0.593 | 0.597 | 0.280 | 0.607 | **0.810** |
|
| 180 |
-
| TF-IDF (
|
| 181 |
-
| **Ensemble (w = 0.25)** | **0.
|
| 182 |
|
| 183 |
```bash
|
| 184 |
pip install transformers torch scikit-learn joblib
|
|
@@ -186,11 +184,13 @@ python ensemble/predict_ensemble.py --model constructelligence/masterformat-clas
|
|
| 186 |
--tfidf ensemble/tfidf.joblib --top-k 3 --divisions "4000 psi concrete slab on grade"
|
| 187 |
```
|
| 188 |
|
| 189 |
-
**Honest caveat.** The ensemble
|
| 190 |
-
|
| 191 |
-
|
| 192 |
-
|
| 193 |
-
|
|
|
|
|
|
|
| 194 |
|
| 195 |
## Intended use
|
| 196 |
|
|
@@ -252,9 +252,8 @@ already present, parse it directly.
|
|
| 252 |
`transformers.js`.
|
| 253 |
|
| 254 |
**Is a better model available?** Yes β this repo ships a **transformer + TF-IDF ensemble** (`ensemble/`) that
|
| 255 |
-
scores **
|
| 256 |
-
|
| 257 |
-
no weight there. Constructelligence's proprietary models are at
|
| 258 |
[constructelligence.co](https://constructelligence.co).
|
| 259 |
|
| 260 |
## Files
|
|
@@ -265,7 +264,8 @@ no weight there. Constructelligence's proprietary models are at
|
|
| 265 |
- `metrics.json` β full held-out evaluation (`val`, `line_items`, `manuals`, `manual_sections`).
|
| 266 |
- `predict.py` β CLI example: batching, `--top-k`, `--divisions`, `--json`, stdin/file input, `--onnx`.
|
| 267 |
- `onnx/` β ONNX fp32 and int8 exports.
|
| 268 |
-
- `ensemble/` β `tfidf.joblib`, `blend.json`, `predict_ensemble.py`: the
|
|
|
|
| 269 |
- `CITATION.cff`, `requirements.txt`.
|
| 270 |
|
| 271 |
## Citation
|
|
|
|
| 68 |
> **6,132 steps / 12 epochs** on a Kaggle T4 over the rebuilt v2 dataset. Top-1 accuracy on the v2 UFGS
|
| 69 |
> validation split is **46.6 %** β **2.3Γ** `mf-0.1` (20.0 %). Line-item accuracy on 194 hand-labelled
|
| 70 |
> estimate items is **59.7 %** (was 28.3 %), and whole-section accuracy is **60.7 %** with **81.0 %** at
|
| 71 |
+
> division level. Because the transformer is trained on specification prose, short estimate line items are its
|
| 72 |
+
> weak spot; the repo therefore also ships a **transformer + TF-IDF ensemble** (`ensemble/`) that lifts
|
| 73 |
+
> line-item top-1 to **70.2 %** and improves manual-chunk and section accuracy; see *Ensemble*.
|
|
|
|
|
|
|
| 74 |
|
| 75 |
MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or
|
| 76 |
endorsed by them.
|
|
|
|
| 153 |
commercial project manuals; out-of-taxonomy gold labels (e.g. `22 40 00`) count only toward the division
|
| 154 |
score, so top-1 is over in-taxonomy items only.
|
| 155 |
|
| 156 |
+
**Reading it:** `mf-0.2` is the best scorer overall at section-level *division* accuracy (0.810). On short
|
| 157 |
+
line items a lexical TF-IDF model is stronger (0.696 alone), so the shipped **transformer + TF-IDF ensemble**
|
| 158 |
+
reaches **0.702** and is the production candidate; see *Ensemble*.
|
|
|
|
| 159 |
|
| 160 |
## Ensemble (closing the line-item gap)
|
| 161 |
|
| 162 |
The transformer is trained on specification prose, so on **short, terse estimate line items** a lexical
|
| 163 |
+
TF-IDF model is still stronger (0.696 vs 0.597 top-1). To close that gap without giving up the transformer's
|
| 164 |
long-text accuracy, this repo ships a **log-probability ensemble** of `mf-0.2` and a TF-IDF + SGD scorer:
|
| 165 |
|
| 166 |
```
|
|
|
|
| 168 |
```
|
| 169 |
|
| 170 |
The 0.25 weight is chosen on the UFGS **validation** split β not on the line-item test set. `ensemble/tfidf.joblib`
|
| 171 |
+
holds the fitted vectorizer (word 1β2 grams, 100k features, plus `char_wb` 3β5 grams, 150k features, `min_df`
|
| 172 |
+
2/3, sublinear tf) and SGD logistic classifier; `ensemble/blend.json` records the weight and all metrics;
|
| 173 |
+
`ensemble/predict_ensemble.py` runs the blend.
|
| 174 |
|
| 175 |
| Scorer | val top-1 | val division | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division |
|
| 176 |
|---|---:|---:|---:|---:|---:|---:|
|
| 177 |
| mf-0.2 (transformer) | 0.466 | 0.593 | 0.597 | 0.280 | 0.607 | **0.810** |
|
| 178 |
+
| TF-IDF (word + char n-grams) | 0.571 | 0.670 | 0.696 | 0.287 | 0.587 | 0.752 |
|
| 179 |
+
| **Ensemble (w = 0.25)** | **0.588** | **0.689** | **0.702** | **0.330** | **0.620** | 0.804 |
|
| 180 |
|
| 181 |
```bash
|
| 182 |
pip install transformers torch scikit-learn joblib
|
|
|
|
| 184 |
--tfidf ensemble/tfidf.joblib --top-k 3 --divisions "4000 psi concrete slab on grade"
|
| 185 |
```
|
| 186 |
|
| 187 |
+
**Honest caveat.** The ensemble beats TF-IDF alone mainly on the transformer's strengths β validation top-1
|
| 188 |
+
(0.588 vs 0.571), manual-chunk top-1 (0.330 vs 0.287) and whole-section top-1 (0.620 vs 0.587) β while line
|
| 189 |
+
items are nearly saturated by TF-IDF (0.702 vs 0.696). Adding bge-small embeddings does **not** help once the
|
| 190 |
+
blend weight is chosen on `val` (its weight goes to ~0); an earlier 0.712 line-item figure came from tuning
|
| 191 |
+
the weights on the 194-item test itself and does not generalise. The transformer remains the best scorer at
|
| 192 |
+
section *division* accuracy (0.810). A reasonable production split is the transformer for whole sections, the
|
| 193 |
+
ensemble for terse line items.
|
| 194 |
|
| 195 |
## Intended use
|
| 196 |
|
|
|
|
| 252 |
`transformers.js`.
|
| 253 |
|
| 254 |
**Is a better model available?** Yes β this repo ships a **transformer + TF-IDF ensemble** (`ensemble/`) that
|
| 255 |
+
scores **70.2 % top-1 on hand-labelled line items** (vs 59.7 % for the transformer alone) and improves
|
| 256 |
+
manual-chunk and whole-section accuracy. Constructelligence's proprietary models are at
|
|
|
|
| 257 |
[constructelligence.co](https://constructelligence.co).
|
| 258 |
|
| 259 |
## Files
|
|
|
|
| 264 |
- `metrics.json` β full held-out evaluation (`val`, `line_items`, `manuals`, `manual_sections`).
|
| 265 |
- `predict.py` β CLI example: batching, `--top-k`, `--divisions`, `--json`, stdin/file input, `--onnx`.
|
| 266 |
- `onnx/` β ONNX fp32 and int8 exports.
|
| 267 |
+
- `ensemble/` β `tfidf.joblib` (word + char n-grams, ~112 MB), `blend.json`, `predict_ensemble.py`: the
|
| 268 |
+
transformer + TF-IDF line-item ensemble.
|
| 269 |
- `CITATION.cff`, `requirements.txt`.
|
| 270 |
|
| 271 |
## Citation
|
ensemble/blend.json
CHANGED
|
@@ -1,27 +1,28 @@
|
|
| 1 |
{
|
| 2 |
"weight": 0.25,
|
| 3 |
-
"
|
|
|
|
| 4 |
"val": {
|
| 5 |
"n": 11598,
|
| 6 |
-
"top1": 0.
|
| 7 |
-
"top3": 0.
|
| 8 |
-
"division": 0.
|
| 9 |
},
|
| 10 |
"line_items": {
|
| 11 |
"n": 194,
|
| 12 |
-
"top1": 0.
|
| 13 |
-
"top3": 0.
|
| 14 |
"division": 0.8195876288659794
|
| 15 |
},
|
| 16 |
"manuals": {
|
| 17 |
"n": 6890,
|
| 18 |
-
"top1": 0.
|
| 19 |
-
"top3": 0.
|
| 20 |
-
"division": 0.
|
| 21 |
},
|
| 22 |
"manual_sections": {
|
| 23 |
"sections": 153,
|
| 24 |
-
"top1": 0.
|
| 25 |
-
"division": 0.
|
| 26 |
}
|
| 27 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"weight": 0.25,
|
| 3 |
+
"config": "word 1-2 (100k) + char_wb 3-5 (150k), SGD alpha 2e-6",
|
| 4 |
+
"features": 190130,
|
| 5 |
"val": {
|
| 6 |
"n": 11598,
|
| 7 |
+
"top1": 0.5875150888084153,
|
| 8 |
+
"top3": 0.7515088808415245,
|
| 9 |
+
"division": 0.6892567684083463
|
| 10 |
},
|
| 11 |
"line_items": {
|
| 12 |
"n": 194,
|
| 13 |
+
"top1": 0.7015706806282722,
|
| 14 |
+
"top3": 0.8219895287958116,
|
| 15 |
"division": 0.8195876288659794
|
| 16 |
},
|
| 17 |
"manuals": {
|
| 18 |
"n": 6890,
|
| 19 |
+
"top1": 0.3301983835415136,
|
| 20 |
+
"top3": 0.4981631153563556,
|
| 21 |
+
"division": 0.5132075471698113
|
| 22 |
},
|
| 23 |
"manual_sections": {
|
| 24 |
"sections": 153,
|
| 25 |
+
"top1": 0.62,
|
| 26 |
+
"division": 0.803921568627451
|
| 27 |
}
|
| 28 |
}
|
ensemble/predict_ensemble.py
CHANGED
|
@@ -1,9 +1,9 @@
|
|
| 1 |
"""Ensemble predictor: the mf-0.2 transformer + a TF-IDF/SGD scorer, blended in log-prob space.
|
| 2 |
|
| 3 |
The transformer alone reaches 0.597 top-1 on the 194 hand-labelled estimate line items; the fitted TF-IDF
|
| 4 |
-
scorer reaches 0.675; the blend reaches 0.
|
| 5 |
-
script reproduces the blend. The weight is chosen on the UFGS validation split
|
| 6 |
-
line-item test set.
|
| 7 |
|
| 8 |
pip install transformers torch scikit-learn joblib
|
| 9 |
|
|
@@ -71,12 +71,16 @@ class Tfidf:
|
|
| 71 |
def __init__(self, path):
|
| 72 |
import joblib
|
| 73 |
d = joblib.load(path)
|
| 74 |
-
|
| 75 |
-
|
|
|
|
|
|
|
| 76 |
# classifier classes_ are indices into the sorted label list; verify against the transformer later
|
| 77 |
self.classes = list(self.clf.classes_)
|
| 78 |
|
| 79 |
def logprobs(self, texts):
|
|
|
|
|
|
|
| 80 |
return norm(self.clf.predict_log_proba(self.vec.transform(texts)))
|
| 81 |
|
| 82 |
|
|
|
|
| 1 |
"""Ensemble predictor: the mf-0.2 transformer + a TF-IDF/SGD scorer, blended in log-prob space.
|
| 2 |
|
| 3 |
The transformer alone reaches 0.597 top-1 on the 194 hand-labelled estimate line items; the fitted TF-IDF
|
| 4 |
+
scorer (word + char n-grams) reaches 0.675; the blend reaches 0.702 (and improves whole-section and
|
| 5 |
+
manual-chunk accuracy). This script reproduces the blend. The weight is chosen on the UFGS validation split
|
| 6 |
+
(see `blend.json`), not on the line-item test set.
|
| 7 |
|
| 8 |
pip install transformers torch scikit-learn joblib
|
| 9 |
|
|
|
|
| 71 |
def __init__(self, path):
|
| 72 |
import joblib
|
| 73 |
d = joblib.load(path)
|
| 74 |
+
if isinstance(d, dict): # legacy {vectorizer, classifier}
|
| 75 |
+
self.pipe, self.vec, self.clf = None, d["vectorizer"], d["classifier"]
|
| 76 |
+
else: # sklearn Pipeline
|
| 77 |
+
self.pipe, self.vec, self.clf = d, None, d.named_steps["clf"]
|
| 78 |
# classifier classes_ are indices into the sorted label list; verify against the transformer later
|
| 79 |
self.classes = list(self.clf.classes_)
|
| 80 |
|
| 81 |
def logprobs(self, texts):
|
| 82 |
+
if self.pipe is not None:
|
| 83 |
+
return norm(self.pipe.predict_log_proba(texts))
|
| 84 |
return norm(self.clf.predict_log_proba(self.vec.transform(texts)))
|
| 85 |
|
| 86 |
|
ensemble/tfidf.joblib
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ea810e3d763fbc5b26d7faf48f9e3af798aa5c7d94203e72f8979b21b69e6eca
|
| 3 |
+
size 112200382
|