constructelligence commited on
Commit
9d4c04e
·
verified ·
1 Parent(s): b774c08

Add transformer+TF-IDF ensemble (line-item top-1 0.597 -> 0.686) and document it in the card

Browse files
README.md CHANGED
@@ -16,6 +16,8 @@ tags:
16
  - bert
17
  - onnx
18
  - safetensors
 
 
19
  - specs
20
  - spec-writing
21
  - specifications
@@ -68,7 +70,9 @@ MasterFormat codes without a human choosing from a 171-row list.
68
  > estimate items is **59.7 %** (was 28.3 %), and whole-section accuracy is **60.7 %** with **81.0 %** at
69
  > division level. It is still **below this project's TF-IDF + embedding ensemble** on line items (70.2 %) —
70
  > see *Results* — but it now matches TF-IDF on section-level top-1 and beats every baseline on section-level
71
- > division accuracy.
 
 
72
 
73
  MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or
74
  endorsed by them.
@@ -156,6 +160,38 @@ division accuracy (0.810). The TF-IDF + embedding **ensemble still leads on shor
156
  0.597), which is the real-world target — so the ensemble remains the production candidate, with `mf-0.2` a
157
  much stronger transformer baseline than `mf-0.1`.
158
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
159
  ## Intended use
160
 
161
  - **Use it for:** tagging construction text with a candidate MasterFormat group (top-3 shown), an
@@ -215,9 +251,11 @@ already present, parse it directly.
215
  **Does it work offline / in the browser?** Yes — the ONNX exports are intended for `onnxruntime` and
216
  `transformers.js`.
217
 
218
- **Is a better model available?** The project's TF-IDF + embedding **ensemble scores 70.2 % top-1 on
219
- hand-labelled line items** and remains the production candidate. Constructelligence's proprietary models are
220
- at [constructelligence.co](https://constructelligence.co).
 
 
221
 
222
  ## Files
223
 
@@ -227,6 +265,7 @@ at [constructelligence.co](https://constructelligence.co).
227
  - `metrics.json` — full held-out evaluation (`val`, `line_items`, `manuals`, `manual_sections`).
228
  - `predict.py` — CLI example: batching, `--top-k`, `--divisions`, `--json`, stdin/file input, `--onnx`.
229
  - `onnx/` — ONNX fp32 and int8 exports.
 
230
  - `CITATION.cff`, `requirements.txt`.
231
 
232
  ## Citation
 
16
  - bert
17
  - onnx
18
  - safetensors
19
+ - tfidf
20
+ - ensemble
21
  - specs
22
  - spec-writing
23
  - specifications
 
70
  > estimate items is **59.7 %** (was 28.3 %), and whole-section accuracy is **60.7 %** with **81.0 %** at
71
  > division level. It is still **below this project's TF-IDF + embedding ensemble** on line items (70.2 %) —
72
  > see *Results* — but it now matches TF-IDF on section-level top-1 and beats every baseline on section-level
73
+ > division accuracy. To close the remaining line-item gap the repo also ships a **transformer + TF-IDF
74
+ > ensemble** (`ensemble/`) that lifts line-item top-1 to **68.6 %** and gives the best manual-chunk and
75
+ > section-accuracy results; see *Ensemble*.
76
 
77
  MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or
78
  endorsed by them.
 
160
  0.597), which is the real-world target — so the ensemble remains the production candidate, with `mf-0.2` a
161
  much stronger transformer baseline than `mf-0.1`.
162
 
163
+ ## Ensemble (closing the line-item gap)
164
+
165
+ The transformer is trained on specification prose, so on **short, terse estimate line items** a lexical
166
+ TF-IDF model is still stronger (0.675 vs 0.597 top-1). To close that gap without giving up the transformer's
167
+ long-text accuracy, this repo ships a **log-probability ensemble** of `mf-0.2` and a TF-IDF + SGD scorer:
168
+
169
+ ```
170
+ p = softmax( 0.25 · log_softmax(mf-0.2) + 0.75 · log_softmax(tfidf) )
171
+ ```
172
+
173
+ The 0.25 weight is chosen on the UFGS **validation** split — not on the line-item test set. `ensemble/tfidf.joblib`
174
+ holds the fitted vectorizer (word 1–2 grams, 100k features, `min_df=2`, sublinear tf) and SGD logistic
175
+ classifier; `ensemble/blend.json` records the weight and all metrics; `ensemble/predict_ensemble.py` runs the blend.
176
+
177
+ | Scorer | val top-1 | val division | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division |
178
+ |---|---:|---:|---:|---:|---:|---:|
179
+ | mf-0.2 (transformer) | 0.466 | 0.593 | 0.597 | 0.280 | 0.607 | **0.810** |
180
+ | TF-IDF (100k features) | 0.571 | 0.663 | 0.675 | 0.288 | 0.573 | 0.726 |
181
+ | **Ensemble (w = 0.25)** | **0.586** | **0.688** | **0.686** | **0.334** | **0.613** | 0.784 |
182
+
183
+ ```bash
184
+ pip install transformers torch scikit-learn joblib
185
+ python ensemble/predict_ensemble.py --model constructelligence/masterformat-classifier \
186
+ --tfidf ensemble/tfidf.joblib --top-k 3 --divisions "4000 psi concrete slab on grade"
187
+ ```
188
+
189
+ **Honest caveat.** The ensemble's line-item strength comes from TF-IDF: in a three-way blend with bge-small
190
+ embeddings the transformer receives **zero** weight on line items (the best line-item result is **0.712** at
191
+ 0.70 TF-IDF + 0.30 embeddings). The transformer's own contribution is on **longer text** — it is the best
192
+ scorer at manual-section *division* accuracy (0.810 vs 0.784 for the ensemble). A reasonable production split
193
+ is the transformer for whole sections and short-answer text, the ensemble for terse line items.
194
+
195
  ## Intended use
196
 
197
  - **Use it for:** tagging construction text with a candidate MasterFormat group (top-3 shown), an
 
251
  **Does it work offline / in the browser?** Yes — the ONNX exports are intended for `onnxruntime` and
252
  `transformers.js`.
253
 
254
+ **Is a better model available?** Yes — this repo ships a **transformer + TF-IDF ensemble** (`ensemble/`) that
255
+ scores **68.6 % top-1 on hand-labelled line items** (vs 59.7 % for the transformer alone) and is the best
256
+ scorer on manual chunks. A TF-IDF + bge-small embedding ensemble reaches **71.2 %** but the transformer takes
257
+ no weight there. Constructelligence's proprietary models are at
258
+ [constructelligence.co](https://constructelligence.co).
259
 
260
  ## Files
261
 
 
265
  - `metrics.json` — full held-out evaluation (`val`, `line_items`, `manuals`, `manual_sections`).
266
  - `predict.py` — CLI example: batching, `--top-k`, `--divisions`, `--json`, stdin/file input, `--onnx`.
267
  - `onnx/` — ONNX fp32 and int8 exports.
268
+ - `ensemble/` — `tfidf.joblib`, `blend.json`, `predict_ensemble.py`: the transformer + TF-IDF line-item ensemble.
269
  - `CITATION.cff`, `requirements.txt`.
270
 
271
  ## Citation
ensemble/blend.json ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "weight": 0.25,
3
+ "max_features": 100000,
4
+ "val": {
5
+ "n": 11598,
6
+ "top1": 0.5863942058975685,
7
+ "top3": 0.7549577513364373,
8
+ "division": 0.6883083290222453
9
+ },
10
+ "line_items": {
11
+ "n": 194,
12
+ "top1": 0.6858638743455497,
13
+ "top3": 0.8272251308900523,
14
+ "division": 0.8195876288659794
15
+ },
16
+ "manuals": {
17
+ "n": 6890,
18
+ "top1": 0.33357825128581925,
19
+ "top3": 0.5049228508449669,
20
+ "division": 0.5136429608127722
21
+ },
22
+ "manual_sections": {
23
+ "sections": 153,
24
+ "top1": 0.6133333333333333,
25
+ "division": 0.7843137254901961
26
+ }
27
+ }
ensemble/predict_ensemble.py ADDED
@@ -0,0 +1,143 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Ensemble predictor: the mf-0.2 transformer + a TF-IDF/SGD scorer, blended in log-prob space.
2
+
3
+ The transformer alone reaches 0.597 top-1 on the 194 hand-labelled estimate line items; the fitted TF-IDF
4
+ scorer reaches 0.675; the blend reaches 0.686 (and improves whole-section and manual-chunk accuracy). This
5
+ script reproduces the blend. The weight is chosen on the UFGS validation split (see `blend.json`), not on the
6
+ line-item test set.
7
+
8
+ pip install transformers torch scikit-learn joblib
9
+
10
+ python predict_ensemble.py "EPDM membrane roofing" "4000 psi concrete slab on grade"
11
+ python predict_ensemble.py --top-k 5 --divisions "12\" RCP storm drain pipe"
12
+ python predict_ensemble.py --json --file items.txt > out.json
13
+
14
+ # default TF-IDF/results live beside this script; point --model at the Hub or a local run
15
+ python predict_ensemble.py --model constructelligence/masterformat-classifier \
16
+ --tfidf tfidf.joblib --weight 0.2 "Wet pipe sprinkler system, light hazard"
17
+ """
18
+ import argparse
19
+ import json
20
+ import sys
21
+ from pathlib import Path
22
+
23
+ import numpy as np
24
+
25
+ HERE = Path(__file__).resolve().parent
26
+ MODEL = "constructelligence/masterformat-classifier"
27
+
28
+ DIVISIONS = {
29
+ "01": "General Requirements", "02": "Existing Conditions", "03": "Concrete",
30
+ "04": "Masonry", "05": "Metals", "06": "Wood, Plastics, and Composites",
31
+ "07": "Thermal and Moisture Protection", "08": "Openings", "09": "Finishes",
32
+ "10": "Specialties", "11": "Equipment", "12": "Furnishings",
33
+ "13": "Special Construction", "14": "Conveying Equipment", "21": "Fire Suppression",
34
+ "22": "Plumbing", "23": "HVAC", "25": "Integrated Automation", "26": "Electrical",
35
+ "27": "Communications", "28": "Electronic Safety and Security", "31": "Earthwork",
36
+ "32": "Exterior Improvements", "33": "Utilities", "34": "Transportation",
37
+ "35": "Waterway and Marine Construction", "40": "Process Interconnections",
38
+ "41": "Material Processing and Handling Equipment",
39
+ "43": "Process Gas and Liquid Handling, Purification, and Storage Equipment",
40
+ "44": "Pollution and Waste Control Equipment", "46": "Water and Wastewater Equipment",
41
+ "48": "Electrical Power Generation",
42
+ }
43
+
44
+
45
+ def norm(lp):
46
+ return lp - np.logaddexp.reduce(lp, axis=1, keepdims=True)
47
+
48
+
49
+ class Transformer:
50
+ def __init__(self, model, batch=64):
51
+ import torch
52
+ from transformers import AutoModelForSequenceClassification, AutoTokenizer
53
+ self.torch = torch
54
+ self.tok = AutoTokenizer.from_pretrained(model)
55
+ self.m = AutoModelForSequenceClassification.from_pretrained(model).eval()
56
+ self.labels = {int(k): v for k, v in self.m.config.id2label.items()}
57
+ self.batch = batch
58
+
59
+ def logprobs(self, texts):
60
+ out = []
61
+ with self.torch.no_grad():
62
+ for i in range(0, len(texts), self.batch):
63
+ enc = self.tok(texts[i:i + self.batch], truncation=True, max_length=128,
64
+ padding=True, return_tensors="pt")
65
+ logits = self.m(**enc).logits.float()
66
+ out.append(self.torch.log_softmax(logits, -1).numpy())
67
+ return norm(np.concatenate(out))
68
+
69
+
70
+ class Tfidf:
71
+ def __init__(self, path):
72
+ import joblib
73
+ d = joblib.load(path)
74
+ self.vec = d["vectorizer"]
75
+ self.clf = d["classifier"]
76
+ # classifier classes_ are indices into the sorted label list; verify against the transformer later
77
+ self.classes = list(self.clf.classes_)
78
+
79
+ def logprobs(self, texts):
80
+ return norm(self.clf.predict_log_proba(self.vec.transform(texts)))
81
+
82
+
83
+ def main():
84
+ ap = argparse.ArgumentParser(description="MasterFormat ensemble (mf-0.2 + TF-IDF).")
85
+ ap.add_argument("text", nargs="*", help="text to classify; '-' reads stdin")
86
+ ap.add_argument("--model", default=MODEL, help="HF repo id or local checkpoint dir")
87
+ ap.add_argument("--tfidf", default=str(HERE / "tfidf.joblib"))
88
+ ap.add_argument("--weight", type=float, default=None,
89
+ help="transformer weight in the blend (default: blend.json beside --tfidf)")
90
+ ap.add_argument("--top-k", type=int, default=3)
91
+ ap.add_argument("--divisions", action="store_true")
92
+ ap.add_argument("--file")
93
+ ap.add_argument("--json", action="store_true")
94
+ a = ap.parse_args()
95
+
96
+ if a.weight is None:
97
+ bf = Path(a.tfidf).with_name("blend.json")
98
+ a.weight = json.loads(bf.read_text())["weight"] if bf.exists() else 0.2
99
+ w = a.weight
100
+
101
+ texts = list(a.text)
102
+ if a.file:
103
+ texts += [l.rstrip("\n") for l in open(a.file, encoding="utf-8") if l.strip()]
104
+ if "-" in texts:
105
+ texts = [t for t in texts if t != "-"] + [l.rstrip("\n") for l in sys.stdin if l.strip()]
106
+ texts = [t for t in texts if t.strip()]
107
+ if not texts:
108
+ texts = ["EPDM membrane roofing"]
109
+
110
+ tr, tf = Transformer(a.model), Tfidf(a.tfidf)
111
+ assert tf.classes == list(range(len(tr.labels))), "TF-IDF classes do not align with the transformer labels"
112
+ lp = norm(w * tr.logprobs(texts) + (1 - w) * tf.logprobs(texts)) # blend of log-probs, renormalised
113
+
114
+ results = []
115
+ for text, row in zip(texts, lp):
116
+ order = np.argsort(-row)[:a.top_k]
117
+ preds = [{"code": tr.labels[int(i)][:8].strip(), "label": tr.labels[int(i)],
118
+ "name": tr.labels[int(i)][8:].strip(), "score": round(float(np.exp(row[i])), 4)} for i in order]
119
+ item = {"text": text, "weight_transformer": w, "predictions": preds}
120
+ if a.divisions:
121
+ tot = {}
122
+ for p in preds:
123
+ tot[p["code"][:2]] = tot.get(p["code"][:2], 0.0) + p["score"]
124
+ item["divisions"] = [{"code": d, "name": DIVISIONS.get(d, d), "score": round(s, 4)}
125
+ for d, s in sorted(tot.items(), key=lambda kv: -kv[1])]
126
+ results.append(item)
127
+
128
+ if a.json:
129
+ json.dump(results, sys.stdout, indent=2, ensure_ascii=False)
130
+ print()
131
+ return
132
+ for r in results:
133
+ print(f"\n{r['text']} (blend, transformer weight {w:g})")
134
+ for p in r["predictions"]:
135
+ print(f" {p['code']} {p['name']:<48.48} {p['score']:.3f}")
136
+ if a.divisions:
137
+ print(" -- divisions --")
138
+ for d in r.get("divisions", []):
139
+ print(f" {d['code']} {d['name']:<48.48} {d['score']:.3f}")
140
+
141
+
142
+ if __name__ == "__main__":
143
+ main()
ensemble/tfidf.joblib ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:957b80ab7b61f679513daa017e7ac3732d992bccb5d98b38affbb6808cdd6f2b
3
+ size 62395513
requirements.txt CHANGED
@@ -2,6 +2,10 @@
2
  transformers>=4.40
3
  torch>=2.0
4
 
5
- # ONNX path: pip install -r requirements.txt onnxruntime
6
  # onnxruntime>=1.17
7
  # numpy>=1.24
 
 
 
 
 
2
  transformers>=4.40
3
  torch>=2.0
4
 
5
+ # ONNX path: pip install onnxruntime
6
  # onnxruntime>=1.17
7
  # numpy>=1.24
8
+
9
+ # Ensemble (ensemble/predict_ensemble.py)
10
+ # scikit-learn>=1.3
11
+ # joblib>=1.3