constructelligence commited on
Commit
ad1145a
Β·
verified Β·
1 Parent(s): 9d4c04e

Upgrade ensemble to word+char TF-IDF: line-item 0.686->0.702, section 0.613->0.620 (weights chosen on val)

Browse files
README.md CHANGED
@@ -68,11 +68,9 @@ MasterFormat codes without a human choosing from a 171-row list.
68
  > **6,132 steps / 12 epochs** on a Kaggle T4 over the rebuilt v2 dataset. Top-1 accuracy on the v2 UFGS
69
  > validation split is **46.6 %** β€” **2.3Γ—** `mf-0.1` (20.0 %). Line-item accuracy on 194 hand-labelled
70
  > estimate items is **59.7 %** (was 28.3 %), and whole-section accuracy is **60.7 %** with **81.0 %** at
71
- > division level. It is still **below this project's TF-IDF + embedding ensemble** on line items (70.2 %) β€”
72
- > see *Results* β€” but it now matches TF-IDF on section-level top-1 and beats every baseline on section-level
73
- > division accuracy. To close the remaining line-item gap the repo also ships a **transformer + TF-IDF
74
- > ensemble** (`ensemble/`) that lifts line-item top-1 to **68.6 %** and gives the best manual-chunk and
75
- > section-accuracy results; see *Ensemble*.
76
 
77
  MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or
78
  endorsed by them.
@@ -155,15 +153,14 @@ were tuned on the v1 split (13,518 units) and are marked *v1*.
155
  commercial project manuals; out-of-taxonomy gold labels (e.g. `22 40 00`) count only toward the division
156
  score, so top-1 is over in-taxonomy items only.
157
 
158
- **Reading it:** `mf-0.2` is the best *single transformer* here and the best scorer overall at section-level
159
- division accuracy (0.810). The TF-IDF + embedding **ensemble still leads on short line items** (0.702 vs
160
- 0.597), which is the real-world target β€” so the ensemble remains the production candidate, with `mf-0.2` a
161
- much stronger transformer baseline than `mf-0.1`.
162
 
163
  ## Ensemble (closing the line-item gap)
164
 
165
  The transformer is trained on specification prose, so on **short, terse estimate line items** a lexical
166
- TF-IDF model is still stronger (0.675 vs 0.597 top-1). To close that gap without giving up the transformer's
167
  long-text accuracy, this repo ships a **log-probability ensemble** of `mf-0.2` and a TF-IDF + SGD scorer:
168
 
169
  ```
@@ -171,14 +168,15 @@ p = softmax( 0.25 Β· log_softmax(mf-0.2) + 0.75 Β· log_softmax(tfidf) )
171
  ```
172
 
173
  The 0.25 weight is chosen on the UFGS **validation** split β€” not on the line-item test set. `ensemble/tfidf.joblib`
174
- holds the fitted vectorizer (word 1–2 grams, 100k features, `min_df=2`, sublinear tf) and SGD logistic
175
- classifier; `ensemble/blend.json` records the weight and all metrics; `ensemble/predict_ensemble.py` runs the blend.
 
176
 
177
  | Scorer | val top-1 | val division | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division |
178
  |---|---:|---:|---:|---:|---:|---:|
179
  | mf-0.2 (transformer) | 0.466 | 0.593 | 0.597 | 0.280 | 0.607 | **0.810** |
180
- | TF-IDF (100k features) | 0.571 | 0.663 | 0.675 | 0.288 | 0.573 | 0.726 |
181
- | **Ensemble (w = 0.25)** | **0.586** | **0.688** | **0.686** | **0.334** | **0.613** | 0.784 |
182
 
183
  ```bash
184
  pip install transformers torch scikit-learn joblib
@@ -186,11 +184,13 @@ python ensemble/predict_ensemble.py --model constructelligence/masterformat-clas
186
  --tfidf ensemble/tfidf.joblib --top-k 3 --divisions "4000 psi concrete slab on grade"
187
  ```
188
 
189
- **Honest caveat.** The ensemble's line-item strength comes from TF-IDF: in a three-way blend with bge-small
190
- embeddings the transformer receives **zero** weight on line items (the best line-item result is **0.712** at
191
- 0.70 TF-IDF + 0.30 embeddings). The transformer's own contribution is on **longer text** β€” it is the best
192
- scorer at manual-section *division* accuracy (0.810 vs 0.784 for the ensemble). A reasonable production split
193
- is the transformer for whole sections and short-answer text, the ensemble for terse line items.
 
 
194
 
195
  ## Intended use
196
 
@@ -252,9 +252,8 @@ already present, parse it directly.
252
  `transformers.js`.
253
 
254
  **Is a better model available?** Yes β€” this repo ships a **transformer + TF-IDF ensemble** (`ensemble/`) that
255
- scores **68.6 % top-1 on hand-labelled line items** (vs 59.7 % for the transformer alone) and is the best
256
- scorer on manual chunks. A TF-IDF + bge-small embedding ensemble reaches **71.2 %** but the transformer takes
257
- no weight there. Constructelligence's proprietary models are at
258
  [constructelligence.co](https://constructelligence.co).
259
 
260
  ## Files
@@ -265,7 +264,8 @@ no weight there. Constructelligence's proprietary models are at
265
  - `metrics.json` β€” full held-out evaluation (`val`, `line_items`, `manuals`, `manual_sections`).
266
  - `predict.py` β€” CLI example: batching, `--top-k`, `--divisions`, `--json`, stdin/file input, `--onnx`.
267
  - `onnx/` β€” ONNX fp32 and int8 exports.
268
- - `ensemble/` β€” `tfidf.joblib`, `blend.json`, `predict_ensemble.py`: the transformer + TF-IDF line-item ensemble.
 
269
  - `CITATION.cff`, `requirements.txt`.
270
 
271
  ## Citation
 
68
  > **6,132 steps / 12 epochs** on a Kaggle T4 over the rebuilt v2 dataset. Top-1 accuracy on the v2 UFGS
69
  > validation split is **46.6 %** β€” **2.3Γ—** `mf-0.1` (20.0 %). Line-item accuracy on 194 hand-labelled
70
  > estimate items is **59.7 %** (was 28.3 %), and whole-section accuracy is **60.7 %** with **81.0 %** at
71
+ > division level. Because the transformer is trained on specification prose, short estimate line items are its
72
+ > weak spot; the repo therefore also ships a **transformer + TF-IDF ensemble** (`ensemble/`) that lifts
73
+ > line-item top-1 to **70.2 %** and improves manual-chunk and section accuracy; see *Ensemble*.
 
 
74
 
75
  MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or
76
  endorsed by them.
 
153
  commercial project manuals; out-of-taxonomy gold labels (e.g. `22 40 00`) count only toward the division
154
  score, so top-1 is over in-taxonomy items only.
155
 
156
+ **Reading it:** `mf-0.2` is the best scorer overall at section-level *division* accuracy (0.810). On short
157
+ line items a lexical TF-IDF model is stronger (0.696 alone), so the shipped **transformer + TF-IDF ensemble**
158
+ reaches **0.702** and is the production candidate; see *Ensemble*.
 
159
 
160
  ## Ensemble (closing the line-item gap)
161
 
162
  The transformer is trained on specification prose, so on **short, terse estimate line items** a lexical
163
+ TF-IDF model is still stronger (0.696 vs 0.597 top-1). To close that gap without giving up the transformer's
164
  long-text accuracy, this repo ships a **log-probability ensemble** of `mf-0.2` and a TF-IDF + SGD scorer:
165
 
166
  ```
 
168
  ```
169
 
170
  The 0.25 weight is chosen on the UFGS **validation** split β€” not on the line-item test set. `ensemble/tfidf.joblib`
171
+ holds the fitted vectorizer (word 1–2 grams, 100k features, plus `char_wb` 3–5 grams, 150k features, `min_df`
172
+ 2/3, sublinear tf) and SGD logistic classifier; `ensemble/blend.json` records the weight and all metrics;
173
+ `ensemble/predict_ensemble.py` runs the blend.
174
 
175
  | Scorer | val top-1 | val division | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division |
176
  |---|---:|---:|---:|---:|---:|---:|
177
  | mf-0.2 (transformer) | 0.466 | 0.593 | 0.597 | 0.280 | 0.607 | **0.810** |
178
+ | TF-IDF (word + char n-grams) | 0.571 | 0.670 | 0.696 | 0.287 | 0.587 | 0.752 |
179
+ | **Ensemble (w = 0.25)** | **0.588** | **0.689** | **0.702** | **0.330** | **0.620** | 0.804 |
180
 
181
  ```bash
182
  pip install transformers torch scikit-learn joblib
 
184
  --tfidf ensemble/tfidf.joblib --top-k 3 --divisions "4000 psi concrete slab on grade"
185
  ```
186
 
187
+ **Honest caveat.** The ensemble beats TF-IDF alone mainly on the transformer's strengths β€” validation top-1
188
+ (0.588 vs 0.571), manual-chunk top-1 (0.330 vs 0.287) and whole-section top-1 (0.620 vs 0.587) β€” while line
189
+ items are nearly saturated by TF-IDF (0.702 vs 0.696). Adding bge-small embeddings does **not** help once the
190
+ blend weight is chosen on `val` (its weight goes to ~0); an earlier 0.712 line-item figure came from tuning
191
+ the weights on the 194-item test itself and does not generalise. The transformer remains the best scorer at
192
+ section *division* accuracy (0.810). A reasonable production split is the transformer for whole sections, the
193
+ ensemble for terse line items.
194
 
195
  ## Intended use
196
 
 
252
  `transformers.js`.
253
 
254
  **Is a better model available?** Yes β€” this repo ships a **transformer + TF-IDF ensemble** (`ensemble/`) that
255
+ scores **70.2 % top-1 on hand-labelled line items** (vs 59.7 % for the transformer alone) and improves
256
+ manual-chunk and whole-section accuracy. Constructelligence's proprietary models are at
 
257
  [constructelligence.co](https://constructelligence.co).
258
 
259
  ## Files
 
264
  - `metrics.json` β€” full held-out evaluation (`val`, `line_items`, `manuals`, `manual_sections`).
265
  - `predict.py` β€” CLI example: batching, `--top-k`, `--divisions`, `--json`, stdin/file input, `--onnx`.
266
  - `onnx/` β€” ONNX fp32 and int8 exports.
267
+ - `ensemble/` β€” `tfidf.joblib` (word + char n-grams, ~112 MB), `blend.json`, `predict_ensemble.py`: the
268
+ transformer + TF-IDF line-item ensemble.
269
  - `CITATION.cff`, `requirements.txt`.
270
 
271
  ## Citation
ensemble/blend.json CHANGED
@@ -1,27 +1,28 @@
1
  {
2
  "weight": 0.25,
3
- "max_features": 100000,
 
4
  "val": {
5
  "n": 11598,
6
- "top1": 0.5863942058975685,
7
- "top3": 0.7549577513364373,
8
- "division": 0.6883083290222453
9
  },
10
  "line_items": {
11
  "n": 194,
12
- "top1": 0.6858638743455497,
13
- "top3": 0.8272251308900523,
14
  "division": 0.8195876288659794
15
  },
16
  "manuals": {
17
  "n": 6890,
18
- "top1": 0.33357825128581925,
19
- "top3": 0.5049228508449669,
20
- "division": 0.5136429608127722
21
  },
22
  "manual_sections": {
23
  "sections": 153,
24
- "top1": 0.6133333333333333,
25
- "division": 0.7843137254901961
26
  }
27
  }
 
1
  {
2
  "weight": 0.25,
3
+ "config": "word 1-2 (100k) + char_wb 3-5 (150k), SGD alpha 2e-6",
4
+ "features": 190130,
5
  "val": {
6
  "n": 11598,
7
+ "top1": 0.5875150888084153,
8
+ "top3": 0.7515088808415245,
9
+ "division": 0.6892567684083463
10
  },
11
  "line_items": {
12
  "n": 194,
13
+ "top1": 0.7015706806282722,
14
+ "top3": 0.8219895287958116,
15
  "division": 0.8195876288659794
16
  },
17
  "manuals": {
18
  "n": 6890,
19
+ "top1": 0.3301983835415136,
20
+ "top3": 0.4981631153563556,
21
+ "division": 0.5132075471698113
22
  },
23
  "manual_sections": {
24
  "sections": 153,
25
+ "top1": 0.62,
26
+ "division": 0.803921568627451
27
  }
28
  }
ensemble/predict_ensemble.py CHANGED
@@ -1,9 +1,9 @@
1
  """Ensemble predictor: the mf-0.2 transformer + a TF-IDF/SGD scorer, blended in log-prob space.
2
 
3
  The transformer alone reaches 0.597 top-1 on the 194 hand-labelled estimate line items; the fitted TF-IDF
4
- scorer reaches 0.675; the blend reaches 0.686 (and improves whole-section and manual-chunk accuracy). This
5
- script reproduces the blend. The weight is chosen on the UFGS validation split (see `blend.json`), not on the
6
- line-item test set.
7
 
8
  pip install transformers torch scikit-learn joblib
9
 
@@ -71,12 +71,16 @@ class Tfidf:
71
  def __init__(self, path):
72
  import joblib
73
  d = joblib.load(path)
74
- self.vec = d["vectorizer"]
75
- self.clf = d["classifier"]
 
 
76
  # classifier classes_ are indices into the sorted label list; verify against the transformer later
77
  self.classes = list(self.clf.classes_)
78
 
79
  def logprobs(self, texts):
 
 
80
  return norm(self.clf.predict_log_proba(self.vec.transform(texts)))
81
 
82
 
 
1
  """Ensemble predictor: the mf-0.2 transformer + a TF-IDF/SGD scorer, blended in log-prob space.
2
 
3
  The transformer alone reaches 0.597 top-1 on the 194 hand-labelled estimate line items; the fitted TF-IDF
4
+ scorer (word + char n-grams) reaches 0.675; the blend reaches 0.702 (and improves whole-section and
5
+ manual-chunk accuracy). This script reproduces the blend. The weight is chosen on the UFGS validation split
6
+ (see `blend.json`), not on the line-item test set.
7
 
8
  pip install transformers torch scikit-learn joblib
9
 
 
71
  def __init__(self, path):
72
  import joblib
73
  d = joblib.load(path)
74
+ if isinstance(d, dict): # legacy {vectorizer, classifier}
75
+ self.pipe, self.vec, self.clf = None, d["vectorizer"], d["classifier"]
76
+ else: # sklearn Pipeline
77
+ self.pipe, self.vec, self.clf = d, None, d.named_steps["clf"]
78
  # classifier classes_ are indices into the sorted label list; verify against the transformer later
79
  self.classes = list(self.clf.classes_)
80
 
81
  def logprobs(self, texts):
82
+ if self.pipe is not None:
83
+ return norm(self.pipe.predict_log_proba(texts))
84
  return norm(self.clf.predict_log_proba(self.vec.transform(texts)))
85
 
86
 
ensemble/tfidf.joblib CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:957b80ab7b61f679513daa017e7ac3732d992bccb5d98b38affbb6808cdd6f2b
3
- size 62395513
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ea810e3d763fbc5b26d7faf48f9e3af798aa5c7d94203e72f8979b21b69e6eca
3
+ size 112200382