Text Classification
Transformers
Joblib
ONNX
Safetensors
English
bert
masterformat
masterformat-classifier
csi-masterformat
construction
construction-technology
sequence-classification
tfidf
ensemble
specs
spec-writing
specifications
takeoff
estimating
cost-code
ufgs
public-domain
Eval Results (legacy)
text-embeddings-inference
Instructions to use constructelligence/masterformat-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use constructelligence/masterformat-classifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="constructelligence/masterformat-classifier")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("constructelligence/masterformat-classifier") model = AutoModelForSequenceClassification.from_pretrained("constructelligence/masterformat-classifier", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Refresh results tables (mf-0.3/mf-0.4 rows, word+char TF-IDF)
Browse files
README.md
CHANGED
|
@@ -127,27 +127,29 @@ the fp32 export is the safer default. Reproduce with `python scripts/export_onnx
|
|
| 127 |
## Results
|
| 128 |
|
| 129 |
All numbers are measured in this project on held-out data. The `line_items`, `manuals` and `manual_sections`
|
| 130 |
-
sets are **fixed**, so those columns are directly comparable for every scorer. `val`
|
| 131 |
-
|
| 132 |
-
were tuned on the v1 split (13,518 units) and are marked *v1*.
|
| 133 |
|
| 134 |
**v2 validation split**
|
| 135 |
|
| 136 |
| Scorer | val top-1 | val top-3 | val division |
|
| 137 |
|---|---:|---:|---:|
|
| 138 |
-
| **mf-0.2 (step 6132)** |
|
|
|
|
|
|
|
| 139 |
| mf-0.1 (step 700) | 0.200 | 0.372 | 0.356 |
|
| 140 |
-
| TF-IDF +
|
| 141 |
|
| 142 |
**Fixed held-out sets (line items 路 manual chunks 路 whole sections)**
|
| 143 |
|
| 144 |
| Scorer | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division |
|
| 145 |
|---|---:|---:|---:|---:|
|
| 146 |
-
| **mf-0.2 (step 6132)** |
|
|
|
|
|
|
|
| 147 |
| mf-0.1 (step 700) | 0.283 | 0.176 | 0.287 | 0.588 |
|
| 148 |
-
| TF-IDF +
|
| 149 |
-
|
|
| 150 |
-
| Ensemble 0.25 embedding + 0.75 TF-IDF *(v1)* | **0.702** | **0.318** | 0.640 | 0.771 |
|
| 151 |
|
| 152 |
`line_items` = 194 hand-labelled estimate line items; `manuals` = 6,890 chunks from 153 sections of two real
|
| 153 |
commercial project manuals; out-of-taxonomy gold labels (e.g. `22 40 00`) count only toward the division
|
|
|
|
| 127 |
## Results
|
| 128 |
|
| 129 |
All numbers are measured in this project on held-out data. The `line_items`, `manuals` and `manual_sections`
|
| 130 |
+
sets are **fixed**, so those columns are directly comparable for every scorer. `val` is the **v2** UFGS split
|
| 131 |
+
(11,598 units).
|
|
|
|
| 132 |
|
| 133 |
**v2 validation split**
|
| 134 |
|
| 135 |
| Scorer | val top-1 | val top-3 | val division |
|
| 136 |
|---|---:|---:|---:|
|
| 137 |
+
| **mf-0.2 (step 6132, published)** | 0.466 | 0.641 | 0.593 |
|
| 138 |
+
| mf-0.3 (v3 data, unfrozen) | **0.556** | **0.710** | **0.662** |
|
| 139 |
+
| mf-0.4 (v2 data, unfrozen) | 0.491 | 0.657 | 0.609 |
|
| 140 |
| mf-0.1 (step 700) | 0.200 | 0.372 | 0.356 |
|
| 141 |
+
| TF-IDF (word + char n-grams) | 0.571 | 0.732 | 0.670 |
|
| 142 |
|
| 143 |
**Fixed held-out sets (line items 路 manual chunks 路 whole sections)**
|
| 144 |
|
| 145 |
| Scorer | line-item top-1 | manual-chunk top-1 | manual-section top-1 | manual-section division |
|
| 146 |
|---|---:|---:|---:|---:|
|
| 147 |
+
| **mf-0.2 (step 6132, published)** | 0.597 | 0.280 | 0.607 | **0.810** |
|
| 148 |
+
| mf-0.3 (v3 data, unfrozen) | 0.597 | 0.259 | 0.547 | 0.765 |
|
| 149 |
+
| mf-0.4 (v2 data, unfrozen) | **0.618** | 0.270 | 0.567 | 0.817 |
|
| 150 |
| mf-0.1 (step 700) | 0.283 | 0.176 | 0.287 | 0.588 |
|
| 151 |
+
| TF-IDF (word + char n-grams) | 0.696 | 0.287 | 0.587 | 0.752 |
|
| 152 |
+
| **Ensemble mf-0.2 + TF-IDF (w = 0.25)** | **0.702** | **0.330** | **0.620** | 0.804 |
|
|
|
|
| 153 |
|
| 154 |
`line_items` = 194 hand-labelled estimate line items; `manuals` = 6,890 chunks from 153 sections of two real
|
| 155 |
commercial project manuals; out-of-taxonomy gold labels (e.g. `22 40 00`) count only toward the division
|