constructelligence's picture
Add evidence caveat, calibration/abstention (91% precision at 50% coverage) and improvement roadmap
e43c8ca verified
|
Raw History Blame Contribute Delete
16.5 kB
metadata
language:
  - en
license: mit
library_name: transformers
pipeline_tag: text-classification
base_model: BAAI/bge-small-en-v1.5
tags:
  - masterformat
  - masterformat-classifier
  - csi-masterformat
  - construction
  - construction-technology
  - text-classification
  - sequence-classification
  - bert
  - onnx
  - safetensors
  - tfidf
  - ensemble
  - specs
  - spec-writing
  - specifications
  - takeoff
  - estimating
  - cost-code
  - ufgs
  - public-domain
metrics:
  - accuracy
model-index:
  - name: masterformat-classifier
    results:
      - task:
          type: text-classification
          name: MasterFormat level-2 group classification
        dataset:
          type: ufgs-val-v2
          name: UFGS held-out text units (11,598)
          split: validation
        metrics:
          - type: accuracy
            value: 0.466
            name: Top-1 accuracy (mf-0.2)
widget:
  - text: EPDM membrane roofing
    example_title: Membrane roofing
  - text: Wet pipe sprinkler system, light hazard
    example_title: Fire suppression
  - text: 8" CMU wall, grout filled at 32" o.c.
    example_title: Unit masonry
  - text: Addressable fire alarm system, devices and panel
    example_title: Fire alarm

MasterFormat classifier (mf-0.2) β€” construction spec & line-item classification

A BERT text classifier that maps construction line items, specification paragraphs, submittals and section titles to one of 171 MasterFormat level-2 groups (e.g. 03 30 00 Cast-in-Place Concrete, 23 30 00 HVAC Air Distribution) across 32 divisions. It is the machine-learning counterpart to a cost-code lookup table: hand it "EPDM membrane roofing" and it returns the MasterFormat group, ranked. English text, one label per input.

Built for estimating, takeoff, spec-writing and RAG pipelines that need to tag free text with CSI MasterFormat codes without a human choosing from a 171-row list.

Completed GPU fine-tune. mf-0.2 is the full run that mf-0.1 (step 700) was an early checkpoint of: 6,132 steps / 12 epochs on a Kaggle T4 over the rebuilt v2 dataset. Top-1 accuracy on the v2 UFGS validation split is 46.6 % β€” 2.3Γ— mf-0.1 (20.0 %). Line-item accuracy on 194 hand-labelled estimate items is 59.7 % (was 28.3 %), and whole-section accuracy is 60.7 % with 81.0 % at division level. Because the transformer is trained on specification prose, short estimate line items are its weak spot; the repo therefore also ships a transformer + TF-IDF ensemble (ensemble/) that lifts line-item top-1 to 70.2 % and improves manual-chunk and section accuracy; see Ensemble.

MasterFormat is a registered trademark of CSI / CSC. This project is independent and not affiliated with or endorsed by them.

Quick start

from transformers import pipeline

clf = pipeline("text-classification", model="constructelligence/masterformat-classifier", top_k=3)
for p in clf("EPDM membrane roofing"):
    print(p["label"], round(p["score"], 3))
# 07 50 00 Membrane Roofing 0.863
# 07 10 00 Dampproofing and Waterproofing 0.068
# 07 30 00 Steep Slope Roofing 0.047

Batched, with the level-1 division roll-up and JSON output:

pip install -r requirements.txt        # transformers + torch
python predict.py --top-k 5 --divisions "Addressable fire alarm system, devices and panel"
echo "12\" RCP storm drain pipe" | python predict.py -
python predict.py --file items.txt --json > out.json

ONNX (no torch, ~34 MB int8):

pip install onnxruntime transformers
python predict.py --onnx onnx/model_quantized.onnx "Wet pipe sprinkler system, light hazard"

The encoder is English-only; inputs should be English.

ONNX exports

onnx/model.onnx (fp32, 134 MB) and onnx/model_quantized.onnx (dynamic int8, 34 MB) are exported with dynamic batch and sequence axes, in the layout transformers.js / Optimum expect β€” inputs input_ids, attention_mask, token_type_ids, output logits. The fp32 export matches PyTorch on a 4-text smoke set (max |Ξ” logit| 1e-5, argmax agreement 4/4). Dynamic int8 quantization shifts the logits more than it did for mf-0.1 β€” the better-trained head is more confident β€” so verify the quantized model on your own inputs; the fp32 export is the safer default. Reproduce with python scripts/export_onnx.py runs/mf-0.2.

Model

  • Base: BAAI/bge-small-en-v1.5 (BERT, 33M parameters, 384-d, MIT).
  • Head: 171-way sequence-classification layer; id2label / label2id are in config.json.
  • Recipe: label smoothing 0.05, max length 64, batch 128, 12 epochs (6,132 steps), fp16 autocast, AdamW with 6 % linear warmup; body LR 5e-5, head LR 1e-3; embeddings and the first 4 encoder layers frozen.
  • Checkpoint: train_state.json records {"step": 6132, "val_acc": 0.4661}.
  • Trained on: Kaggle T4, ~16.5 min wall-clock, from BAAI/bge-small-en-v1.5 (not resumed from mf-0.1).

Results

All numbers are measured in this project on held-out data. The line_items, manuals and manual_sections sets are fixed, so those columns are directly comparable for every scorer. val is the v2 UFGS split (11,598 units).

v2 validation split

Scorer val top-1 val top-3 val division
mf-0.2 (step 6132, published) 0.466 0.641 0.593
mf-0.3 (v3 data, unfrozen) 0.556 0.710 0.662
mf-0.4 (v2 data, unfrozen) 0.491 0.657 0.609
mf-0.1 (step 700) 0.200 0.372 0.356
TF-IDF (word + char n-grams) 0.571 0.732 0.670

Fixed held-out sets (line items Β· manual chunks Β· whole sections)

Scorer line-item top-1 manual-chunk top-1 manual-section top-1 manual-section division
mf-0.2 (step 6132, published) 0.597 0.280 0.607 0.810
mf-0.3 (v3 data, unfrozen) 0.597 0.259 0.547 0.765
mf-0.4 (v2 data, unfrozen) 0.618 0.270 0.567 0.817
mf-0.1 (step 700) 0.283 0.176 0.287 0.588
TF-IDF (word + char n-grams) 0.696 0.287 0.587 0.752
Ensemble mf-0.2 + TF-IDF (w = 0.25) 0.702 0.330 0.620 0.804

line_items = 194 hand-labelled estimate line items; manuals = 6,890 chunks from 153 sections of two real commercial project manuals; out-of-taxonomy gold labels (e.g. 22 40 00) count only toward the division score, so top-1 is over in-taxonomy items only.

Evidence caveat. The line-item and section sets are small. At n = 191 the 95 % Wilson interval on the line-item top-1 is [0.633, 0.762] (Β±6.4 pts), and whole-section is Β±~8 pts β€” so differences of a few points between models here are noise. Treat the val and manual-chunk rows (Β±1 pt) as the measurable ones, and the line-item/section rows as indicative. Enlarging these eval sets is the top item on the roadmap.

Reading it: mf-0.2 is the best scorer overall at section-level division accuracy (0.810). On short line items a lexical TF-IDF model is stronger (0.696 alone), so the shipped transformer + TF-IDF ensemble reaches 0.702 and is the production candidate; see Ensemble.

Ensemble (closing the line-item gap)

The transformer is trained on specification prose, so on short, terse estimate line items a lexical TF-IDF model is still stronger (0.696 vs 0.597 top-1). To close that gap without giving up the transformer's long-text accuracy, this repo ships a log-probability ensemble of mf-0.2 and a TF-IDF + SGD scorer:

p = softmax( 0.25 Β· log_softmax(mf-0.2)  +  0.75 Β· log_softmax(tfidf) )

The 0.25 weight is chosen on the UFGS validation split β€” not on the line-item test set. ensemble/tfidf.joblib holds the fitted vectorizer (word 1–2 grams, 100k features, plus char_wb 3–5 grams, 150k features, min_df 2/3, sublinear tf) and SGD logistic classifier; ensemble/blend.json records the weight and all metrics; ensemble/predict_ensemble.py runs the blend.

Scorer val top-1 val division line-item top-1 manual-chunk top-1 manual-section top-1 manual-section division
mf-0.2 (transformer) 0.466 0.593 0.597 0.280 0.607 0.810
TF-IDF (word + char n-grams) 0.571 0.670 0.696 0.287 0.587 0.752
Ensemble (w = 0.25) 0.588 0.689 0.702 0.330 0.620 0.804
pip install transformers torch scikit-learn joblib
python ensemble/predict_ensemble.py --model constructelligence/masterformat-classifier \
    --tfidf ensemble/tfidf.joblib --top-k 3 --divisions "4000 psi concrete slab on grade"

Honest caveat. The ensemble beats TF-IDF alone mainly on the transformer's strengths β€” validation top-1 (0.588 vs 0.571), manual-chunk top-1 (0.330 vs 0.287) and whole-section top-1 (0.620 vs 0.587) β€” while line items are nearly saturated by TF-IDF (0.702 vs 0.696). Adding bge-small embeddings does not help once the blend weight is chosen on val (its weight goes to ~0); an earlier 0.712 line-item figure came from tuning the weights on the 194-item test itself and does not generalise. The transformer remains the best scorer at section division accuracy (0.810). A reasonable production split is the transformer for whole sections, the ensemble for terse line items.

Calibration and abstention

The ensemble is already close to calibrated; a single temperature T = 0.90 (fit on val) lowers the expected calibration error on val from 0.060 to 0.022. More useful in production: send the lowest-confidence predictions to a human instead of filing them.

Coverage β†’ precision on the 194 line items (predictions sorted by confidence):

auto-filed (coverage) 100 % 90 % 80 % 70 % 50 % 30 %
precision 0.69 0.75 0.81 0.85 0.91 0.97

Routing the least-confident 30 % to review lifts precision to 91 %; routing 50 % gives 97 %. So the model is usable today as an assist that flags its own uncertainty, even though it is not accurate enough to file unattended. Reproduce with python cloud/stats.py.

Roadmap

Further gains are evidence-bound, not architecture-bound (see the repo's IMPROVEMENT-PLAN.md):

  1. Enlarge the eval (β‰₯1,000 line items, β‰₯400 sections) so changes are measurable at Β±3 pts β€” currently line-item changes below ~10 pts cannot be validated.
  2. Real line-item training data (public bid tabulations, agency item catalogs) β€” the transformer's line-item ceiling is a data problem; synthetic augmentation was tested and did not transfer (mf-0.3).
  3. Domain-adaptive pretraining on construction prose beyond UFGS, then retrain.
  4. Distillation + hierarchical (division β†’ group) head to beat the TF-IDF blend with a single model.
  5. Level-3 / full-section codes and an explicit out-of-taxonomy fallback.

Intended use

  • Use it for: tagging construction text with a candidate MasterFormat group (top-3 shown), an auto-classification step for estimating or spec workflows, and as a strong base to fine-tune or distil.
  • Do not use it for: unattended production takeoff, bid pricing, code compliance, or anything where a wrong cost code has financial or contractual consequences. Keep a human in the loop.
  • Not a substitute for review. Classifies text content only β€” if the input already contains a MasterFormat number, read the number instead.

Training data and taxonomy

  • Source: UFGS (Unified Facilities Guide Specifications) .SEC files β€” US federal works in the public domain. Paragraphs and titles are parsed into labelled text units; boilerplate shared by multiple sections is dropped, cross-references are stripped so section numbers cannot leak labels, and long paragraphs are cut on sentence boundaries into 8–60-word windows.
  • Split: held out by a hash of the normalised source text (10 %), so a paragraph and its augmentations stay on the same side. 65,496 train / 11,598 val rows (the v2 set).
  • Taxonomy: 171 level-2 groups over 32 MasterFormat divisions; group numbers follow the MasterFormat numbering convention and the short names are this project's own (see config.json).

Limitations

  • Below the ensemble on line items. 59.7 % vs 70.2 % on 194 hand-labelled estimate items; do not deploy as the sole classifier.
  • Domain skew. UFGS over-represents heavy-civil, water/wastewater and process work relative to commercial building estimates; the training set is class-balanced, so raw predictions do not reflect building-project priors.
  • Out-of-taxonomy inputs (sections whose level-2 group is not among the 171) can only be scored at division level.
  • Short, terse line items are the hardest inputs; division (2-digit) accuracy is consistently higher than group (6-digit) accuracy.
  • Quantized ONNX drifts. The int8 export is smaller but less faithful than mf-0.1's; prefer fp32.

Bias, risks and safety

  • Estimating bias. Class-balanced training over a public-domain federal corpus does not represent any particular firm's cost structure or regional practice. Do not treat output as a standard or an authority.
  • Trademark. MasterFormat is a registered trademark of CSI / CSC; this model is not endorsed by them and its group names are the project's own short descriptions, not CSI's official titles.
  • Privacy. The model runs locally; no input text leaves your machine unless you call a hosted endpoint.

FAQ

What is MasterFormat? The CSI/CSC MasterFormat is the North American standard for organising construction specifications and cost data into numbered divisions and sections. This model predicts the level-2 group (a 6-digit code such as 03 30 00 Cast-in-Place Concrete), not the full section number.

How is this different from mf-0.1? mf-0.1 was step 700 of an interrupted CPU run; mf-0.2 is the completed 12-epoch GPU fine-tune on the rebuilt dataset. Same architecture, ~2.3Γ— the validation accuracy.

Can it classify a whole specification section? Yes β€” average the model's log-probabilities over a section's chunks. On two real project manuals that gives 60.7 % top-1 and 81.0 % at division level.

Can it read a MasterFormat number out of the text? No. It classifies the description. If the number is already present, parse it directly.

Does it work offline / in the browser? Yes β€” the ONNX exports are intended for onnxruntime and transformers.js.

Is a better model available? Yes β€” this repo ships a transformer + TF-IDF ensemble (ensemble/) that scores 70.2 % top-1 on hand-labelled line items (vs 59.7 % for the transformer alone) and improves manual-chunk and whole-section accuracy. Constructelligence's proprietary models are at constructelligence.co.

Files

  • model.safetensors, config.json, tokenizer.json, tokenizer_config.json, vocab.txt, special_tokens_map.json β€” standard transformers checkpoint.
  • train_state.json β€” step and validation accuracy of the saved checkpoint.
  • metrics.json β€” full held-out evaluation (val, line_items, manuals, manual_sections).
  • predict.py β€” CLI example: batching, --top-k, --divisions, --json, stdin/file input, --onnx.
  • onnx/ β€” ONNX fp32 and int8 exports.
  • ensemble/ β€” tfidf.joblib (word + char n-grams, ~112 MB), blend.json, predict_ensemble.py: the transformer + TF-IDF line-item ensemble.
  • IMPROVEMENT-PLAN.md β€” the prioritised roadmap (enlarged eval, data, domain-adaptive pretraining, distillation).
  • CITATION.cff, requirements.txt.

Citation

@misc{constructelligence_masterformat_classifier,
  title        = {MasterFormat Classifier (mf-0.2): construction spec and line-item classification},
  author       = {Constructelligence},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/constructelligence/masterformat-classifier}},
  note         = {Fine-tuned from BAAI/bge-small-en-v1.5; 171-way MasterFormat level-2 classifier}
}

Licence and attribution

Released under the MIT licence, matching the base model. UFGS source text is public domain. MasterFormat is a registered trademark of CSI / CSC; this project is not affiliated with or endorsed by them. Constructelligence's production models are available at constructelligence.co.