Image-to-Text
English
custom
ocr
document-ai
information-extraction
nameplate
nameplate-ocr
equipment-nameplate
data-center
data-centre
datacenter
mission-critical
mep
electrical
equipment-schedule
schedule-verification
asset-register
commissioning
quality-assurance
ups
pdu
switchgear
generator
ats
transformer
busway
battery
crah
crac
chiller
cooling
construction
aec
edge-ai
on-device
privacy-preserving
tesseract
browser
|
Download README.md from constructelligence/data-centre-nameplates: direct link, hf CLI and curl.
- Browser
- Download file 14.4 kB
-
https://huggingface.co/constructelligence/data-centre-nameplates/resolve/main/README.md
- Command line
-
hf download hf://constructelligence/data-centre-nameplates/README.md
-
curl -L -o README.md https://huggingface.co/constructelligence/data-centre-nameplates/resolve/main/README.md
14.4 kB
| license: apache-2.0 | |
| language: | |
| - en | |
| library_name: custom | |
| pipeline_tag: image-to-text | |
| tags: | |
| - ocr | |
| - document-ai | |
| - information-extraction | |
| - nameplate | |
| - nameplate-ocr | |
| - equipment-nameplate | |
| - data-center | |
| - data-centre | |
| - datacenter | |
| - mission-critical | |
| - mep | |
| - electrical | |
| - equipment-schedule | |
| - schedule-verification | |
| - asset-register | |
| - commissioning | |
| - quality-assurance | |
| - ups | |
| - pdu | |
| - switchgear | |
| - generator | |
| - ats | |
| - transformer | |
| - busway | |
| - battery | |
| - crah | |
| - crac | |
| - chiller | |
| - cooling | |
| - construction | |
| - aec | |
| - edge-ai | |
| - on-device | |
| - privacy-preserving | |
| - tesseract | |
| - browser | |
| # Data Centre Nameplates — equipment-nameplate OCR checked against the schedule | |
| **Read an equipment nameplate and know if it is the unit the schedule asked for.** Point a phone at the | |
| nameplate of a UPS, PDU, switchgear, generator, transformer, busway, battery, CRAH or chiller; the model returns | |
| the plate's fields — manufacturer, model, serial, voltage, phase, Hz, kVA, kW, amps, MCA, MOCP, refrigerant, | |
| manufacture date — and **checks each one against the equipment schedule**, flagging a mismatch instead of a | |
| silent pass. It builds an asset register and exports it to CSV. | |
| Built for the part of a data-centre job where the gear *is* the job. The expensive mistakes are a unit | |
| delivered at the wrong voltage or rating, or one that quietly never makes it onto the asset register. | |
| > **Runs on the device.** OCR and matching run in the browser (or Node) with no server and no upload. A | |
| > nameplate photo of live infrastructure never leaves the phone that took it. | |
| **▶ Try it live:** [huggingface.co/spaces/constructelligence/data-centre-nameplates](https://huggingface.co/spaces/constructelligence/data-centre-nameplates) — a browser demo with a sample schedule and a sample plate. | |
| | | | | |
| |---|---| | |
| | **Task** | Image → structured nameplate fields → schedule verification | | |
| | **Fields** | 13 (manufacturer, model, serial, voltage, phase, Hz, kVA, kW, amps, MCA, MOCP, refrigerant, mfg date) | | |
| | **Equipment** | 13 types (UPS, generator, ATS/STS, switchgear, transformer, PDU/RPP, panelboard, busway, battery, CDU, chiller, CRAH/CRAC, pump) | | |
| | **OCR** | Tesseract.js 5 (English LSTM), pretrained — no fine-tuning | | |
| | **Extraction** | Deterministic labelled-field parser with an OCR error model | | |
| | **Runtime** | Browser · Node · edge — no GPU, no server, no data egress | | |
| | **Licence** | Apache-2.0 | | |
| --- | |
| ## Why this exists | |
| General OCR reads the text on a nameplate. A data-centre commissioning or QA team needs the text to answer a | |
| narrower question: **is this the right unit, and does it match what was specified?** That means three things a | |
| plain OCR call does not do: | |
| 1. **Turn plate text into fields, not paragraphs.** "INPUT: 480Y/277 VAC 3 PH 60 HZ" becomes | |
| `voltage=480Y/277 V`, `phase=3`, `hz=60`; "RATING: 750 KVA / 675 KW" becomes `kva=750`, `kw=675`. | |
| 2. **Survive the way OCR misreads plate lettering.** Serifed `V`→`Vv`, `K`→`X` (`XVA`), `HZ`→`Hw`, codes split | |
| after their punctuation (`NPX- 750- 480`). The parser repairs these before matching. | |
| 3. **Compare to the schedule the way codes actually differ.** Model and serial codes are compared with `O/0`, | |
| `I/L/1`, `S/5`, `B/8`, `Z/2` folded and punctuation ignored, so a misread is not a false mismatch, while a | |
| genuine voltage variant (`NPX-750-415` vs `NPX-750-480`) still fails. | |
| The result is a pass/mismatch verdict per plate, with the fields that disagree named. | |
| ## How it works | |
| ``` | |
| photo ──► OCR (Tesseract.js) ──► field parser ──► schedule matcher ──► verdict + asset register CSV | |
| │ │ │ | |
| lines + confidence labelled-value rules code folding, tolerance checks | |
| ``` | |
| 1. **OCR** (`tesseract.js@5`, English) returns lines with a per-line confidence. | |
| 2. **Field parser** scans lines for labelled values (`MODEL`, `S/N`, `MVA`, `MCA`, `MOCP`, `FLA`, `VOLTS`, | |
| `REFRIGERANT`, `MFG DATE`), reading voltages as `480Y/277`, `13.8 kV`, `208/120`, and rejecting look-alikes | |
| (`480/277` with no voltage word is not a voltage; `R-513A` and a charge weight are not amps). Lines OCR | |
| scored near zero (brushed metal, background texture) are dropped before parsing. | |
| 3. **Equipment typing** recognises the family from plate wording (UPS, genset, ATS, switchgear, transformer, | |
| PDU/RPP, panelboard, busway, battery, CDU, chiller, CRAH/CRAC, pump). | |
| 4. **Schedule matcher** ranks candidate tags — a serial match wins outright; otherwise model similarity | |
| dominates, with manufacturer and equipment type as tie-breakers. Data halls repeat, so among equal matches | |
| the first **unscanned** tag is offered. | |
| 5. **Field checks** compare the reading to the chosen row and return `ok`, `mismatch` or `unread` per field, and | |
| an overall status: `verified`, `partial`, `mismatch`, `unmatched` or `missing` (scheduled but not scanned). | |
| ## Fields extracted | |
| | Field | Example | Notes | | |
| |---|---|---| | |
| | `manufacturer` | Northline Power Systems | ~44 data-centre makers recognised as whole words; otherwise the company-looking line | | |
| | `model` | NPX-750-480 | `MODEL`, `MOD.`, `M/N`, `CAT. NO.`, `P/N`, `PART NO.` | | |
| | `serial` | NL26A01937 | `SERIAL`, `SER.`, `S/N`, `5/N`, `SN` | | |
| | `voltage` | 480Y/277 V | every voltage on the plate (transformers and UPSes carry more than one) | | |
| | `phase` | 3 | `1`/`3`, `PH`, `PHASE`, `Ø` | | |
| | `hz` | 60 | 50 or 60, `HZ`/`FREQUENCY` | | |
| | `kva` | 750 | thousands separators handled (`3,750`); European decimals too | | |
| | `kw` | 675 | `KW` / `EKW` | | |
| | `amps` | 1002 | labelled (`FLA`, `RLA`, `CURRENT`, `AMPS`) first | | |
| | `mca` | 612 | minimum circuit ampacity | | |
| | `mocp` | 800 | `MOCP` / `MAX FUSE` / `MAX BREAKER` | | |
| | `refrigerant` | R-513A | R-410A, 134A, 1234ZE, 1234YF, 513A, 514A, 1233ZD, 454B, 407C, 32, 22 | | |
| | `mfgDate` | 03/2026 | `MFG`, `MANUFACTURED`, `MFD`, `BUILD`, `PRODUCTION` | | |
| ## Equipment types recognised | |
| `ups` · `generator` · `ats` (also STS) · `switchgear` · `transformer` · `pdu` (also RPP) · `panelboard` · | |
| `busway` · `battery` · `cdu` · `chiller` · `crah` (also CRAC) · `pump` | |
| Manufacturers recognised include Schneider Electric, Square D, APC, Eaton, ABB, Siemens, Vertiv, Liebert, | |
| Caterpillar, Cummins, Kohler, Rolls-Royce, MTU, Generac, Trane, Carrier, Daikin, York, Johnson Controls, | |
| Stulz, Munters, nVent, Legrand, Mitsubishi Electric, Toshiba, General Electric, Hubbell, Socomec, Piller, | |
| Starline, Raritan, Server Technology, Emerson, Russelectric, ASCO, Hitachi, Riello, Rittal, Motivair, CoolIT, | |
| Hitec, EnerSys, Leclanché and Samsung SDI. | |
| ## Results — OCR and extraction, end to end | |
| Six drawn plates (UPS, diesel genset, chiller, dry-type transformer, PDU, CRAH), each pushed through the | |
| degradations a phone photo actually has, then run through the app's own OCR + parser and scored **field by | |
| field**. **56 checks per condition** (six plates × their fields), 13 conditions. | |
| | Condition | Field accuracy | ms/plate | | |
| |---|:---:|---:| | |
| | clean | **100%** (56/56) | 1236 | | |
| | small (640 px) | **100%** (56/56) | 897 | | |
| | blur | **100%** (56/56) | 540 | | |
| | noise | **100%** (56/56) | 748 | | |
| | low contrast | **100%** (56/56) | 557 | | |
| | glare | **100%** (56/56) | 532 | | |
| | tilt (7°) | **100%** (56/56) | 612 | | |
| | keystone | **100%** (56/56) | 474 | | |
| | dark anodised plate | **100%** (56/56) | 808 | | |
| | JPEG | **100%** (56/56) | 524 | | |
| | plate small in a cluttered scene | **100%** (56/56) | 895 | | |
| | sideways | **100%** (56/56) | 554 | | |
| | phone (blur + glare + noise + tilt + scene + JPEG) | **98.2%** (55/56) | 1607 | | |
| | **Overall** | **99.9% (727/728)** | **768** | | |
| The single miss is one `hz` value on the hardest "phone" composite. The benchmark measures what **OCR** loses, | |
| not parser coverage: the ground truth is the parser's own reading of the clean plate text, so every reduction | |
| is an OCR failure the parser could not repair. Reproduce it with `node tests/nameplates-bench.mjs`. | |
| **Read this honestly.** These are crisp, machine-drawn plates. A real nameplate photographed on a live unit is | |
| harder: cast shadows, embossed lettering, print over brushed metal, a plate half out of frame. Expect the | |
| degradation suite, not the clean row, and treat every low-confidence field as one a person must confirm. | |
| ## Fine-tuning — the higher-accuracy server model | |
| The engine above is the on-device path. Alongside it, a **Donut** model | |
| ([`naver-clova-ix/donut-base`](https://huggingface.co/naver-clova-ix/donut-base)) is fine-tuned to read a plate | |
| **straight into fields** (`<s_nameplate><s_type>ups</s_type><s_manufacturer>…`) for a server / endpoint where | |
| the browser engine falls short. Training data is synthetic plates with real-photo degradations, plus corrected | |
| real plates from the field. Once trained, the weights are published in a sibling repository and linked here. | |
| Nothing on this page claims the accuracy of that model yet — it is the roadmap, not a result. | |
| ## Intended use | |
| - **Commissioning and QA walks:** photograph each unit as it is installed; confirm the delivered unit matches | |
| the equipment schedule; capture an asset register as you go. | |
| - **Receiving and delivery checks:** flag the unit that arrived at the wrong voltage or rating before it is set. | |
| - **Register completion:** track scheduled tags that have no plate scanned against them yet. | |
| - **Takeoff and submittal review:** pull model/rating fields off a plate photo into a spreadsheet. | |
| ### Out of scope — do not use it for | |
| - **Safety, code compliance or energisation decisions.** It reads a plate; it does not verify that a unit is | |
| safe to energise, correctly protected, or code-compliant. | |
| - **An as-built or a legal record on its own.** OCR misreads happen; low-confidence fields are highlighted for a | |
| person to check, and any field that disagrees with the schedule must be confirmed by eye. | |
| - **Serial-number evidence in a dispute** without human verification. Two plates can share a misread serial; | |
| the register flags duplicate serials but cannot tell a repeated plate from a copied number. | |
| ## Limitations | |
| - **English nameplates**, Latin script. Non-Latin or handwritten plates are out of scope. | |
| - **Glare, blur and angle degrade OCR.** The parser repairs common misreads, but a plate that is small in the | |
| frame, badly glared or very low contrast will lose fields (those fields come back `unread`, not invented). | |
| - **Deterministic parser, not a language model.** It reports only what it can match to a rule; it will not | |
| infer a missing digit. | |
| - **The schedule must be structured** (CSV with a Tag or Model column). Free-form schedule PDFs are not parsed | |
| here. | |
| ## Quickstart | |
| The engine is plain ES modules with no build step. `nameplate-model.js` is the extractor and verifier; | |
| `schedule-import.js` is the CSV reader it depends on. | |
| ```js | |
| import { parseNameplate, parseEquipmentSchedule, candidatesFor, checkAgainst, statusOf } from './nameplate-model.js'; | |
| // 1. Parse plate text (from Tesseract, or typed by hand) | |
| const reading = parseNameplate(`NORTHLINE POWER SYSTEMS | |
| UNINTERRUPTIBLE POWER SUPPLY | |
| MODEL NO: NPX-750-480 | |
| SERIAL NO: NL26A01937 | |
| INPUT: 480Y/277 VAC 3 PH 60 HZ | |
| RATING: 750 KVA / 675 KW | |
| INPUT CURRENT: 1002 A | |
| MFG DATE: 03/2026`); | |
| console.log(reading.type); // 'ups' | |
| console.log(reading.fields.voltage); // { value: '480Y/277 V', values: [480, 277], conf, line } | |
| // 2. Match against the equipment schedule and check each field | |
| const { items } = parseEquipmentSchedule(scheduleCsv); | |
| const [best] = candidatesFor(reading, items); | |
| const checks = checkAgainst(reading, best.item); | |
| console.log(statusOf(checks, best.item.tag)); // 'verified' | 'partial' | 'mismatch' | ... | |
| ``` | |
| In a browser, feed Tesseract's `data.lines` (`[{ text, confidence, bbox }]`) straight into `parseNameplate` | |
| to keep OCR confidence and the plate line each value came from. | |
| ## Evaluation | |
| - **Extraction/verification logic:** 12 unit tests over the parser, voltage reader, code folding, schedule | |
| parsing, candidate ranking and mismatch detection — all passing (`node --test tests/nameplate-model.test.mjs`). | |
| - **OCR end-to-end:** a field-by-field benchmark draws plates for six equipment types and puts each through the | |
| degradations real phone photos have — small, blurred, noisy, low-contrast, glare, tilt, keystone, dark | |
| anodised plate, JPEG, plate-small-in-scene and sideways — then scores the app's own OCR + parsing pipeline | |
| field by field. Run it with `node tests/nameplates-bench.mjs`. | |
| ## Repository contents | |
| | File | What it is | | |
| |---|---| | |
| | `nameplate-model.js` | the extractor, matcher and register builder (ES module) | | |
| | `schedule-import.js` | the CSV schedule reader it depends on | | |
| | `progress-project.js` | a helper the schedule reader imports | | |
| | `sample_schedule.csv` | a 10-row data-centre equipment schedule | | |
| | `sample_plates.json` | two sample plates: one that verifies, one flagged on model and voltage | | |
| | `eval/` | the OCR benchmark harness output | | |
| | `config.json` | machine-readable fields, types and statuses | | |
| ## Provenance | |
| An open extraction and verification engine from Constructelligence, the same code path that powers the | |
| **Nameplates** mode of BuildVision. No customer data is used; sample plates use fictional makers and models. | |
| It is the document-AI sibling of the vision models | |
| [`construction-site-safety-hazards`](https://huggingface.co/constructelligence/construction-site-safety-hazards) | |
| and [`electrical-circuit-connectivity`](https://huggingface.co/constructelligence/electrical-circuit-connectivity). | |
| ## Safety | |
| This is a **prompt to look**, not a finding. It is not safety-rated, not a compliance decision, and not a | |
| substitute for inspection by a competent person. Verify every flagged field against the physical plate before | |
| acting on it. | |
| ## Licence and trademarks | |
| Code released under **Apache-2.0**. Brand and product names (Schneider Electric, Eaton, Vertiv, Caterpillar, | |
| Trane, …) are trademarks of their owners and are matched only to identify user-supplied plates; this project is | |
| independent and not affiliated with or endorsed by them. | |
| ## Citation | |
| ```bibtex | |
| @misc{constructelligence_nameplates, | |
| title = {Data Centre Nameplates: equipment-nameplate OCR checked against the equipment schedule}, | |
| author = {Constructelligence}, | |
| year = {2026}, | |
| howpublished = {\url{https://huggingface.co/constructelligence/data-centre-nameplates}} | |
| } | |
| ``` | |