Image-to-Text
English
custom
ocr
document-ai
information-extraction
nameplate
nameplate-ocr
equipment-nameplate
data-center
data-centre
datacenter
mission-critical
mep
electrical
equipment-schedule
schedule-verification
asset-register
commissioning
quality-assurance
ups
pdu
switchgear
generator
ats
transformer
busway
battery
crah
crac
chiller
cooling
construction
aec
edge-ai
on-device
privacy-preserving
tesseract
browser
File size: 14,448 Bytes
5e2d1eb 65f33bb 5e2d1eb b224bd3 5e2d1eb 94e56dd 5e2d1eb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 | ---
license: apache-2.0
language:
- en
library_name: custom
pipeline_tag: image-to-text
tags:
- ocr
- document-ai
- information-extraction
- nameplate
- nameplate-ocr
- equipment-nameplate
- data-center
- data-centre
- datacenter
- mission-critical
- mep
- electrical
- equipment-schedule
- schedule-verification
- asset-register
- commissioning
- quality-assurance
- ups
- pdu
- switchgear
- generator
- ats
- transformer
- busway
- battery
- crah
- crac
- chiller
- cooling
- construction
- aec
- edge-ai
- on-device
- privacy-preserving
- tesseract
- browser
---
# Data Centre Nameplates — equipment-nameplate OCR checked against the schedule
**Read an equipment nameplate and know if it is the unit the schedule asked for.** Point a phone at the
nameplate of a UPS, PDU, switchgear, generator, transformer, busway, battery, CRAH or chiller; the model returns
the plate's fields — manufacturer, model, serial, voltage, phase, Hz, kVA, kW, amps, MCA, MOCP, refrigerant,
manufacture date — and **checks each one against the equipment schedule**, flagging a mismatch instead of a
silent pass. It builds an asset register and exports it to CSV.
Built for the part of a data-centre job where the gear *is* the job. The expensive mistakes are a unit
delivered at the wrong voltage or rating, or one that quietly never makes it onto the asset register.
> **Runs on the device.** OCR and matching run in the browser (or Node) with no server and no upload. A
> nameplate photo of live infrastructure never leaves the phone that took it.
**▶ Try it live:** [huggingface.co/spaces/constructelligence/data-centre-nameplates](https://huggingface.co/spaces/constructelligence/data-centre-nameplates) — a browser demo with a sample schedule and a sample plate.
| | |
|---|---|
| **Task** | Image → structured nameplate fields → schedule verification |
| **Fields** | 13 (manufacturer, model, serial, voltage, phase, Hz, kVA, kW, amps, MCA, MOCP, refrigerant, mfg date) |
| **Equipment** | 13 types (UPS, generator, ATS/STS, switchgear, transformer, PDU/RPP, panelboard, busway, battery, CDU, chiller, CRAH/CRAC, pump) |
| **OCR** | Tesseract.js 5 (English LSTM), pretrained — no fine-tuning |
| **Extraction** | Deterministic labelled-field parser with an OCR error model |
| **Runtime** | Browser · Node · edge — no GPU, no server, no data egress |
| **Licence** | Apache-2.0 |
---
## Why this exists
General OCR reads the text on a nameplate. A data-centre commissioning or QA team needs the text to answer a
narrower question: **is this the right unit, and does it match what was specified?** That means three things a
plain OCR call does not do:
1. **Turn plate text into fields, not paragraphs.** "INPUT: 480Y/277 VAC 3 PH 60 HZ" becomes
`voltage=480Y/277 V`, `phase=3`, `hz=60`; "RATING: 750 KVA / 675 KW" becomes `kva=750`, `kw=675`.
2. **Survive the way OCR misreads plate lettering.** Serifed `V`→`Vv`, `K`→`X` (`XVA`), `HZ`→`Hw`, codes split
after their punctuation (`NPX- 750- 480`). The parser repairs these before matching.
3. **Compare to the schedule the way codes actually differ.** Model and serial codes are compared with `O/0`,
`I/L/1`, `S/5`, `B/8`, `Z/2` folded and punctuation ignored, so a misread is not a false mismatch, while a
genuine voltage variant (`NPX-750-415` vs `NPX-750-480`) still fails.
The result is a pass/mismatch verdict per plate, with the fields that disagree named.
## How it works
```
photo ──► OCR (Tesseract.js) ──► field parser ──► schedule matcher ──► verdict + asset register CSV
│ │ │
lines + confidence labelled-value rules code folding, tolerance checks
```
1. **OCR** (`tesseract.js@5`, English) returns lines with a per-line confidence.
2. **Field parser** scans lines for labelled values (`MODEL`, `S/N`, `MVA`, `MCA`, `MOCP`, `FLA`, `VOLTS`,
`REFRIGERANT`, `MFG DATE`), reading voltages as `480Y/277`, `13.8 kV`, `208/120`, and rejecting look-alikes
(`480/277` with no voltage word is not a voltage; `R-513A` and a charge weight are not amps). Lines OCR
scored near zero (brushed metal, background texture) are dropped before parsing.
3. **Equipment typing** recognises the family from plate wording (UPS, genset, ATS, switchgear, transformer,
PDU/RPP, panelboard, busway, battery, CDU, chiller, CRAH/CRAC, pump).
4. **Schedule matcher** ranks candidate tags — a serial match wins outright; otherwise model similarity
dominates, with manufacturer and equipment type as tie-breakers. Data halls repeat, so among equal matches
the first **unscanned** tag is offered.
5. **Field checks** compare the reading to the chosen row and return `ok`, `mismatch` or `unread` per field, and
an overall status: `verified`, `partial`, `mismatch`, `unmatched` or `missing` (scheduled but not scanned).
## Fields extracted
| Field | Example | Notes |
|---|---|---|
| `manufacturer` | Northline Power Systems | ~44 data-centre makers recognised as whole words; otherwise the company-looking line |
| `model` | NPX-750-480 | `MODEL`, `MOD.`, `M/N`, `CAT. NO.`, `P/N`, `PART NO.` |
| `serial` | NL26A01937 | `SERIAL`, `SER.`, `S/N`, `5/N`, `SN` |
| `voltage` | 480Y/277 V | every voltage on the plate (transformers and UPSes carry more than one) |
| `phase` | 3 | `1`/`3`, `PH`, `PHASE`, `Ø` |
| `hz` | 60 | 50 or 60, `HZ`/`FREQUENCY` |
| `kva` | 750 | thousands separators handled (`3,750`); European decimals too |
| `kw` | 675 | `KW` / `EKW` |
| `amps` | 1002 | labelled (`FLA`, `RLA`, `CURRENT`, `AMPS`) first |
| `mca` | 612 | minimum circuit ampacity |
| `mocp` | 800 | `MOCP` / `MAX FUSE` / `MAX BREAKER` |
| `refrigerant` | R-513A | R-410A, 134A, 1234ZE, 1234YF, 513A, 514A, 1233ZD, 454B, 407C, 32, 22 |
| `mfgDate` | 03/2026 | `MFG`, `MANUFACTURED`, `MFD`, `BUILD`, `PRODUCTION` |
## Equipment types recognised
`ups` · `generator` · `ats` (also STS) · `switchgear` · `transformer` · `pdu` (also RPP) · `panelboard` ·
`busway` · `battery` · `cdu` · `chiller` · `crah` (also CRAC) · `pump`
Manufacturers recognised include Schneider Electric, Square D, APC, Eaton, ABB, Siemens, Vertiv, Liebert,
Caterpillar, Cummins, Kohler, Rolls-Royce, MTU, Generac, Trane, Carrier, Daikin, York, Johnson Controls,
Stulz, Munters, nVent, Legrand, Mitsubishi Electric, Toshiba, General Electric, Hubbell, Socomec, Piller,
Starline, Raritan, Server Technology, Emerson, Russelectric, ASCO, Hitachi, Riello, Rittal, Motivair, CoolIT,
Hitec, EnerSys, Leclanché and Samsung SDI.
## Results — OCR and extraction, end to end
Six drawn plates (UPS, diesel genset, chiller, dry-type transformer, PDU, CRAH), each pushed through the
degradations a phone photo actually has, then run through the app's own OCR + parser and scored **field by
field**. **56 checks per condition** (six plates × their fields), 13 conditions.
| Condition | Field accuracy | ms/plate |
|---|:---:|---:|
| clean | **100%** (56/56) | 1236 |
| small (640 px) | **100%** (56/56) | 897 |
| blur | **100%** (56/56) | 540 |
| noise | **100%** (56/56) | 748 |
| low contrast | **100%** (56/56) | 557 |
| glare | **100%** (56/56) | 532 |
| tilt (7°) | **100%** (56/56) | 612 |
| keystone | **100%** (56/56) | 474 |
| dark anodised plate | **100%** (56/56) | 808 |
| JPEG | **100%** (56/56) | 524 |
| plate small in a cluttered scene | **100%** (56/56) | 895 |
| sideways | **100%** (56/56) | 554 |
| phone (blur + glare + noise + tilt + scene + JPEG) | **98.2%** (55/56) | 1607 |
| **Overall** | **99.9% (727/728)** | **768** |
The single miss is one `hz` value on the hardest "phone" composite. The benchmark measures what **OCR** loses,
not parser coverage: the ground truth is the parser's own reading of the clean plate text, so every reduction
is an OCR failure the parser could not repair. Reproduce it with `node tests/nameplates-bench.mjs`.
**Read this honestly.** These are crisp, machine-drawn plates. A real nameplate photographed on a live unit is
harder: cast shadows, embossed lettering, print over brushed metal, a plate half out of frame. Expect the
degradation suite, not the clean row, and treat every low-confidence field as one a person must confirm.
## Fine-tuning — the higher-accuracy server model
The engine above is the on-device path. Alongside it, a **Donut** model
([`naver-clova-ix/donut-base`](https://huggingface.co/naver-clova-ix/donut-base)) is fine-tuned to read a plate
**straight into fields** (`<s_nameplate><s_type>ups</s_type><s_manufacturer>…`) for a server / endpoint where
the browser engine falls short. Training data is synthetic plates with real-photo degradations, plus corrected
real plates from the field. Once trained, the weights are published in a sibling repository and linked here.
Nothing on this page claims the accuracy of that model yet — it is the roadmap, not a result.
## Intended use
- **Commissioning and QA walks:** photograph each unit as it is installed; confirm the delivered unit matches
the equipment schedule; capture an asset register as you go.
- **Receiving and delivery checks:** flag the unit that arrived at the wrong voltage or rating before it is set.
- **Register completion:** track scheduled tags that have no plate scanned against them yet.
- **Takeoff and submittal review:** pull model/rating fields off a plate photo into a spreadsheet.
### Out of scope — do not use it for
- **Safety, code compliance or energisation decisions.** It reads a plate; it does not verify that a unit is
safe to energise, correctly protected, or code-compliant.
- **An as-built or a legal record on its own.** OCR misreads happen; low-confidence fields are highlighted for a
person to check, and any field that disagrees with the schedule must be confirmed by eye.
- **Serial-number evidence in a dispute** without human verification. Two plates can share a misread serial;
the register flags duplicate serials but cannot tell a repeated plate from a copied number.
## Limitations
- **English nameplates**, Latin script. Non-Latin or handwritten plates are out of scope.
- **Glare, blur and angle degrade OCR.** The parser repairs common misreads, but a plate that is small in the
frame, badly glared or very low contrast will lose fields (those fields come back `unread`, not invented).
- **Deterministic parser, not a language model.** It reports only what it can match to a rule; it will not
infer a missing digit.
- **The schedule must be structured** (CSV with a Tag or Model column). Free-form schedule PDFs are not parsed
here.
## Quickstart
The engine is plain ES modules with no build step. `nameplate-model.js` is the extractor and verifier;
`schedule-import.js` is the CSV reader it depends on.
```js
import { parseNameplate, parseEquipmentSchedule, candidatesFor, checkAgainst, statusOf } from './nameplate-model.js';
// 1. Parse plate text (from Tesseract, or typed by hand)
const reading = parseNameplate(`NORTHLINE POWER SYSTEMS
UNINTERRUPTIBLE POWER SUPPLY
MODEL NO: NPX-750-480
SERIAL NO: NL26A01937
INPUT: 480Y/277 VAC 3 PH 60 HZ
RATING: 750 KVA / 675 KW
INPUT CURRENT: 1002 A
MFG DATE: 03/2026`);
console.log(reading.type); // 'ups'
console.log(reading.fields.voltage); // { value: '480Y/277 V', values: [480, 277], conf, line }
// 2. Match against the equipment schedule and check each field
const { items } = parseEquipmentSchedule(scheduleCsv);
const [best] = candidatesFor(reading, items);
const checks = checkAgainst(reading, best.item);
console.log(statusOf(checks, best.item.tag)); // 'verified' | 'partial' | 'mismatch' | ...
```
In a browser, feed Tesseract's `data.lines` (`[{ text, confidence, bbox }]`) straight into `parseNameplate`
to keep OCR confidence and the plate line each value came from.
## Evaluation
- **Extraction/verification logic:** 12 unit tests over the parser, voltage reader, code folding, schedule
parsing, candidate ranking and mismatch detection — all passing (`node --test tests/nameplate-model.test.mjs`).
- **OCR end-to-end:** a field-by-field benchmark draws plates for six equipment types and puts each through the
degradations real phone photos have — small, blurred, noisy, low-contrast, glare, tilt, keystone, dark
anodised plate, JPEG, plate-small-in-scene and sideways — then scores the app's own OCR + parsing pipeline
field by field. Run it with `node tests/nameplates-bench.mjs`.
## Repository contents
| File | What it is |
|---|---|
| `nameplate-model.js` | the extractor, matcher and register builder (ES module) |
| `schedule-import.js` | the CSV schedule reader it depends on |
| `progress-project.js` | a helper the schedule reader imports |
| `sample_schedule.csv` | a 10-row data-centre equipment schedule |
| `sample_plates.json` | two sample plates: one that verifies, one flagged on model and voltage |
| `eval/` | the OCR benchmark harness output |
| `config.json` | machine-readable fields, types and statuses |
## Provenance
An open extraction and verification engine from Constructelligence, the same code path that powers the
**Nameplates** mode of BuildVision. No customer data is used; sample plates use fictional makers and models.
It is the document-AI sibling of the vision models
[`construction-site-safety-hazards`](https://huggingface.co/constructelligence/construction-site-safety-hazards)
and [`electrical-circuit-connectivity`](https://huggingface.co/constructelligence/electrical-circuit-connectivity).
## Safety
This is a **prompt to look**, not a finding. It is not safety-rated, not a compliance decision, and not a
substitute for inspection by a competent person. Verify every flagged field against the physical plate before
acting on it.
## Licence and trademarks
Code released under **Apache-2.0**. Brand and product names (Schneider Electric, Eaton, Vertiv, Caterpillar,
Trane, …) are trademarks of their owners and are matched only to identify user-supplied plates; this project is
independent and not affiliated with or endorsed by them.
## Citation
```bibtex
@misc{constructelligence_nameplates,
title = {Data Centre Nameplates: equipment-nameplate OCR checked against the equipment schedule},
author = {Constructelligence},
year = {2026},
howpublished = {\url{https://huggingface.co/constructelligence/data-centre-nameplates}}
}
```
|