--- license: agpl-3.0 library_name: onnx pipeline_tag: object-detection tags: - yolov8 - yolo-world - onnx - object-detection - open-images - open-vocabulary - in-browser - onnxruntime-web --- # danmu-detector Two ONNX detectors for in-browser furniture detection via `onnxruntime-web`, meant to be run **together**. Built for Danmu, a local-first interior decoration app — everything runs on-device, and works offline once the files are cached. They are paired because they fail on *different* classes. Measured on a real four-photo room containing 19 catalogued objects: | | recall | |---|---| | `yolov8n-oiv7` alone | 7/19 (37%) | | `yolov8s-worldv2-danmu` alone | 10/19 (53%) | | **both** | **13/19 (68%)** | The fixed-label model reliably finds monitors and windows. The open-vocabulary one finds refrigerators, ceiling fans, wardrobes, lamps and curtains — none of which the first detects **at all**: those class heads peak at 0.002–0.03 on real room photos against 0.38–0.44 for classes that do fire, and they stay dead at every model size (s, m and x all score the same 7/19 as nano). That is a vocabulary limit, not a capacity one, which is what the second model fixes. Nothing here depends on an external repository; the reference implementation below is complete. ## Files | File | Size | Output | Notes | |---|---|---|---| | `yolov8n-oiv7.onnx` | 14.2 MB | `1x605x8400` | YOLOv8 nano, Open Images V7, 601 fixed classes | | `yolov8n-oiv7.names.json` | 12 KB | — | class index → name, 601 entries | | `yolov8s-worldv2-danmu.onnx` | 50.4 MB | `1x48x8400` | YOLO-World small, 44 furniture prompts baked in | Both are opset 12, input `1x3x640x640` NCHW, and share the standard YOLOv8 detect-head layout: channels are `cx, cy, w, h` then per-class scores (`605 - 4 = 601`, `48 - 4 = 44`) across 8400 anchors. **No objectness channel and no sigmoid** — scores are ready to threshold. Boxes are in letterboxed 640-space and must be unpadded. Both use only standard ONNX ops (no custom domains), so they run under `onnxruntime-web` on WASM or WebGPU. ## The baked vocabulary `yolov8s-worldv2-danmu.onnx` was produced with Ultralytics `set_classes()`, which runs the CLIP text encoder once at export and freezes the embeddings into the graph — so no text encoder is needed at runtime. **Class index N is `PROMPTS[N]`**, in exactly this order: ```js const PROMPTS = [ 'sofa', 'couch', 'armchair', 'chair', 'office chair', 'stool', 'table', 'coffee table', 'dining table', 'desk', 'bed', 'mattress', 'nightstand', 'wardrobe', 'closet', 'chest of drawers', 'storage cabinet', 'shelf', 'bookshelf', 'shoe rack', 'mirror', 'curtain', 'window curtain', 'window blind', 'picture frame', 'wall art', 'poster', 'lamp', 'light bulb', 'ceiling light', 'ceiling fan', 'electric fan', 'refrigerator', 'potted plant', 'door', 'wooden door', 'computer monitor', 'television', 'window', 'laptop', 'washing machine', 'microwave oven', 'clothes rack', 'hanging clothes', ]; ``` Do not reorder that list — it is the model's channel order, and an off-by-one mislabels every detection silently. Several phrases deliberately collapse to one concept (`clothes rack`, `hanging clothes`, `storage cabinet` → wardrobe). Real rooms contain clothes rails and stacked fabric cubes, not the canonical `Wardrobe` a fixed-label model was trained on, and naming what is actually there is the whole advantage of an open-vocabulary model. To retarget it, re-export with your own prompt list. ## Loading ```js const BASE = 'https://huggingface.co/DearthAI/danmu-detector/resolve/main/'; const providers = []; if (typeof navigator !== 'undefined' && 'gpu' in navigator) providers.push('webgpu'); providers.push('wasm'); const oiv = await ort.InferenceSession.create(BASE + 'yolov8n-oiv7.onnx', { executionProviders: providers }); const world = await ort.InferenceSession.create(BASE + 'yolov8s-worldv2-danmu.onnx', { executionProviders: providers }); const names = await (await fetch(BASE + 'yolov8n-oiv7.names.json')).json(); ``` Use the `/resolve/` URL, not `/blob/` — `/blob/` returns the HTML viewer page, which passes a HEAD status check and then hands your runtime a page of markup. ## Preprocessing Letterbox to 640×640 preserving aspect ratio, pad with grey, normalize to `0..1`, feed **planar** RGB (all R, then all G, then all B — not interleaved). ```js const INPUT = 640; function toTensor(img, ox, oy, cw, ch) { // crop rect in source pixels const scale = INPUT / Math.max(cw, ch); const sw = Math.round(cw * scale), sh = Math.round(ch * scale); const dw = Math.floor((INPUT - sw) / 2), dh = Math.floor((INPUT - sh) / 2); const c = document.createElement('canvas'); c.width = c.height = INPUT; const ctx = c.getContext('2d'); ctx.fillStyle = '#727272'; // letterbox grey ctx.fillRect(0, 0, INPUT, INPUT); ctx.drawImage(img, ox, oy, cw, ch, dw, dh, sw, sh); const px = ctx.getImageData(0, 0, INPUT, INPUT).data; const area = INPUT * INPUT; const data = new Float32Array(3 * area); for (let i = 0; i < area; i++) { data[i] = px[i * 4] / 255; data[area + i] = px[i * 4 + 1] / 255; data[2 * area + i] = px[i * 4 + 2] / 255; } return { data, scale, dw, dh }; } ``` ## Tile — it matters more than model size Letterboxing a 2000px wide-angle photo to 640 shrinks mid-sized furniture below what these models resolve. Running the whole frame **plus 2×2 tiles at 15% overlap** took nano from 4/19 to 7/19 for zero extra download, and the world model from 7/19 to 10/19. Overlap keeps objects straddling a seam intact. ```js function tilesFor(iw, ih) { const ox = iw * 0.15, oy = ih * 0.15; const crops = [{ ox: 0, oy: 0, cw: iw, ch: ih }]; for (let r = 0; r < 2; r++) for (let c = 0; c < 2; c++) { const x0 = Math.max(0, c * iw / 2 - ox), y0 = Math.max(0, r * ih / 2 - oy); const x1 = Math.min(iw, (c + 1) * iw / 2 + ox); const y1 = Math.min(ih, (r + 1) * ih / 2 + oy); crops.push({ ox: x0, oy: y0, cw: x1 - x0, ch: y1 - y0 }); } return crops; } ``` ## Decoding, and merging the two models Collect candidates from **both models over every tile into one pool in normalized whole-image coordinates**, then run a single NMS. That one pass resolves per-tile duplicates, cross-tile seam duplicates, and the same object found by both models. ```js const CONF = 0.35, IOU_T = 0.45, MAX_PER_IMAGE = 12; const pool = []; for (const crop of tilesFor(iw, ih)) { const pre = toTensor(img, crop.ox, crop.oy, crop.cw, crop.ch); const toX = v => (crop.ox + (v - pre.dw) / pre.scale) / iw; const toY = v => (crop.oy + (v - pre.dh) / pre.scale) / ih; for (const [sess, labelOf] of [ [oiv, i => names[String(i)]], [world, i => PROMPTS[i]], ]) { const t = new ort.Tensor('float32', pre.data.slice(), [1, 3, INPUT, INPUT]); const out = (await sess.run({ [sess.inputNames[0]]: t }))[sess.outputNames[0]]; const [, channels, anchors] = out.dims; const nc = channels - 4, d = out.data; const at = (ch, a) => d[ch * anchors + a]; for (let a = 0; a < anchors; a++) { let best = 0, bestC = -1; for (let ci = 0; ci < nc; ci++) { const s = at(4 + ci, a); if (s > best) { best = s; bestC = ci; } } if (best < CONF || bestC < 0) continue; pool.push({ x: toX(at(0, a)), y: toY(at(1, a)), w: at(2, a) / pre.scale / iw, h: at(3, a) / pre.scale / ih, conf: best, label: labelOf(bestC), }); } } } ``` Then NMS in that same normalized space, and convert centre-form to top-left: ```js function iou(a, b) { const x1 = Math.max(a.x - a.w / 2, b.x - b.w / 2); const y1 = Math.max(a.y - a.h / 2, b.y - b.h / 2); const x2 = Math.min(a.x + a.w / 2, b.x + b.w / 2); const y2 = Math.min(a.y + a.h / 2, b.y + b.h / 2); const inter = Math.max(0, x2 - x1) * Math.max(0, y2 - y1); return inter / (a.w * a.h + b.w * b.h - inter + 1e-9); } const kept = []; for (const b of [...pool].sort((p, q) => q.conf - p.conf)) { if (kept.every(k => iou(k, b) < IOU_T)) kept.push(b); if (kept.length >= MAX_PER_IMAGE) break; } const boxes = kept.map(b => ({ label: b.label, conf: b.conf, x: b.x - b.w / 2, y: b.y - b.h / 2, w: b.w, h: b.h, })); ``` Subtracting `dw`/`dh` before dividing is the step that is easy to miss — skip it and every box drifts toward the image centre on non-square inputs. ## Tuning notes Measured, so you don't have to rediscover them: - **Confidence 0.35.** Dropping to 0.20 gained one object and added a spurious `sofa(0.29)`. - **`MAX_PER_IMAGE` 12.** Raising it to 30 added one box and changed no score. - **Bigger is not better.** `yolov8s/m/x-oiv7` (46 / 105 / 275 MB) all score the same 7/19 as the 14 MB nano. Spend the bytes on the second model instead. - **Still missed** at 13/19: doors in some framings, wall art, and curtains in a frame that is almost entirely curtain. Treat the output as a head start on manual correction, not a finished result. ## Reproducing ```bash pip install --index-url https://download.pytorch.org/whl/cpu torch pip install ultralytics onnx onnxslim onnxruntime clip-anytorch ftfy python - <<'PY' from ultralytics import YOLO YOLO('yolov8n-oiv7.pt').export(format='onnx', imgsz=640, opset=12) m = YOLO('yolov8s-worldv2.pt') m.set_classes(PROMPTS) # the 44 phrases above, in order m.export(format='onnx', imgsz=640, opset=12) PY ``` CPU-only torch is enough — export needs no GPU, and it keeps the install near 200 MB instead of 2.5 GB. ## Licence **AGPL-3.0.** Both are Ultralytics YOLO weights, licensed AGPL-3.0 by [Ultralytics](https://github.com/ultralytics/ultralytics). The AGPL obligations attach to these artifacts and travel with them. They are hosted here, separately from the application that consumes them, precisely so the licences stay distinct: that app is MIT, and it fetches these weights at runtime rather than bundling or redistributing them. If you vendor these files into your own project, AGPL-3.0 applies to you — a permissive licence on the surrounding code does not cover them.