danmu-detector / README.md
DearthAI's picture
Update README.md
abe26bd verified
|
Raw
History Blame Contribute Delete
10.2 kB
---
license: agpl-3.0
library_name: onnx
pipeline_tag: object-detection
tags:
- yolov8
- yolo-world
- onnx
- object-detection
- open-images
- open-vocabulary
- in-browser
- onnxruntime-web
---
# danmu-detector
Two ONNX detectors for in-browser furniture detection via `onnxruntime-web`,
meant to be run **together**. Built for Danmu, a local-first interior decoration
app — everything runs on-device, and works offline once the files are cached.
They are paired because they fail on *different* classes. Measured on a real
four-photo room containing 19 catalogued objects:
| | recall |
|---|---|
| `yolov8n-oiv7` alone | 7/19 (37%) |
| `yolov8s-worldv2-danmu` alone | 10/19 (53%) |
| **both** | **13/19 (68%)** |
The fixed-label model reliably finds monitors and windows. The open-vocabulary
one finds refrigerators, ceiling fans, wardrobes, lamps and curtains — none of
which the first detects **at all**: those class heads peak at 0.002–0.03 on real
room photos against 0.38–0.44 for classes that do fire, and they stay dead at
every model size (s, m and x all score the same 7/19 as nano). That is a
vocabulary limit, not a capacity one, which is what the second model fixes.
Nothing here depends on an external repository; the reference implementation
below is complete.
## Files
| File | Size | Output | Notes |
|---|---|---|---|
| `yolov8n-oiv7.onnx` | 14.2 MB | `1x605x8400` | YOLOv8 nano, Open Images V7, 601 fixed classes |
| `yolov8n-oiv7.names.json` | 12 KB | — | class index → name, 601 entries |
| `yolov8s-worldv2-danmu.onnx` | 50.4 MB | `1x48x8400` | YOLO-World small, 44 furniture prompts baked in |
Both are opset 12, input `1x3x640x640` NCHW, and share the standard YOLOv8
detect-head layout: channels are `cx, cy, w, h` then per-class scores
(`605 - 4 = 601`, `48 - 4 = 44`) across 8400 anchors. **No objectness channel and
no sigmoid** — scores are ready to threshold. Boxes are in letterboxed 640-space
and must be unpadded.
Both use only standard ONNX ops (no custom domains), so they run under
`onnxruntime-web` on WASM or WebGPU.
## The baked vocabulary
`yolov8s-worldv2-danmu.onnx` was produced with Ultralytics `set_classes()`,
which runs the CLIP text encoder once at export and freezes the embeddings into
the graph — so no text encoder is needed at runtime. **Class index N is
`PROMPTS[N]`**, in exactly this order:
```js
const PROMPTS = [
'sofa', 'couch', 'armchair',
'chair', 'office chair', 'stool',
'table', 'coffee table', 'dining table',
'desk',
'bed', 'mattress',
'nightstand',
'wardrobe', 'closet', 'chest of drawers', 'storage cabinet',
'shelf', 'bookshelf', 'shoe rack',
'mirror',
'curtain', 'window curtain', 'window blind',
'picture frame', 'wall art', 'poster',
'lamp', 'light bulb', 'ceiling light',
'ceiling fan', 'electric fan',
'refrigerator',
'potted plant',
'door', 'wooden door',
'computer monitor',
'television',
'window',
'laptop',
'washing machine',
'microwave oven',
'clothes rack', 'hanging clothes',
];
```
Do not reorder that list — it is the model's channel order, and an off-by-one
mislabels every detection silently.
Several phrases deliberately collapse to one concept (`clothes rack`, `hanging
clothes`, `storage cabinet` → wardrobe). Real rooms contain clothes rails and
stacked fabric cubes, not the canonical `Wardrobe` a fixed-label model was
trained on, and naming what is actually there is the whole advantage of an
open-vocabulary model. To retarget it, re-export with your own prompt list.
## Loading
```js
const BASE = 'https://huggingface.co/DearthAI/danmu-detector/resolve/main/';
const providers = [];
if (typeof navigator !== 'undefined' && 'gpu' in navigator) providers.push('webgpu');
providers.push('wasm');
const oiv = await ort.InferenceSession.create(BASE + 'yolov8n-oiv7.onnx', { executionProviders: providers });
const world = await ort.InferenceSession.create(BASE + 'yolov8s-worldv2-danmu.onnx', { executionProviders: providers });
const names = await (await fetch(BASE + 'yolov8n-oiv7.names.json')).json();
```
Use the `/resolve/` URL, not `/blob/``/blob/` returns the HTML viewer page,
which passes a HEAD status check and then hands your runtime a page of markup.
## Preprocessing
Letterbox to 640×640 preserving aspect ratio, pad with grey, normalize to `0..1`,
feed **planar** RGB (all R, then all G, then all B — not interleaved).
```js
const INPUT = 640;
function toTensor(img, ox, oy, cw, ch) { // crop rect in source pixels
const scale = INPUT / Math.max(cw, ch);
const sw = Math.round(cw * scale), sh = Math.round(ch * scale);
const dw = Math.floor((INPUT - sw) / 2), dh = Math.floor((INPUT - sh) / 2);
const c = document.createElement('canvas');
c.width = c.height = INPUT;
const ctx = c.getContext('2d');
ctx.fillStyle = '#727272'; // letterbox grey
ctx.fillRect(0, 0, INPUT, INPUT);
ctx.drawImage(img, ox, oy, cw, ch, dw, dh, sw, sh);
const px = ctx.getImageData(0, 0, INPUT, INPUT).data;
const area = INPUT * INPUT;
const data = new Float32Array(3 * area);
for (let i = 0; i < area; i++) {
data[i] = px[i * 4] / 255;
data[area + i] = px[i * 4 + 1] / 255;
data[2 * area + i] = px[i * 4 + 2] / 255;
}
return { data, scale, dw, dh };
}
```
## Tile — it matters more than model size
Letterboxing a 2000px wide-angle photo to 640 shrinks mid-sized furniture below
what these models resolve. Running the whole frame **plus 2×2 tiles at 15%
overlap** took nano from 4/19 to 7/19 for zero extra download, and the world
model from 7/19 to 10/19. Overlap keeps objects straddling a seam intact.
```js
function tilesFor(iw, ih) {
const ox = iw * 0.15, oy = ih * 0.15;
const crops = [{ ox: 0, oy: 0, cw: iw, ch: ih }];
for (let r = 0; r < 2; r++)
for (let c = 0; c < 2; c++) {
const x0 = Math.max(0, c * iw / 2 - ox), y0 = Math.max(0, r * ih / 2 - oy);
const x1 = Math.min(iw, (c + 1) * iw / 2 + ox);
const y1 = Math.min(ih, (r + 1) * ih / 2 + oy);
crops.push({ ox: x0, oy: y0, cw: x1 - x0, ch: y1 - y0 });
}
return crops;
}
```
## Decoding, and merging the two models
Collect candidates from **both models over every tile into one pool in
normalized whole-image coordinates**, then run a single NMS. That one pass
resolves per-tile duplicates, cross-tile seam duplicates, and the same object
found by both models.
```js
const CONF = 0.35, IOU_T = 0.45, MAX_PER_IMAGE = 12;
const pool = [];
for (const crop of tilesFor(iw, ih)) {
const pre = toTensor(img, crop.ox, crop.oy, crop.cw, crop.ch);
const toX = v => (crop.ox + (v - pre.dw) / pre.scale) / iw;
const toY = v => (crop.oy + (v - pre.dh) / pre.scale) / ih;
for (const [sess, labelOf] of [
[oiv, i => names[String(i)]],
[world, i => PROMPTS[i]],
]) {
const t = new ort.Tensor('float32', pre.data.slice(), [1, 3, INPUT, INPUT]);
const out = (await sess.run({ [sess.inputNames[0]]: t }))[sess.outputNames[0]];
const [, channels, anchors] = out.dims;
const nc = channels - 4, d = out.data;
const at = (ch, a) => d[ch * anchors + a];
for (let a = 0; a < anchors; a++) {
let best = 0, bestC = -1;
for (let ci = 0; ci < nc; ci++) {
const s = at(4 + ci, a);
if (s > best) { best = s; bestC = ci; }
}
if (best < CONF || bestC < 0) continue;
pool.push({
x: toX(at(0, a)), y: toY(at(1, a)),
w: at(2, a) / pre.scale / iw,
h: at(3, a) / pre.scale / ih,
conf: best, label: labelOf(bestC),
});
}
}
}
```
Then NMS in that same normalized space, and convert centre-form to top-left:
```js
function iou(a, b) {
const x1 = Math.max(a.x - a.w / 2, b.x - b.w / 2);
const y1 = Math.max(a.y - a.h / 2, b.y - b.h / 2);
const x2 = Math.min(a.x + a.w / 2, b.x + b.w / 2);
const y2 = Math.min(a.y + a.h / 2, b.y + b.h / 2);
const inter = Math.max(0, x2 - x1) * Math.max(0, y2 - y1);
return inter / (a.w * a.h + b.w * b.h - inter + 1e-9);
}
const kept = [];
for (const b of [...pool].sort((p, q) => q.conf - p.conf)) {
if (kept.every(k => iou(k, b) < IOU_T)) kept.push(b);
if (kept.length >= MAX_PER_IMAGE) break;
}
const boxes = kept.map(b => ({
label: b.label, conf: b.conf,
x: b.x - b.w / 2, y: b.y - b.h / 2, w: b.w, h: b.h,
}));
```
Subtracting `dw`/`dh` before dividing is the step that is easy to miss — skip it
and every box drifts toward the image centre on non-square inputs.
## Tuning notes
Measured, so you don't have to rediscover them:
- **Confidence 0.35.** Dropping to 0.20 gained one object and added a spurious
`sofa(0.29)`.
- **`MAX_PER_IMAGE` 12.** Raising it to 30 added one box and changed no score.
- **Bigger is not better.** `yolov8s/m/x-oiv7` (46 / 105 / 275 MB) all score the
same 7/19 as the 14 MB nano. Spend the bytes on the second model instead.
- **Still missed** at 13/19: doors in some framings, wall art, and curtains in a
frame that is almost entirely curtain. Treat the output as a head start on
manual correction, not a finished result.
## Reproducing
```bash
pip install --index-url https://download.pytorch.org/whl/cpu torch
pip install ultralytics onnx onnxslim onnxruntime clip-anytorch ftfy
python - <<'PY'
from ultralytics import YOLO
YOLO('yolov8n-oiv7.pt').export(format='onnx', imgsz=640, opset=12)
m = YOLO('yolov8s-worldv2.pt')
m.set_classes(PROMPTS) # the 44 phrases above, in order
m.export(format='onnx', imgsz=640, opset=12)
PY
```
CPU-only torch is enough — export needs no GPU, and it keeps the install near
200 MB instead of 2.5 GB.
## Licence
**AGPL-3.0.** Both are Ultralytics YOLO weights, licensed AGPL-3.0 by
[Ultralytics](https://github.com/ultralytics/ultralytics). The AGPL obligations
attach to these artifacts and travel with them.
They are hosted here, separately from the application that consumes them,
precisely so the licences stay distinct: that app is MIT, and it fetches these
weights at runtime rather than bundling or redistributing them. If you vendor
these files into your own project, AGPL-3.0 applies to you — a permissive
licence on the surrounding code does not cover them.