File size: 10,215 Bytes
25624e1
 
7dfb78f
 
 
 
abe26bd
7dfb78f
 
 
abe26bd
7dfb78f
 
25624e1
7dfb78f
 
 
abe26bd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ab4d6d
 
 
7dfb78f
 
 
abe26bd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7dfb78f
abe26bd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7dfb78f
9ab4d6d
7dfb78f
 
 
9ab4d6d
 
 
 
 
abe26bd
 
9ab4d6d
 
 
 
 
 
 
 
 
abe26bd
9ab4d6d
 
 
 
abe26bd
 
 
9ab4d6d
 
 
 
 
 
 
abe26bd
9ab4d6d
 
 
 
 
abe26bd
 
 
9ab4d6d
abe26bd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ab4d6d
 
 
abe26bd
 
 
 
 
 
9ab4d6d
 
abe26bd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ab4d6d
 
 
 
abe26bd
9ab4d6d
 
 
 
 
 
 
 
 
 
 
 
abe26bd
 
 
9ab4d6d
abe26bd
 
 
9ab4d6d
7dfb78f
 
abe26bd
9ab4d6d
 
abe26bd
9ab4d6d
abe26bd
9ab4d6d
abe26bd
 
 
 
 
 
 
 
7dfb78f
 
 
 
abe26bd
 
7dfb78f
abe26bd
 
 
9ab4d6d
abe26bd
 
 
 
9ab4d6d
 
abe26bd
 
 
7dfb78f
 
abe26bd
7dfb78f
abe26bd
7dfb78f
9ab4d6d
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
---
license: agpl-3.0
library_name: onnx
pipeline_tag: object-detection
tags:
  - yolov8
  - yolo-world
  - onnx
  - object-detection
  - open-images
  - open-vocabulary
  - in-browser
  - onnxruntime-web
---

# danmu-detector

Two ONNX detectors for in-browser furniture detection via `onnxruntime-web`,
meant to be run **together**. Built for Danmu, a local-first interior decoration
app β€” everything runs on-device, and works offline once the files are cached.

They are paired because they fail on *different* classes. Measured on a real
four-photo room containing 19 catalogued objects:

| | recall |
|---|---|
| `yolov8n-oiv7` alone | 7/19 (37%) |
| `yolov8s-worldv2-danmu` alone | 10/19 (53%) |
| **both** | **13/19 (68%)** |

The fixed-label model reliably finds monitors and windows. The open-vocabulary
one finds refrigerators, ceiling fans, wardrobes, lamps and curtains β€” none of
which the first detects **at all**: those class heads peak at 0.002–0.03 on real
room photos against 0.38–0.44 for classes that do fire, and they stay dead at
every model size (s, m and x all score the same 7/19 as nano). That is a
vocabulary limit, not a capacity one, which is what the second model fixes.

Nothing here depends on an external repository; the reference implementation
below is complete.

## Files

| File | Size | Output | Notes |
|---|---|---|---|
| `yolov8n-oiv7.onnx` | 14.2 MB | `1x605x8400` | YOLOv8 nano, Open Images V7, 601 fixed classes |
| `yolov8n-oiv7.names.json` | 12 KB | β€” | class index β†’ name, 601 entries |
| `yolov8s-worldv2-danmu.onnx` | 50.4 MB | `1x48x8400` | YOLO-World small, 44 furniture prompts baked in |

Both are opset 12, input `1x3x640x640` NCHW, and share the standard YOLOv8
detect-head layout: channels are `cx, cy, w, h` then per-class scores
(`605 - 4 = 601`, `48 - 4 = 44`) across 8400 anchors. **No objectness channel and
no sigmoid** β€” scores are ready to threshold. Boxes are in letterboxed 640-space
and must be unpadded.

Both use only standard ONNX ops (no custom domains), so they run under
`onnxruntime-web` on WASM or WebGPU.

## The baked vocabulary

`yolov8s-worldv2-danmu.onnx` was produced with Ultralytics `set_classes()`,
which runs the CLIP text encoder once at export and freezes the embeddings into
the graph β€” so no text encoder is needed at runtime. **Class index N is
`PROMPTS[N]`**, in exactly this order:

```js
const PROMPTS = [
  'sofa', 'couch', 'armchair',
  'chair', 'office chair', 'stool',
  'table', 'coffee table', 'dining table',
  'desk',
  'bed', 'mattress',
  'nightstand',
  'wardrobe', 'closet', 'chest of drawers', 'storage cabinet',
  'shelf', 'bookshelf', 'shoe rack',
  'mirror',
  'curtain', 'window curtain', 'window blind',
  'picture frame', 'wall art', 'poster',
  'lamp', 'light bulb', 'ceiling light',
  'ceiling fan', 'electric fan',
  'refrigerator',
  'potted plant',
  'door', 'wooden door',
  'computer monitor',
  'television',
  'window',
  'laptop',
  'washing machine',
  'microwave oven',
  'clothes rack', 'hanging clothes',
];
```

Do not reorder that list β€” it is the model's channel order, and an off-by-one
mislabels every detection silently.

Several phrases deliberately collapse to one concept (`clothes rack`, `hanging
clothes`, `storage cabinet` β†’ wardrobe). Real rooms contain clothes rails and
stacked fabric cubes, not the canonical `Wardrobe` a fixed-label model was
trained on, and naming what is actually there is the whole advantage of an
open-vocabulary model. To retarget it, re-export with your own prompt list.

## Loading

```js
const BASE = 'https://huggingface.co/DearthAI/danmu-detector/resolve/main/';

const providers = [];
if (typeof navigator !== 'undefined' && 'gpu' in navigator) providers.push('webgpu');
providers.push('wasm');

const oiv   = await ort.InferenceSession.create(BASE + 'yolov8n-oiv7.onnx', { executionProviders: providers });
const world = await ort.InferenceSession.create(BASE + 'yolov8s-worldv2-danmu.onnx', { executionProviders: providers });
const names = await (await fetch(BASE + 'yolov8n-oiv7.names.json')).json();
```

Use the `/resolve/` URL, not `/blob/` β€” `/blob/` returns the HTML viewer page,
which passes a HEAD status check and then hands your runtime a page of markup.

## Preprocessing

Letterbox to 640Γ—640 preserving aspect ratio, pad with grey, normalize to `0..1`,
feed **planar** RGB (all R, then all G, then all B β€” not interleaved).

```js
const INPUT = 640;

function toTensor(img, ox, oy, cw, ch) {      // crop rect in source pixels
  const scale = INPUT / Math.max(cw, ch);
  const sw = Math.round(cw * scale), sh = Math.round(ch * scale);
  const dw = Math.floor((INPUT - sw) / 2), dh = Math.floor((INPUT - sh) / 2);

  const c = document.createElement('canvas');
  c.width = c.height = INPUT;
  const ctx = c.getContext('2d');
  ctx.fillStyle = '#727272';                  // letterbox grey
  ctx.fillRect(0, 0, INPUT, INPUT);
  ctx.drawImage(img, ox, oy, cw, ch, dw, dh, sw, sh);

  const px = ctx.getImageData(0, 0, INPUT, INPUT).data;
  const area = INPUT * INPUT;
  const data = new Float32Array(3 * area);
  for (let i = 0; i < area; i++) {
    data[i]            = px[i * 4]     / 255;
    data[area + i]     = px[i * 4 + 1] / 255;
    data[2 * area + i] = px[i * 4 + 2] / 255;
  }
  return { data, scale, dw, dh };
}
```

## Tile β€” it matters more than model size

Letterboxing a 2000px wide-angle photo to 640 shrinks mid-sized furniture below
what these models resolve. Running the whole frame **plus 2Γ—2 tiles at 15%
overlap** took nano from 4/19 to 7/19 for zero extra download, and the world
model from 7/19 to 10/19. Overlap keeps objects straddling a seam intact.

```js
function tilesFor(iw, ih) {
  const ox = iw * 0.15, oy = ih * 0.15;
  const crops = [{ ox: 0, oy: 0, cw: iw, ch: ih }];
  for (let r = 0; r < 2; r++)
    for (let c = 0; c < 2; c++) {
      const x0 = Math.max(0, c * iw / 2 - ox), y0 = Math.max(0, r * ih / 2 - oy);
      const x1 = Math.min(iw, (c + 1) * iw / 2 + ox);
      const y1 = Math.min(ih, (r + 1) * ih / 2 + oy);
      crops.push({ ox: x0, oy: y0, cw: x1 - x0, ch: y1 - y0 });
    }
  return crops;
}
```

## Decoding, and merging the two models

Collect candidates from **both models over every tile into one pool in
normalized whole-image coordinates**, then run a single NMS. That one pass
resolves per-tile duplicates, cross-tile seam duplicates, and the same object
found by both models.

```js
const CONF = 0.35, IOU_T = 0.45, MAX_PER_IMAGE = 12;
const pool = [];

for (const crop of tilesFor(iw, ih)) {
  const pre = toTensor(img, crop.ox, crop.oy, crop.cw, crop.ch);
  const toX = v => (crop.ox + (v - pre.dw) / pre.scale) / iw;
  const toY = v => (crop.oy + (v - pre.dh) / pre.scale) / ih;

  for (const [sess, labelOf] of [
    [oiv,   i => names[String(i)]],
    [world, i => PROMPTS[i]],
  ]) {
    const t = new ort.Tensor('float32', pre.data.slice(), [1, 3, INPUT, INPUT]);
    const out = (await sess.run({ [sess.inputNames[0]]: t }))[sess.outputNames[0]];
    const [, channels, anchors] = out.dims;
    const nc = channels - 4, d = out.data;
    const at = (ch, a) => d[ch * anchors + a];

    for (let a = 0; a < anchors; a++) {
      let best = 0, bestC = -1;
      for (let ci = 0; ci < nc; ci++) {
        const s = at(4 + ci, a);
        if (s > best) { best = s; bestC = ci; }
      }
      if (best < CONF || bestC < 0) continue;
      pool.push({
        x: toX(at(0, a)), y: toY(at(1, a)),
        w: at(2, a) / pre.scale / iw,
        h: at(3, a) / pre.scale / ih,
        conf: best, label: labelOf(bestC),
      });
    }
  }
}
```

Then NMS in that same normalized space, and convert centre-form to top-left:

```js
function iou(a, b) {
  const x1 = Math.max(a.x - a.w / 2, b.x - b.w / 2);
  const y1 = Math.max(a.y - a.h / 2, b.y - b.h / 2);
  const x2 = Math.min(a.x + a.w / 2, b.x + b.w / 2);
  const y2 = Math.min(a.y + a.h / 2, b.y + b.h / 2);
  const inter = Math.max(0, x2 - x1) * Math.max(0, y2 - y1);
  return inter / (a.w * a.h + b.w * b.h - inter + 1e-9);
}

const kept = [];
for (const b of [...pool].sort((p, q) => q.conf - p.conf)) {
  if (kept.every(k => iou(k, b) < IOU_T)) kept.push(b);
  if (kept.length >= MAX_PER_IMAGE) break;
}
const boxes = kept.map(b => ({
  label: b.label, conf: b.conf,
  x: b.x - b.w / 2, y: b.y - b.h / 2, w: b.w, h: b.h,
}));
```

Subtracting `dw`/`dh` before dividing is the step that is easy to miss β€” skip it
and every box drifts toward the image centre on non-square inputs.

## Tuning notes

Measured, so you don't have to rediscover them:

- **Confidence 0.35.** Dropping to 0.20 gained one object and added a spurious
  `sofa(0.29)`.
- **`MAX_PER_IMAGE` 12.** Raising it to 30 added one box and changed no score.
- **Bigger is not better.** `yolov8s/m/x-oiv7` (46 / 105 / 275 MB) all score the
  same 7/19 as the 14 MB nano. Spend the bytes on the second model instead.
- **Still missed** at 13/19: doors in some framings, wall art, and curtains in a
  frame that is almost entirely curtain. Treat the output as a head start on
  manual correction, not a finished result.

## Reproducing

```bash
pip install --index-url https://download.pytorch.org/whl/cpu torch
pip install ultralytics onnx onnxslim onnxruntime clip-anytorch ftfy

python - <<'PY'
from ultralytics import YOLO
YOLO('yolov8n-oiv7.pt').export(format='onnx', imgsz=640, opset=12)

m = YOLO('yolov8s-worldv2.pt')
m.set_classes(PROMPTS)          # the 44 phrases above, in order
m.export(format='onnx', imgsz=640, opset=12)
PY
```

CPU-only torch is enough β€” export needs no GPU, and it keeps the install near
200 MB instead of 2.5 GB.

## Licence

**AGPL-3.0.** Both are Ultralytics YOLO weights, licensed AGPL-3.0 by
[Ultralytics](https://github.com/ultralytics/ultralytics). The AGPL obligations
attach to these artifacts and travel with them.

They are hosted here, separately from the application that consumes them,
precisely so the licences stay distinct: that app is MIT, and it fetches these
weights at runtime rather than bundling or redistributing them. If you vendor
these files into your own project, AGPL-3.0 applies to you β€” a permissive
licence on the surrounding code does not cover them.