File size: 18,470 Bytes
9b92c75
 
 
 
 
 
 
 
 
 
9c7fbbf
 
 
9b92c75
b054fd7
9c7fbbf
 
 
 
9b92c75
 
 
 
 
 
 
 
 
 
3547bdb
 
 
 
9b92c75
 
 
 
 
 
 
911143d
3547bdb
9b92c75
 
 
 
 
 
 
 
 
 
 
 
 
 
911143d
 
9b92c75
 
 
 
 
73256ab
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9b92c75
 
3547bdb
 
9b92c75
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
911143d
9b92c75
 
 
 
 
 
911143d
9b92c75
 
 
4bc49a3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9b92c75
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a889641
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9b92c75
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
---
license: apache-2.0
tags:
  - object-detection
  - pytorch
  - computer-vision
  - research
---

# ObjectModel-v1
<table align="center">
<tr>
<td align="center">

<a href="https://huggingface.co/spaces/bench-labs/blog?post=ObjectModel-v1.html">📝 Blog Post</a>

</td>
</tr>
</table>
ObjectModel-v1 is a clean-room, compact, end-to-end object detector from Bench Labs. It is a
research implementation, not a benchmark claim. The model tests
whether global semantic reasoning can be compressed into a small fixed latent memory while
precise geometry is recovered by query-conditioned sampling from full-resolution pyramid
features.

The model is NMS-free. It predicts a fixed set of objects and is trained with Hungarian
bipartite matching.

This is v1: a single from-scratch training run with no pretraining, no hyperparameter search
across seeds, and none of the planned ablations run yet. Treat everything below as a first
checkpoint in an ongoing process, not a finished result. v2 is expected to do meaningfully better,
whether that comes from more training compute, an architecture change informed by v1's ablations,
or both.

## Status

- Architecture, COCO data path, losses, training, evaluation, profiling, and ONNX export are implemented.
- Synthetic forward/loss/backward and data tests are included.
- Full COCO2017 training is complete: 100 epochs on a single RTX 5090, peak val AP 0.358 at epoch 95.
  See [Training Progress](#training-progress) below for the full curve.
- No "state of the art" or "beats YOLO/DETR" claim is made. Peak AP (0.358) is below the 40-50
  range originally set as a competitiveness bar against 20-40M-parameter real-time detectors.
  This is a real, single-seed, un-pretrained result, not a benchmark claim. See the
  [Minimum Validation Protocol](#minimum-validation-protocol) for what a competitive claim actually
  requires: controlled baselines, three seeds, and ablations, none of which have been run yet.

## Training Progress

<img src="assets/ap_progress.svg" alt="Validation AP by epoch across the full 100-epoch run, rising from near-zero to a peak of 0.358 at epoch 95" width="900">

| | |
|---|---|
| Parameters | 40.8M |
| Epochs | 100 / 100 (complete) |
| Peak val AP (IoU 0.50:0.95) | 0.358 (epoch 95) |
| Final val AP (epoch 100) | 0.356 |
| Val AP50 / AP75 | 0.544 / 0.381 |
| Val AP small / medium / large | 0.188 / 0.387 / 0.493 |
| Val AR@100 | 0.573 |
| Train loss | 4.11 (from 15.77 at epoch 1) |
| Throughput | ~90-110 img/s, batch 32, RTX 5090 |
| Single-frame inference | 30.7 ms / 32.5 FPS (batch 1, eager, RTX 5090) |

Train loss over the same run, log-shaped as expected for a detector trained from scratch:

<img src="assets/loss_progress.svg" alt="Train loss by epoch, dropping sharply in the first few epochs then declining gradually from about 6 to 4.1 over the rest of training" width="900">

AP50 (IoU 0.50, a looser localization threshold) against AP75 (IoU 0.75, stricter). Both climb
together early on, then AP75 plateaus lower, meaning coarse localization improved faster than
precise localization did:

<img src="assets/ap50_75_progress.svg" alt="AP50 and AP75 by epoch as two lines, AP50 rising to about 0.54 and AP75 to about 0.38" width="900">

AP broken out by object size. Large objects are detected far more reliably than small ones
throughout training, a common pattern for this class of detector and one of the things the
required ablations would need to explain:

<img src="assets/ap_by_size_progress.svg" alt="AP by object size (small, medium, large) by epoch as three lines, large objects highest at about 0.49, medium at about 0.39, small lowest at about 0.18" width="900">

Peak AP came at epoch 95 (0.3578), not the final epoch. That is normal late-training fluctuation,
and `best.pt` correctly holds the epoch-95 weights rather than epoch 100's. Against the 40-50 AP
range set as the competitiveness bar, this run falls short. It is a real result from a genuine
from-scratch run, on the low end of what the project was aiming for.
The gap is exactly what the [Minimum Validation Protocol](#minimum-validation-protocol) exists to
characterize properly: whether more decoder layers, more latents, longer training, or pretraining
would close it is unknown without actually running those ablations.

## Examples

Six detections from the final EMA checkpoint (epoch 95, peak AP) on COCO val2017 images,
confidence &ge; 0.35. Picked for variety, not cherry-picked for perfection.

| | |
|---|---|
| ![A tray of Japanese desserts and a sake bottle, each item individually boxed](assets/examples/food_tray.jpg) | ![A child on a skateboard with a parent, skateboard detected at high confidence](assets/examples/skatepark.jpg) |
| Dessert tray, bottle, bowl, and a dozen individually boxed items | Skate park, parent and child, skateboard at 0.93 confidence |
| ![A street market with a person at 0.94 confidence and umbrella awning detected](assets/examples/street_market.jpg) | ![A rodeo scene with multiple people correctly boxed in a crowd](assets/examples/rodeo.jpg) |
| Street market, person at 0.94, market umbrella | Rodeo, crowd of people, one animal still mislabeled |
| ![A parent holding a baby at a table, with cup, chair, and bottle detected](assets/examples/family.jpg) | ![A rainbow kite in flight above a beach with people and a distant boat](assets/examples/kite_beach.jpg) |
| Family at a table, cups, chair, bottle, two people | Beach, kite at 0.80, people along the shore, a distant boat |

## Live Tracking Demo

<img src="assets/tracking_demo.gif" alt="ObjectModel-v1 detections chained through a SORT tracker on real pedestrian footage, boxes and ids tracking people across frames" width="420">

[Full-length video (20s, MP4)](assets/tracking_demo.mp4)

ObjectModel-v1 itself has no temporal component. Every frame is detected independently. The clip
above chains detections through a from-scratch SORT-style tracker (`src/objectmodel_v1/tracking.py`:
constant-velocity Kalman motion model plus IoU/Hungarian frame-to-frame association) to give boxes
a persistent id and a short motion trail. Source footage is `vtest.avi`, OpenCV's standard
pedestrian test clip (BSD-3, ships with OpenCV), genuine video the model never trained on.

This is the same clip run at three points in training, tracker unchanged throughout. Only the
detector's checkpoint improved:

| Checkpoint | Ids issued over 20s | Longest-lived ids |
|---|---|---|
| Epoch 13 (AP 0.227) | ~89 | none survive past a few seconds |
| Epoch 26 (AP 0.288) | ~84 | 3 ids survive nearly the full clip |
| Epoch 95 (AP 0.358, peak) | 92 | 4 ids survive nearly the full clip |

Track count issued does not fall much, since new people keep entering frame throughout the clip
and each one earns a new id, which is correct behavior. Track persistence for people already in
frame improved consistently instead. A separate test on a genuinely different scene, an eye-level
warehouse clip not shown here, also surfaced a real and distinct limitation: the detector still
occasionally hallucinates objects on plain background surfaces, reading a support pillar as
"refrigerator", even at this peak checkpoint. Tracking quality rides on detection quality, and
detection quality on out-of-domain footage (camera angles, lighting, and compression that COCO's
photos do not really cover) is visibly weaker than on COCO's own validation images.

## Restricted Zone Detection

<img src="assets/zone_demo.gif" alt="A marked polygon zone over a plaza walkway, with tracked people highlighted and a red ALERT banner firing when someone enters the zone" width="420">

[Full-length video (20s, MP4)](assets/zone_demo.mp4)

A small layer on top of the tracker: `src/objectmodel_v1/zones.py` defines a polygon in the same
pixel coordinates as the tracked boxes, and fires an event only on the transition into or out of
it, not on every frame a track spends inside. It keys off each box's bottom-center point (roughly
where feet touch the ground) rather than the box centroid, since a zone drawn on a floor plane
should care where someone is standing, not where their torso is.

Correctness is checked on a synthetic sequence, not just eyeballed on video: a track walking
through a rectangle produces exactly one `entered` and one `exited` event, none of the frames
spent inside re-fire, a track that never enters never fires, and a fresh id gets independent
state. That test lives with the module.

On the same plaza clip used above, the zone above the walkway fired 91 enter/exit events over 20
seconds. Read that number carefully: most of it is not 91 different real crossings. It is the
same tracker id churn already described in the section above, the same person's track resetting
and re-entering the zone as a "new" id, plus a handful of misclassified objects (a backpack or
handbag read as its own tracked object) that shouldn't have triggered an event at all. The event
mechanics are verified correct; the real-world event count is exactly as reliable as the
underlying detector and tracker are, which right now is "usable for a demo, not for anything
where a false alert has a real cost."

```python
from objectmodel_v1.zones import RestrictedZoneMonitor

zone = [(300, 150), (650, 150), (650, 320), (300, 320)]  # pixel-space polygon
monitor = RestrictedZoneMonitor(zone)

for frame_index, frame in enumerate(video_frames):
    boxes, labels, scores = detect(frame)
    tracks = tracker.update(boxes, labels, scores)
    for event in monitor.update(tracks, frame_index):
        print(event.track_id, event.kind, event.position)  # "entered" or "exited"
```

## Architecture

```text
image
  -> compact convolutional backbone (strides 8/16/32)
  -> top-down pyramid fusion
  -> pooled multi-scale tokens
  -> fixed latent memory (global semantics)
  -> learned object queries
       -> query self-attention
       -> cross-attention to latent memory
       -> local sampling around the current query box
       -> iterative class and box prediction
  -> object set (no anchors, no NMS)
```

The local sampling radius scales with each query's current width and height. Early decoder
layers can search broadly; later layers focus naturally as boxes are refined. During training,
an optional dense auxiliary head adds one-to-many spatial supervision. It is discarded for
inference and must be evaluated as an ablation, not assumed to help.

## Installation

Use Python 3.11 or another PyTorch-supported Python version:

```bash
python3.11 -m venv .venv
.venv/bin/pip install --upgrade pip
.venv/bin/pip install -e '.[coco,export,dev]'
```

For a specific CUDA build, install the matching PyTorch wheel first using the command from
<https://pytorch.org/get-started/locally/>, then install ObjectModel-v1.

## Data

The default configuration expects COCO 2017:

```text
/path/to/coco/
  annotations/instances_train2017.json
  annotations/instances_val2017.json
  train2017/*.jpg
  val2017/*.jpg
```

Category IDs are mapped to contiguous training labels and converted back during evaluation.
Images without target objects are supported.

## Commands

Profile the model before allocating training compute:

```bash
objectmodel-profile --config configs/objectmodel_v1.yaml --device cuda
```

Overfit a small dataset first. A full single-GPU command is:

```bash
objectmodel-train \
  --config configs/objectmodel_v1.yaml \
  --data-root /path/to/coco \
  --output outputs/objectmodel_v1
```

Distributed training:

```bash
torchrun --standalone --nproc_per_node=8 -m objectmodel_v1.train \
  --config configs/objectmodel_v1.yaml \
  --data-root /path/to/coco \
  --output outputs/objectmodel_v1
```

Resume and override configuration values:

```bash
objectmodel-train \
  --config outputs/objectmodel_v1/config.yaml \
  --data-root /path/to/coco \
  --output outputs/objectmodel_v1 \
  --resume outputs/objectmodel_v1/last.pt \
  --set train.batch_size=8
```

Evaluate the EMA checkpoint with canonical `pycocotools` metrics:

```bash
objectmodel-eval \
  --config outputs/objectmodel_v1/config.yaml \
  --checkpoint outputs/objectmodel_v1/best.pt \
  --data-root /path/to/coco
```

Export raw logits and normalized `cxcywh` boxes to ONNX:

```bash
objectmodel-export \
  --config outputs/objectmodel_v1/config.yaml \
  --checkpoint outputs/objectmodel_v1/best.pt \
  --output outputs/objectmodel_v1/objectmodel-v1.onnx
```

Track detections across video frames (see [Live Tracking Demo](#live-tracking-demo); this is a
post-processing layer over independent per-frame detections, not a model capability):

```python
from objectmodel_v1.tracking import SortTracker

tracker = SortTracker(iou_threshold=0.3, max_age=5, min_hits=2)
for frame in video_frames:
    boxes, labels, scores = detect(frame)  # your decode_predictions() call
    for t in tracker.update(boxes, labels, scores):
        print(t.id, t.box, t.label, t.score)
```

## Minimum Validation Protocol

Before describing ObjectModel-v1 as competitive, run all models on the same COCO train2017 and
val2017 data, image resolution, augmentation budget, training epochs, and hardware. Report:

- COCO AP, AP50, AP75, APS, APM, and APL.
- Parameters, FLOPs/MACs, FP32/FP16/INT8 artifact sizes.
- End-to-end batch-1 median and p95 latency, including preprocessing and decoding.
- Peak training and inference memory, GPU-hours, epochs, and images seen.
- Three seeds for the principal result, with mean and standard deviation.
- Results both from random initialization and with the same permitted pretraining.

Required ablations:

| Experiment | Question |
|---|---|
| latent memory vs flattened feature attention | Does compression preserve useful global context? |
| local sampler disabled | Does high-resolution geometric evidence improve localization? |
| fixed vs box-scaled offsets | Does coarse-to-fine sampling matter? |
| dense auxiliary head disabled | Does added supervision improve convergence? |
| 1/2/3 latent layers | Where is the accuracy/latency optimum? |
| 32/64/96 latents | How aggressively can global context be compressed? |
| 3/4/6 decoder layers | What is the anytime speed/accuracy curve? |

Suggested external baselines are RT-DETR-R18, D-FINE-N/S, LW-DETR-T/S, and YOLOX-S. Use
their official implementations and report their license and measurement setup separately.

## Possible Extensions and v2 Directions

Two different questions worth separating: what this pipeline could be extended to do without
touching the detector, and what would actually make the detector itself better in v2.

**Extensions on top of the current detector and tracker.** The zone monitor above is one instance
of a general pattern: detector output plus `SortTracker`'s persistent ids is enough to build most
counting and monitoring logic without retraining anything.

- Unique counting: tracks already carry stable ids, so counting distinct people or objects through
  a scene (footfall, event attendance, traffic counts) needs no new detection work, just tallying
  ids instead of per-frame boxes.
- Movement heatmaps: the demos already compute per-track position trails for the overlays. The same
  data aggregated across a full video gives a map of where traffic actually concentrates.
- Dwell time and loitering: track how long a given id stays inside a zone instead of only its
  entry and exit. A per-id timer, a small addition to `zones.py`'s pattern.
- Line crossing: detect a track crossing a line in a given direction instead of entering a polygon.
  A doorway in/out counter. Same shape as `RestrictedZoneMonitor`, a line-side test instead of
  polygon containment.
- Class-based counting: the detector already classifies 80 COCO categories, so counting vehicles,
  bicycles, or any other class is the same unique-counting idea applied to a different label.

Behavior or anomaly detection (wrong-way movement, erratic motion) and multi-camera re-identification
are real extensions but not close additions. Both need work this repository does not do yet:
proper motion modeling beyond a Kalman filter, and cross-camera identity matching.

**What v2 would actually need to change.** Three things came out of this run as specific, not
generic:

- Small objects are detected far worse than large ones (APS 0.188 vs APL 0.493, see
  [Training Progress](#training-progress)). That gap is bigger than the usual small-vs-large
  spread in this class of detector, and points at the fixed latent memory or the local sampler's
  pyramid resolution as the first places to look, not just "train longer."
  The latent-layer and latent-count ablations above exist specifically to test that.
- Detection quality drops visibly on footage that does not look like COCO's own photography
  (see [Live Tracking Demo](#live-tracking-demo)), which limits every downstream extension listed
  above. Broader training data or augmentation is the direct lever; the required ablations don't
  cover this gap, which is worth adding to that list rather than assuming architecture changes
  alone fix it.
- v1 trained from random initialization only. The Minimum Validation Protocol already calls for a
  pretrained comparison, and it is the single change most likely to move AP the most, based on how
  much of a difference pretraining typically makes at this parameter count in the baselines cited
  under Research Basis.

None of this is scheduled work, it's the concrete list of what the data from this run actually
points at, as opposed to a generic "train bigger, train longer."

## Research Basis

ObjectModel-v1 builds on published, independently attributable ideas:

- DETR: set prediction and Hungarian matching.
- Conditional and Deformable DETR: spatially conditioned/local sparse attention.
- RT-DETR: efficient separation of multi-scale encoding and query decoding.
- D-FINE: evidence that fine-grained iterative localization is valuable.
- DEIM: evidence that one-to-one matching benefits from denser training supervision.
- LW-DETR: evidence that compact transformer detectors can compete with real-time CNNs.

ObjectModel-v1's specific hypothesis is the combination of a **fixed compressed global memory** and
**box-scaled local pyramid sampling**. Publication novelty requires a broader prior-art search
and empirical ablations; this repository does not claim that the combination is patent-new.

## License

Apache License 2.0. Dataset images, annotations, pretrained weights, and external baselines
retain their own licenses and are not included.