File size: 25,993 Bytes
af02367
 
c7c1d80
 
 
 
 
 
 
 
 
 
 
 
 
af02367
c7c1d80
a873f95
c7c1d80
03d386d
c7c1d80
 
 
 
 
 
 
 
13fbc30
c7c1d80
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1bf7c97
 
 
 
 
c7c1d80
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cf1bb91
 
 
 
c7c1d80
 
 
 
 
cf1bb91
c7c1d80
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
be8c031
c7c1d80
cf1bb91
c7c1d80
 
 
 
 
be8c031
 
 
c7c1d80
be8c031
c7c1d80
 
 
 
 
 
 
 
 
 
 
be8c031
c7c1d80
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
---
license: apache-2.0
library_name: tf-lite
pipeline_tag: object-detection
tags:
  - object-detection
  - tflite
  - int8
  - quantization
  - edge
  - robotics
  - ftc
  - limelight
  - mobile
  - ssd
---

# RepLiteDet-S0: a compact FTC 2026-2027 BioBuzz Pollen Detector Model for the Limelight 3A

A drop-in 8-bit TFLite model for FTC teams working in season 2026-2027 BioBuzz to detect the game element Pollen. Built by FTC 8565. The complete training scripts follow in the next release.

![architecture](charts/replitedet_dynamic_flow_v2.gif)

*Backbone -> FPN-Lite fusion -> shared SSD head -> decode + NMS. The diagram shows the
training graph (12,750 anchors); the exported detector drops the P5/P6 heads and runs
P3 + P4 = 12,000 anchors.*


## Brief Comparison to the LimeLight3A Default Model

| | RepLiteDet-S0 (ours) | SSD-MNV2, our pipeline | SSD-MNV2, Limelight online trainer |
|---|---|---|---|
| mAP@0.5 | **0.9365** | 0.8516 | 0.4605 |
| mAP@0.75 | **0.7581** | 0.5746 | 0.3044 |
| recall `<10 px` | **0.8213** | 0.6823 | 0.3285 |
| false positives | **410** | 845 | 863 |
| size | **0.64 MB** | 4.96 MB | 4.96 MB |
| latency, **on a Limelight 3A** | **52.8 ms** | 86.5 ms | 86.5 ms* |

**+10.0% mAP@0.5 and +20.4% recall on objects under 10 px than the same network trained the
conventional way, at 0.64 MB against 4.96 MB and 39% less latency measured on a Limelight
3A.** Against the model the standard cloud-training route produces from the same dataset at
the same budget: **+103% mAP@0.5 and 2.5x the sub-10 px recall**. Ours runs at 320x320
while both baselines run at 300x300, so it is doing more work per frame, not less.

The weak axis is box tightness on large close-up objects; the numbers for that are below,
not hidden.


## Side-by-side, on real frames

**Read the boxes like this:**

- **Green** -- a detection that matches a real object (IoU >= 0.5). This is a hit.
- **Yellow** -- a real object the model **missed**: it is there in the ground truth and the
  model did not find it at its own threshold.
- **Red** -- a **false positive**: a box that matches no real object, either because it is on
  the wrong thing entirely or because it is on the right thing but positioned or sized badly
  enough to fail the IoU test.

Ground-truth boxes are not drawn -- a fourth box per object made the frames unreadable.
Nothing is hidden: misses show yellow, unmatched detections show red, and each tile prints
how many of the frame's objects that model found, at its own best-F1 threshold.

![contrast](panels/panel_contrast.png)

**These frames are selected, and the rule is printed on the figure**: every frame where we
find at least 2 more objects than the retrained baseline does. Across the four, **19 of 20**
found by us, **9** by the baseline, **1** by the online trainer. It shows what the difference
looks like when it appears -- the tables above are what it is worth on average. Same three
models, each at its own threshold.


## Deploying to a Limelight 3A

Upload the 8-bit `.tflite` and the labels file through the Limelight web interface and select the **CPU** runtime, per Limelight's own detector documentation. Both files are required; a model without its labels file will not run.


## What these numbers mean

The camera feeds a 320x320 image. A ball across the field is under 10 px wide; one at intake
range is 50 px or more. Every metric below is measured on 3,546 real objects.

**"Overlap" always means shared area divided by combined area (IoU).** A box twice the size
of the ball scores 0.5 even though it contains all of it -- too big is penalised exactly as
hard as too small.

| metric | what it measures | on the robot |
|---|---|---|
| **mAP@0.5** | did it find the ball, with a box overlapping the real one by at least 50% | "there is a ball over there" -- enough to drive toward it |
| **mAP@0.75** | same, at a 75% bar -- a loose box now counts as a miss | enough to line up an intake, not just point the robot |
| **mAP@[.5:.95]** | the average across ten strictness levels from 50% to 95% | one number for "finds it AND boxes it well" |
| **mean best IoU** | for each real ball, the overlap of the best box we drew -- averaged over **every** ball, so one we never found counts as ~0 | box size is how you estimate distance; a loose box is a wrong distance |
| **TP / FP** | correct detections / boxes matching no real ball | FP is the robot reacting to something that is not there |
| **best-F1** | the balance of "found most of them" against "was not wrong often" | the single number if you only get one |
| **own threshold** | the confidence cutoff at which that model scores best | each model has a different one; comparing two models at the same cutoff is unfair |
| **NMS** | non-maximum suppression: the model scores thousands of candidate boxes and a real object lights up several neighbours at once, so NMS keeps the highest-scoring box and deletes anything overlapping it by more than a threshold | too permissive stacks three boxes on one ball; too aggressive merges two balls into one. Ours ships at 0.35 |
| **recall** | of the balls that really are there, the fraction the model finds. Says nothing about box quality -- a box that scrapes the 50% bar counts the same as a perfect one | a miss is a ball your robot never knows about |
| **recall by band** | the same number, computed separately for each object-size range | size is distance: `<10 px` is the far half of the field, `>45 px` is intake range |

### Worked example: the two `<10 px` numbers

The eval pool holds **554 balls under 10 px** -- the far half of the field. Two numbers describe how we do on them, and they answer different questions.

**recall 0.8213 = how many we found.** Of those 554 balls, the model produced a box that
counted as a hit on **455**, and missed **99**. It says nothing about box quality: a box
that scrapes past the 50% bar counts the same as a perfect one. This is the number that
decides whether the robot knows a far ball exists at all.

**mean best IoU 0.6669 = how well we boxed them, averaged over all 554.** For each ball we
take the best box we drew and measure the overlap, then average -- *including* the 99 we
never found, which contribute about 0. So this one number blends finding and boxing: the
misses drag it down. Back out the 82% we found and the boxes on those sit near **0.81**.

Read together: **we find 4 far balls in 5, and box the ones we find well.** The same pair
for the retrained baseline is 378 of 554 found (recall 0.6823) at mean IoU 0.5716 -- and
backing out its recall the same way gives ~0.84 on what it finds. **Almost the whole gap in
that band is finding the ball, not boxing it.**

So `<10 px` recall **0.8213 vs 0.3285** means: of the balls far enough away to be under
10 pixels, we find about **4 in 5** and the cloud-trained model finds about **1 in 3**. In a
match, that is seeing the far side of the field during auto instead of driving in blind.


## What you get

| file | what it is | release road map |
|---|---|---|
| `replitedet_s0_int8.tflite` | the model -- 0.64 MB, full-INT8, uploads straight to a Limelight 3A | **available today** |
| `pollen_labels.txt` | the labels file the device requires beside it | **available today** |
| `README.md` | every number here, and how it was measured | **available today** |
| `ckpt-*` / `checkpoint` | the float checkpoint, so you can retrain on YOUR data | **available today** |
| `*_vp2.config` | the exact training recipe we used -- nothing held back | **available today** |
| training scripts and pipeline | our training pipeline, with the tuned learning-rate schedule | *releasing later* |

## Design

Built backwards from one target: an FTC robot, watching a field, through a Limelight 3A.

- **43.9%** of training objects are under 18 px, so the neck fuses high-resolution early
  features with deeper ones. That is where the `<10 px` numbers come from.
- The device runs **full-integer INT8 only**. No float, float16 or weight-only fallback.
- The head is a **stock SSD head** with in-graph `TFLite_Detection_PostProcess`, so the
  export is drop-in: upload the tflite and labels, pick the CPU runtime. A more modern head
  would score better and would not be drop-in.
- **0.86 M parameters.** COCO pre-training is done on *our* architecture,
  **class-agnostically** (every annotation -> one `object` class), so we keep the
  representation and carry none of the 80-class head width: **0.64 MB vs 5.20 MB**.


## Where this model stands

The baseline is the network the Limelight training route produces -- plain SSD MobileNetV2
300x300, alpha 1.0 -- **retrained by us**, from its own COCO checkpoint, on our data, at the
same 24,000 steps, with a correctly batch-scaled learning rate and its exported NMS threshold
swept the same way ours was. The third comparison is that same route's own trainer, run on
the same data at the same budget.

### Why the baseline is a legitimate comparison: it is the same network, verified

Same architecture, our rebuild against LimeLight Trainer download:

| | our retrained baseline | the Limelight-trained artifact |
|---|---|---|
| file size | **5,202,224 bytes** | **5,202,224 bytes** |
| CONV_2D / DEPTHWISE_CONV_2D | 55 / 17 | 55 / 17 |
| ADD / RESHAPE / CONCATENATION | 10 / 13 / 2 | 10 / 13 / 2 |
| QUANTIZE / DEQUANTIZE / LOGISTIC | 1 / 2 / 1 | 1 / 2 / 1 |
| `TFLite_Detection_PostProcess` | 1 | 1 |

Byte-identical size, identical operator histogram: plain SSD MobileNetV2 300x300, alpha 1.0,
six levels, 1,917 anchors, 1x1 convolutional box predictor. Architecture verified; their
training *recipe* is not visible to us, which is what the three models below separate.

### Separating the architecture from the training pipeline

A better network, or the same network trained better? Three models, one variable each.

| model | architecture | trained by | batch | steps |
|---|---|---|---|---|
| **RepLiteDet-S0** (this release) | RepLiteDet-S0 | our pipeline | 8 | 12k / 24k |
| SSD-MobileNetV2 300 | SSD-MobileNetV2 300 | our pipeline | 8 | 12k / 24k |
| SSD-MobileNetV2 300 | SSD-MobileNetV2 300 | **the Limelight online trainer** | 16 | 12k / 24k |

* **RepLiteDet-S0 vs SSD-MobileNetV2 in our pipeline** isolates the architecture.
* **SSD-MobileNetV2 in our pipeline vs the same network from the Limelight online trainer**
  isolates the training pipeline, confounded with batch 16 vs 8 -- which favours the online
  trainer, since at equal steps batch 16 sees twice the images.

**SSD-MobileNetV2 300, trained two ways.** Same network, same data, **24,000 steps both
sides** (the online-trainer run is job `eb8961b1`, its log recording `Steps: 24000, Batch
size: 16`), each read at its own best-F1 threshold:

![the Limelight online trainer run](charts/24k-limelight-trained.png)

*The online-trainer run as configured: 24,000 steps, target Coral / CPU, variant **Default**
-- we changed nothing but the step count, which is the point. 141 minutes on their H100. The
trainer's own advice is on the same screen: 4,000 steps for testing, 20,000 for robot use.*

| | mAP@0.5 | mAP@0.75 | mAP@[.5:.95] | best-F1 | mean IoU | TP | FP |
|---|---|---|---|---|---|---|---|
| trained by **our pipeline** | **0.8516** | **0.5746** | **0.5426** | **0.8531** | **0.7849** | **3266** | 845 |
| trained by the **Limelight online trainer** | 0.4605 | 0.3044 | 0.2875 | 0.5092 | 0.5244 | 1506 | 863 |

**+85% mAP@0.5 from the training pipeline alone**, with the confound against us: batch 16
sees twice the images ours does at equal steps. The online-trainer model finds 1,506 of
3,546 objects to our 3,266.

Neither pipeline gains from the second 12k. The online trainer at 12k -> 24k: mAP@0.5 0.4509
-> 0.4605 but best-F1 0.5152 -> 0.5092 and TP 1758 -> 1506; it becomes more conservative, not
better. One axis it wins outright: **12 spurious detections on 118 empty frames, against
RepLiteDet-S0's 17 and 90 for the same network in our pipeline.** It hallucinates least
because it detects least.

### By source -- the two that carry the weight

Per source, not averaged: `pollen_clean` is small-object-heavy, `clement_pdp` is the easier
one, and a mean across them would read as uncertainty when it is really the gap between two
distributions. `pollenv1` (153 GT, 4.3%, our own smoke-test set) is reported further down and
included in the pooled numbers, but is too small to head a table.

| source | GT | ours mAP@0.5 | ours mAP95 | baseline mAP@0.5 | baseline mAP95 |
|---|---|---|---|---|---|
| `clement_pdp` | 1131 | 0.9910 | 0.8482 | 0.9578 | 0.7552 |
| `pollen_clean` | 2262 | 0.8985 | 0.5579 | 0.7961 | 0.4408 |

### Full pool (all 3,546 GT), each model at its OWN best-F1 threshold

| model | own thr (n=1) | mAP@0.5 | mAP@[.5:.95] | best-F1 | mean IoU | FP |
|---|---|---|---|---|---|---|
| **RepLiteDet-S0 (ours)** | 0.2 | 0.9365 | 0.6694 | 0.9029 | 0.8314 | 410 |
| SSD-MobileNetV2 300 (baseline) | 0.05 | 0.8516 | 0.5426 | 0.8531 | 0.7849 | 845 |
| SSD-MobileNetV2 300, Limelight online trainer | 0.2 | 0.4605 | 0.2875 | 0.5092 | 0.5244 | 863 |

### By object size -- the reason this model exists

Size is distance: `<10 px` is the far half of the field, `>45 px` is intake range.

**Recall -- the fraction of objects in that band the model finds.**

| band | GT | ours | SSD-MNV2, our pipeline | SSD-MNV2, LL online trainer |
|---|---|---|---|---|
| `<10` | 554 | **0.8213** | 0.6823 | 0.3285 |
| `10-18` | 1001 | **0.9850** | 0.9351 | 0.3487 |
| `18-28` | 495 | **0.9960** | 0.9758 | 0.8121 |
| `28-45` | 654 | **0.9908** | 0.9786 | 0.7003 |
| `>45` | 842 | **0.9846** | 0.9857 | 0.9276 |

**Mean best IoU -- how well the box lines up, averaged over every object in the band
(a missed one counts ~0, so this blends finding with boxing).**

| band | ours | SSD-MNV2, our pipeline | SSD-MNV2, LL online trainer |
|---|---|---|---|
| `<10` | **0.6669** | 0.5716 | 0.3209 |
| `10-18` | **0.8164** | 0.7536 | 0.2575 |
| `18-28` | **0.9009** | 0.8549 | 0.6858 |
| `28-45` | **0.9076** | 0.8752 | 0.6054 |
| `>45` | **0.8573** | 0.8512 | 0.8176 |

We lead every band on both channels except `>45 px` recall, where the baseline is ahead by
0.0011 -- inside noise, and the band we call out as our open problem.

### Per source, including the one we exclude from the headline

| source | GT | RepLiteDet-S0 | SSD-MNV2, our pipeline | SSD-MNV2, LL online trainer |
|---|---|---|---|---|
| `clement_pdp` | 1131 | 0.9910 | 0.9578 | 0.5389 |
| `pollen_clean` | 2262 | 0.8985 | 0.7961 | 0.4289 |
| `pollenv1` *(smoke-test set, 4.3% of GT)* | 153 | 0.9894 | 0.9034 | 0.4823 |

`pollenv1` is an early set we collected ourselves to smoke-test the pipeline. At 153 of 3,546 GT (4.3%) it is too small to support a claim on its own, so we report it and keep it out of the headline -- but it stays in the pooled numbers.


## Context: how this compares to YOLO

Context, not a claim. **At float, YOLO is the stronger detector** -- 0.85 mAP@[.5:.95]
against our 0.67 -- and that is the bar we aim at. It is listed because it is the obvious
thing a team reaches for.

| model (same data, same 24k budget, same eval harness) | mAP@0.5 | mAP@0.75 | mAP@[.5:.95] |
|---|---|---|---|
| yolov8n, **float** | 0.9937 | 0.9523 | 0.8465 |
| yolo11n, **float** | 0.9926 | 0.9521 | 0.8455 |
| yolov8n, INT8 TFLite | 0.8640 | 0.5422 | 0.5158 |
| yolo11n, INT8 TFLite | 0.8396 | 0.4672 | 0.4699 |
| **this model, INT8 TFLite** | 0.9365 | 0.7581 | 0.6694 |

**A Limelight 3A cannot run it.** The device accepts only full-integer INT8 with a `uint8`
NHWC input -- float32, float16 and weight-only (`w8a32`/`w8a16`) are rejected at load. A
stock YOLO export also emits one `[1, 4+nc, N]` tensor where the device needs 3-4 separate
float32 outputs; the INT8 rows above required a custom export wrapper we wrote.

Through the one mode the device permits, YOLO loses **0.33-0.38 mAP@[.5:.95]**, and the loss
is box tightness, not detection -- same objects found, looser boxes, confidences compressed
into ~[0, 0.5]. Independently reported: YOLOv8 n/s/m/l/x lose **6.1-7.5 mAP points** and
YOLO11 n/s/m/l/x **6.6-8.7** under LiteRT full-integer INT8, against 0.0-0.2 for
dynamic-range:

> Ozturk et al., *Benchmarking Quantized YOLO Models for Real-Time Object Detection on
> Edge Devices*, arXiv:2511.13453. https://arxiv.org/pdf/2511.13453

Same weights, four export widths, identical calibration:

| yolov8n, same weights | mAP@0.5 | mAP@0.75 | box precision @320px |
|---|---|---|---|
| float32 | 0.9937 | 0.9523 | 0.0000018 px |
| int8 weights, float activations (`w8a32`) | 0.9940 | 0.9189 | 0.0000003 px |
| int8 weights, int16 activations (`w8a16`) | 0.9936 | 0.9329 | 0.0114 px |
| **full-integer int8 (the only mode this device runs)** | 0.8640 | **0.5422** | **1.584 px** |

The last column is the mechanism. YOLO packs box coordinates and class scores into **one
output tensor sharing a single 8-bit scale**; at full-integer INT8 that scale is worth
**1.58 px of box resolution** at 320 px input, on a dataset whose median object is under
18 px. It also caps every confidence at 0.4998. Widen the activations and the damage is
gone -- but those are the modes the device rejects.

> Two limits on these rows: (1) it is a property of the *TFLite full-integer export path*,
> not of YOLO -- the same architectures quantize near-losslessly on Hailo (-1.6%), TensorRT
> (-0.05) and OpenVINO (-2.7%); (2) known mitigations exist (splitting the head's outputs,
> excluding the score sigmoid) and we did not implement them, so a better INT8 YOLO than this
> one certainly exists.


## Latency

**Measured on a Limelight 3A** -- one-minute running average, on the device:

| model | input | on device | size |
|---|---|---|---|
| **RepLiteDet-S0 (ours)** | 320x320 | **52.8 ms** | **0.64 MB** |
| SSD-MobileNetV2 300 | 300x300 | 86.5 ms | 4.96 MB |

**39% faster on the device, while running a larger input** -- 320x320 against their 300x300,
so it is doing more work per frame, not less.

*The two SSD-MobileNetV2 rows in the headline table are the **same network**; the device
figure was measured once for that architecture and applies to both.*


## False positives, broken down -- read this before reacting to the raw count

**410 false positives** across 3546 objects, at its own best-F1 threshold. The three kinds behave very differently on a robot:

| kind | count | share | what it actually is |
|---|---|---|---|
| **localization** | 319 | 78% | found a REAL object, but the box is too loose to count as a hit (IoU < 0.5) |
| background | 91 | 22% | a detection where there is no object -- a genuine hallucination |
| duplicate | 0 | 0% | a second box on an object already detected |

Duplicates are eliminated by the exported NMS threshold, not by training. Hallucinations are
the smallest category:

### Spurious detections on empty frames

Images containing no objects at all, which isolates hallucination:

| model | false positives | per empty image |
|---|---|---|
| **this model** | 17 | **0.1441** |
| SSD-MobileNetV2 300, our pipeline | 90 | **0.7627** |
| SSD-MNV2 300, Limelight online trainer | 12 | **0.1017** |

Fewer than either comparison model. This is the number that predicts how often a robot acts
on nothing.

### Localization FP is the open problem

Most of our false positives are loose boxes on real objects, and that share has **not**
improved across milestones while duplicates went to zero and background FP fell. Same deficit
shows up as mAP@0.75 and as the weakness on very large close-up objects.

Cause is architectural, not a training shortfall: the anchor ladder tops out near **181 px**,
so larger objects have no anchor that fits. **Fixing the ladder is the next release's main
objective.**


## The data this was trained and measured on

Not published. The distribution is, so you can judge transfer to your field and camera.

| split | images | ground-truth boxes | negative images |
|---|---|---|---|
| train | 1,931 | 4,352 | 289 |
| valid | 670 | 1,953 | 70 |
| test | 618 | 1,593 | 59 |
| **evaluation pool (valid+test)** | **1,288** | **3,546** | **129** |

Every number in this card is measured on pooled **valid+test**. Size = shorter box side at a
320x320 reference:

| size band | GT boxes in eval | share |
|---|---|---|
| `<10 px` | 554 | 15.6% |
| `10-18 px` | 1001 | 28.2% |
| `18-28 px` | 495 | 14.0% |
| `28-45 px` | 654 | 18.4% |
| `>45 px` | 842 | 23.7% |

Training set: mean 39.3 px, median 23.4 px, **43.9% under 18 px** -- which is why this card
reports recall per band rather than pooled.

| source | eval GT | what it is |
|---|---|---|
| `clement_pdp` | 1,131 | the easier source: larger objects, median ~31 px |
| `pollen_clean` | 2,262 | the hard one: median ~13 px, ~26% under 10 px |
| `pollenv1` | 153 | an early set we collected ourselves to smoke-test the pipeline |

**Leakage check.** Video-derived, so split in chunks with guard frames dropped at
boundaries: **0 of 670** valid and **0 of 618** test images collide with training at a
perceptual-hash distance <= 3.


## Limits

- **Latency is a one-minute running average measured on a Limelight 3A**, taken with the vision pipeline running as it normally would -- what a robot actually sees, rather than an isolated benchmark figure.

- **Training data not published.** Reproducible in shape, not in absolute value.

- **Single class.** Multi-class retraining needs `num_classes` changed and a re-export.

- **We slightly lose on very large, close-up objects.** `>45 px` recall 0.9846 vs the baseline's
  0.9857, that is the only band it leads. Its ladder runs six levels to 1x1; ours tops out near
  181 px, so an object filling the view has no anchor that fits. Fixing in next release.

- **Every comparison here is 24,000 training steps on the same images**, each model read at its own best-F1 threshold. The Limelight online trainer runs at batch 16 against our 8, which favours it.

- **Post-processing was tuned, and the tuning costs band recall.** The exported NMS IoU is
  **0.35**, swept on the eval pool. Against 0.45 it is better on everything pooled --
  mAP@0.5 0.9365 vs 0.9311, best-F1 0.9029 vs 0.8837, FP 410 vs 508 -- and it pays for that
  in two size bands, `<10px` recall 0.8213 vs 0.8321 and `>45px` 0.9846 vs 0.9964, because
  pollen occurs in crowded clusters that tighter suppression merges. If your scene is
  cluster-heavy, 0.45 is a one-flag re-export and no retraining -- ask us, or do it yourself once the export path ships.


## Where this goes

For FTC 2026-2027 BioBuzz game element `Pollen` on a Limelight 3A: small enough to leave headroom, faster than the
alternatives measured here, more accurate than the same network trained conventionally, and
drop-in for the detector pipeline you already have. Box tightness on large objects is the
open axis and the next release's target.

Built to be retuned, and the checkpoint plus training pipeline are the next thing we ship.
If you run this one on your own hardware, we would like to hear what you measure.


## Retraining on your own data

The float checkpoint and the exact config are here. The training pipeline we used is not
yet -- it follows in the next release. What that means practically:

**This is not a stock TFOD architecture.** The config declares a feature extractor
`ssd_replitedet_s0_fpn_keras`, and the anchor generator is selected by the `_vp2` filename
suffix. Neither exists in a stock Object Detection API install, so `model_builder.build` on
this config fails with an unknown-feature-extractor error until our code is on the path.

If you would rather not wait, everything you need to write your own trainer is published:
the config carries the full recipe, the checkpoint carries the layer names and shapes, and
the anchor geometry -- the one thing NOT recoverable from either -- is written out below.

### The anchor geometry, in full

The config's `multiscale_anchor_generator` block declares 6 anchors per location, but our
generator overrides where they sit. Build stock anchors from that proto and you get a
silently worse model, because the head weights were trained against these:

| level | stride | grid | 6 anchors per cell, as (size px @ offset within the cell) |
|---|---|---|---|
| P3 | 8 | 40x40 | 16 @ (2,2), 16 @ (6,2), 16 @ (2,6), 16 @ (6,6), 32 @ (4,4), 45 @ (4,4) |
| P4 | 16 | 20x20 | 64, 90.5, 90.5, 128, 128, 181 -- all square, all centred at (8,8) |
| P5 | 32 | 10x10 | 128, 181, 181, 256, 256, 362 -- centred (head dropped at export) |
| P6 | 64 | 5x5 | 256, 362, 362, 512, 512, 724 -- centred (head dropped at export) |

Sizes are square, in pixels at the 320x320 input. P4-P6 are `base x {1, sqrt2, 2} x
{1, sqrt2}` with bases 64 / 128 / 256, which is why 90.5 and 128 each appear twice. The four
shifted 16 px anchors at P3 are the **virtual P2**: they give sub-cell coverage at stride 8
without paying for a stride-4 level.

The **shipped detector keeps P3 + P4 only** -- 9,600 + 2,400 = **12,000 anchors**. The P5
and P6 heads are dropped at export; the training graph has all four levels and 12,750.

Two things that will cost you a day if you miss them:

1. **Keep the config filename ending in `_vp2.config`.** Once our pipeline is out, the
   anchor generator is selected by filename. Rename it and you train stock anchors against
   a vP2 checkpoint, with no error message and a large accuracy loss.
2. **`fine_tune_checkpoint_type: full`** loads the detection heads too. `detection` loads
   only the backbone and leaves the heads random -- correct if you change the number of
   classes, wrong if you do not.


## License

**Apache License 2.0** -- see `LICENSE`. This covers the model weights, the
TFLite export and the training config shipped here. It does not cover the training data,
which is not published, nor the third-party baseline model referenced for comparison, which
is not redistributed here.


Models produced by the Limelight online trainer are **not** redistributed here. Train your own if you want to reproduce that row.