Chris Leo commited on
Commit
694fdc6
Β·
verified Β·
1 Parent(s): 321c5d0

scorevision: push artifact

Browse files
Files changed (1) hide show
  1. README.md +193 -0
README.md ADDED
@@ -0,0 +1,193 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Detect-crime Miner β€” Recipe to Beat the King
2
+
3
+ Target element: `manak0/Detect-crime` on subnet 423 (open-source / public track).
4
+
5
+ Read [ANALYSIS.md](ANALYSIS.md) first β€” it documents the king's model (the manak0 baseline)
6
+ and where the gap lives.
7
+
8
+ Current king of record (2026-05-04 leaderboard): hotkey `5CSeBY…tv9f`, score **0.576**.
9
+ Crime is **uncontested**: there is no `Detect-crime-winner` HF repo, and the king's score
10
+ is within rounding of the published baseline's `overall_iou` (0.597). Anybody who lands a
11
+ modest improvement takes the throne.
12
+
13
+ ## Layout
14
+
15
+ ```
16
+ crime_miner/
17
+ β”œβ”€β”€ ANALYSIS.md ← analysis of the king + scoring + constraints
18
+ β”œβ”€β”€ README.md ← this file
19
+ β”œβ”€β”€ miner.py ← deployable inference (multi-scale TTA + WBF + CLAHE)
20
+ β”œβ”€β”€ chute_config.yml ← chute resource spec (16 GB GPU, matches king's)
21
+ β”œβ”€β”€ class_names.txt ← target class order β€” DO NOT REORDER
22
+ └── training/
23
+ β”œβ”€β”€ DATASET.md ← dataset sources + pipeline (start here)
24
+ β”œβ”€β”€ build_dataset.py ← end-to-end builder: manako + Roboflow + COCO bat
25
+ β”œβ”€β”€ poll_manako.py ← background poller for in-domain frames + king's preds
26
+ β”œβ”€β”€ train.py ← two-stage YOLOv11 training (silver β†’ clean fine-tune)
27
+ β”œβ”€β”€ verify_dataset.py ← QA over assembled YOLO dirs
28
+ β”œβ”€β”€ export_onnx.py ← export with NMS baked in -> [1, 300, 6]
29
+ └── requirements.txt
30
+ ```
31
+
32
+ ## What the miner does differently
33
+
34
+ `miner.py` keeps the king's I/O contract (single `weights.onnx` β†’ `TVFrameResult`) but adds
35
+ six concrete improvements over the auto-generated `subnet_bridge` template the king ships:
36
+
37
+ 1. **Letterboxed input at 1280** instead of stretch-resized 640. Small objects (balaclava
38
+ ~30 px, glove ~25 px, spray paint can ~20 px) survive β€” the king's stretch resize
39
+ destroys them. This alone lifts recall on the four catastrophic classes.
40
+ 2. **Per-class confidence floors**. King uses one global 0.25 across all six classes; we
41
+ set `balaclava=0.05, bat=0.10, glove=0.05, graffiti=0.20, hoodie=0.20, spray paint=0.10`.
42
+ Synthetic-benchmark recalls were 0.034 / 0.143 / 0.064 / 0.321 / 0.274 / 0.161 β€” the
43
+ bottleneck is recall, and the FFPI cap has plenty of headroom (~6.5 preds/img today).
44
+ 3. **Multi-scale TTA** at `{1280, 1536} Γ— {orig, hflip}` = 4 forward passes, collapsed to 2
45
+ when the ONNX export is static-shape. Pro_6000 has the budget (latency p95 = 10 s).
46
+ 4. **Weighted Box Fusion** across TTA streams. WBF averages cluster boxes weighted by
47
+ score, which yields tighter localizations than always picking the highest-confidence
48
+ proposal β€” and tighter boxes mean more cases cross the IoUβ‰₯0.5 bar that the scorer uses.
49
+ 5. **CLAHE on dark frames only** (luma gate). Crime CCTV is night-heavy. King applies no
50
+ preprocessing.
51
+ 6. **Class-aware NMS at IoU=0.45**. King uses class-agnostic NMS, which suppresses
52
+ balaclava-on-hoodie or glove-near-bat overlaps. Class-aware keeps both.
53
+
54
+ Total ONNX inference cost on Pro_6000 with YOLOv11s + 2-scale TTA is well under 1 s/frame.
55
+
56
+ ## How to deploy
57
+
58
+ You need: a `weights.onnx` exported in `[1, 300, 6]` layout (NMS baked in) β€” produced by
59
+ `training/export_onnx.py` after training, OR you can ship the king's raw ONNX directly to
60
+ test the inference improvements alone.
61
+
62
+ ### Option 1 β€” drop-in test with the king's weights
63
+
64
+ Sanity-check that the inference improvements alone help, before training:
65
+
66
+ ```bash
67
+ cp /root/turbovision_crime/king_models/Detect-crime/weights.onnx ./weights.onnx
68
+ python miner.py # smoke test on /tmp/crime_proof.png
69
+ ```
70
+
71
+ Expected: with the king's weights but our miner.py, you should already see a noticeable lift
72
+ on the rare classes (recall driven up by the lower per-class conf floors and the 1280 input
73
+ that the dynamic-shape ONNX accepts). The king's published ONNX is **static** at 640Γ—640,
74
+ so the dynamic letterbox path won't help unless you re-export β€” see below.
75
+
76
+ ### Option 2 β€” train a real beating model
77
+
78
+ See [training/DATASET.md](training/DATASET.md) for full data-pipeline notes. Quick path:
79
+
80
+ ```bash
81
+ cd training
82
+ pip install -r requirements.txt
83
+
84
+ # 1) Start the manako poller in the background to accumulate in-domain frames
85
+ # (each rotation surfaces a fresh challenge ~every few minutes during active scoring).
86
+ python poll_manako.py --out ../manako_pool --interval 120 --forever &
87
+
88
+ # 2) Build the silver dataset. Combine manako frames (king-labeled), Roboflow
89
+ # per-class detection sets, and (optional) COCO baseball bat. Roboflow needs
90
+ # ROBOFLOW_API_KEY in env.
91
+ python build_dataset.py \
92
+ --out ../data \
93
+ --king-onnx /root/turbovision_crime/king_models/Detect-crime/weights.onnx \
94
+ --manako --manako-polls 30 --manako-poll-delay 120 \
95
+ --roboflow balaclava=brainster/balaclava-detection-v3 \
96
+ --roboflow glove=ppe-detection/gloves-v1 \
97
+ --roboflow graffiti=graffiti-detection/graffiti-v3 \
98
+ --roboflow "spray paint=tools/spray-paint-can-v1" \
99
+ --coco-bat /path/to/coco/instances_train2017.json /path/to/coco/train2017 \
100
+ --extra-dir ../manako_pool/images \
101
+ --min-conf 0.10 --keep-empty --intra-threads 16
102
+
103
+ # 3) Verify the assembled dataset
104
+ python verify_dataset.py --data ../data/data.yaml --visualize 20
105
+
106
+ # 4) Stage A: silver pretrain
107
+ python train.py --data ../data/data.yaml --weights yolo11s.pt \
108
+ --imgsz 1280 --batch 16 --stage A --epochs 200 --name crime_a
109
+
110
+ # 5) Build a clean set: hand-verify (or LLM-verify) ~300 manako frames into
111
+ # ../data_clean/data.yaml with the same YOLO layout.
112
+
113
+ # 6) Stage B: clean fine-tune
114
+ python train.py --data ../data_clean/data.yaml \
115
+ --weights ../runs/detect/crime_a/weights/best.pt \
116
+ --imgsz 1280 --batch 16 --stage B --epochs 50 --name crime_b
117
+
118
+ # 7) Export with NMS baked in -> [1, 300, 6]
119
+ python export_onnx.py --weights ../runs/detect/crime_b/weights/best.pt \
120
+ --imgsz 1280 --out ../weights.onnx
121
+ ```
122
+
123
+ ### Option 3 β€” deploy via the turbovision CLI
124
+
125
+ ```bash
126
+ cd /root/turbovision_crime
127
+ sv -vv deploy-os-miner --model-path scratch/crime_miner --element-id manak0/Detect-crime
128
+ ```
129
+
130
+ The CLI uploads `miner.py`, `weights.onnx`, `class_names.txt`, `chute_config.yml` to your
131
+ HF repo, builds the chute, and commits the on-chain pointer.
132
+
133
+ ## Tuning knobs (top of `miner.py`)
134
+
135
+ | Constant | Default | Effect of raising | Effect of lowering |
136
+ |---|---|---|---|
137
+ | `PER_CLASS_CONF[0]` (balaclava) | 0.05 | fewer FPs (good for FFPI) | more recall (better AP, better IoU) |
138
+ | `PER_CLASS_CONF[2]` (glove) | 0.05 | as above | as above |
139
+ | `PER_CLASS_CONF[4]` (hoodie) | 0.20 | fewer hoodie FPs | more boxes (may hurt precision) |
140
+ | `TTA_SIZES` | (1280, 1536) | better small-object recall | faster inference |
141
+ | `WBF_IOU` | 0.55 | more conservative fusion | tighter clusters |
142
+ | `NMS_IOU` | 0.45 | keeps more near-duplicates | stricter dedup |
143
+ | `MAX_DET` | 100 | more boxes survive ranking | tighter cap |
144
+ | `CLAHE_DARK_THRESHOLD` | 70 | CLAHE on more frames | only the very dark ones |
145
+
146
+ When tuning, validate against `runs/detect/crime_b/val_batch*.jpg` and the manako latest
147
+ challenge image β€” don't hill-climb on the synthetic benchmark alone (it's only 50 frames).
148
+
149
+ ## Why these specific choices
150
+
151
+ - **The IoU pillar dominates the live score** (dashboard 0.576 β‰ˆ baseline `overall_iou`
152
+ 0.597). IoU is the *label-agnostic* AUC-F1 placement metric β€” what matters most is
153
+ whether *any* well-placed box exists for each GT. So the optimal strategy is to flood
154
+ predictions for the rare classes; the FFPI cap (10 FP/image, currently ~6.5 preds/img
155
+ baseline) gives generous headroom.
156
+ - **mAP@50 matters too** because secondary pillars are likely weighted in. mAP@50 is
157
+ per-class-averaged with strict label match. Raising recall on the four near-zero classes
158
+ even modestly (0.03 β†’ 0.20 on balaclava) lifts the per-class mean by ~0.03 alone.
159
+ - **WBF over hard NMS**: tighter localizations β†’ more boxes clearing the IoUβ‰₯0.5 bar.
160
+ - **Class-aware NMS**: balaclava overlaps with hoodie geometry; bat overlaps with glove
161
+ on a held bat. Class-agnostic NMS would silently kill one of each pair.
162
+ - **CLAHE only on dark frames**: applying CLAHE to bright frames hurts hoodie/graffiti
163
+ texture. Luma gate keeps it surgical.
164
+
165
+ ## Verifying you're actually beating the king
166
+
167
+ Before committing on-chain:
168
+
169
+ 1. Pull the latest annotated challenge image+predictions:
170
+ ```bash
171
+ curl -sL "https://console.scorevision.io/api/v2/elements/manak0%2FDetect-crime?lookback_days=7" \
172
+ | jq '.latestAnnotatedChallenge'
173
+ ```
174
+ 2. Run your `miner.py` on that image; visually verify your boxes β‰₯ king's, especially on
175
+ balaclava, glove, and spray paint.
176
+ 3. Run `sv -vv run-once` (per `MINER.md`) to score yourself end-to-end on a real challenge
177
+ without committing β€” confirms the chute deploys correctly and your output format matches.
178
+ 4. Only after the offline score is repeatedly above 0.62 (the king + a comfortable margin)
179
+ should you deploy and commit.
180
+
181
+ ## Open questions / pending work
182
+
183
+ - **Live pillar weights for `Detect-crime`** β€” confirm by reading the active manifest with
184
+ `sv -vv elements list` once `.env` is configured. The recipe above assumes IoU-dominated
185
+ scoring; if mAP/precision/recall pillars are weighted higher, the per-class confidence
186
+ floors should be raised (less recall, more precision).
187
+ - **Real GT vs SAM3 PGT** β€” confirm whether `elements[].ground_truth = true` in the live
188
+ manifest. If real GT (Manako-internal), the synthetic_fixed dataset on HF is the closest
189
+ proxy and we should overfit it carefully. If SAM3 PGT, the live targets are whatever
190
+ SAM3 detects when prompted with the 6 class names β€” slightly fuzzier.
191
+ - **Manako data pull** β€” `poll_manako.py` is built but untested for `Detect-crime`. The
192
+ endpoint shape is the same as petrol-station's; if Manako gates the API for low-traffic
193
+ elements, fall back to using the king's ONNX as the silver labeler over Roboflow data.