File size: 20,695 Bytes
1c0c94d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
# AVIS β€” Automated Violation Intelligence System
### Design & Concept Note Β· codename *Gridlock*

> **Thesis.** Most teams will claim they detect all seven violations at high accuracy from a
> single photo. That claim is false, and judges who know vision will see through it. AVIS
> wins by being the system that is *rigorous about evidence*: it detects what a photograph
> can actually prove, attaches an explicit **evidence-sufficiency level** and confidence to
> every finding, uses a vision-language model to verify ambiguous cases and to **abstain**
> when the image doesn't support a verdict, and emits a court-ready, explainable evidence
> package. Honesty + explainability + human-in-the-loop is the differentiator.

This document is the single source of truth for the project's objective and design. It
replaces the earlier draft (`context.md`), which over-engineered the infrastructure and
over-claimed what is detectable from one image.

---

## 1. Problem & what makes it genuinely hard

Traffic cameras produce huge volumes of images; manual review is slow, inconsistent, and
expensive. The task: automatically detect road users, identify and classify violations,
recognise plates, and generate annotated evidence β€” robust to low light, rain, shadows,
motion blur, density, and image quality.

The non-obvious difficulty β€” and the thing the original plan glossed over β€” is that
**violations differ enormously in how provable they are from a single still image:**

- Some are **appearance facts** visible in one frame (no helmet, three people on a bike).
- Some are **spatial facts** that need to know the scene geometry (a vehicle past the
  stop line) β€” solvable from one image *only with per-camera calibration*.
- Some are **temporal/behavioural facts** (running a red light, driving the wrong way,
  *parking* = staying put) that fundamentally require motion or duration. Real red-light
  cameras use video buffers and multiple angles and *still* misfire on cars that brake
  hard at the line. A single photo can produce a *candidate*, never proof.

A credible system must encode this distinction instead of pretending it doesn't exist.

---

## 2. Design philosophy (the objective, reframed)

1. **Evidence sufficiency first.** Every violation type is tagged with what evidence a
   single image can provide. The system never auto-confirms a violation the photo can't
   prove β€” it produces a *candidate* and routes it for verification or human review.
2. **Deterministic CV for detection; AI for judgement.** Fast, reliable open models do the
   detecting and measuring. A vision-language model (VLM) is used **only** to verify
   ambiguous candidates, to abstain when evidence is weak, and to write the human-readable
   justification. This keeps the pipeline cheap, fast, and explainable.
3. **Explainable & accountable by construction.** Every output carries: the reason, the
   confidence, the evidence-sufficiency level, the legal reference, and a tamper-evident
   hash. No black-box verdicts; a human can always be put in the loop.
4. **Right-sized engineering.** A clean modular monolith that runs on a laptop/CPU and free
   APIs β€” with a clearly documented path to horizontal scale. We spend our time on the CV
   and the demo, not on orchestration plumbing.
5. **Free and self-hostable.** Open-source models + one free API tier (Gemini). No billing.

---

## 3. The Violation Detectability Matrix  *(the core intellectual contribution)*

Each violation is assigned a **tier** that determines its evidence-sufficiency level and
how the pipeline routes it. This table drives the Rule Engine and the routing policy.

| Violation | Tier | What a single image proves | Inputs needed | Default route |
|---|---|---|---|---|
| **Helmet non-compliance** | **A β€” Appearance** | Reliable: rider head visible, helmet/no-helmet | Detector + helmet classifier | Auto-confirm if high conf, else VLM |
| **Triple riding** | **A β€” Appearance** | Reliable: count riders linked to one motorcycle | Detector + rider↔bike association | Auto-confirm if clean count, else VLM |
| **Seatbelt non-compliance** | **B β€” Hard appearance** | Weak: windshield glare/resolution make this unreliable | Windshield crop classifier *or* VLM | VLM-verify or human; mark low-confidence |
| **Stop-line violation** | **C β€” Spatial (needs calibration)** | Strong *candidate*: vehicle body past the stop-line zone | Detector + per-camera stop-line polygon | VLM-verify; human if no calibration |
| **Red-light violation** | **C/D β€” Spatial + temporal** | Candidate only: light=red (from image) AND vehicle past stop line; true running is temporal | Detector + light-state + stop-line polygon | Always VLM-verify + flag "confirm with sequence" |
| **Illegal parking** | **C/D β€” Spatial + temporal** | Candidate only: vehicle inside no-parking zone; "parked" = duration, unprovable from one frame | Detector + no-parking polygon | Candidate; flag "needs dwell-time confirmation" |
| **Wrong-side driving** | **D β€” Temporal** | Not provable: direction of travel needs motion; orientation-vs-lane is a weak proxy | Detector + lane direction (+ orientation) | Low-confidence candidate; recommend video |
| **License-plate recognition** | *(supporting)* | Reliable when plate is legible | Plate detector + OCR + regex | Always attempted; attach to any violation |

**Consequence for the build:** Tier A is our headline, demo-grade capability. Tier B is
best-effort with honest confidence. Tiers C/D are positioned as **candidate generation +
calibration-assisted evidence**, never as "we solved red-light running from a JPEG." This
framing is defensible and impressive; over-claiming is not.

---

## 4. System overview

A **production-credible modular monolith**: one FastAPI codebase running a linear pipeline of
pure-ish stages, with the heavy CV work handed to a **worker via a Redis queue** so the API
stays responsive. This is production-grade *and* demo-safe β€” real horizontal scalability (add
workers) without splitting into many microservices. Everything is free/self-hosted and comes
up with one `docker compose up`. Storage and the queue sit behind interfaces, so a developer
can run a lightweight local mode (in-process tasks + filesystem) when they don't want to
start the full stack.

```
  upload β†’ FastAPI api ──enqueue──▢ Redis ──▢ Worker(s): pipeline stages 1–10
                                                  β”‚
   1 Ingest+QualityGate β†’ 2 Preprocess β†’ 3 Detect β†’ 4 SceneGraph
   β†’ 5 Attributes(helmet/seatbelt/plate) β†’ 6 RuleEngine β†’ 7 Fuse+Route
   β†’ 8 VLM Verify (Gemini, only when routed) β†’ 9 Legal map β†’ 10 Evidence Composer
                                                  β”‚
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  Postgres (JSONB: metadata + evidence graph)                  MinIO (S3: images)
        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                                β–Ό
                                  React + Vite dashboard
                        (upload Β· review queue Β· analytics Β· search)
   scale: add workers · swap MinIO→S3, Postgres→managed · GPU detection workers
```

### Stage responsibilities
1. **Ingest + Quality Gate** β€” store the image + metadata (camera_id, timestamp, geo if
   given); run a blur/exposure/resolution check. If the image is unusable, mark
   `undeterminable` and stop β€” *abstaining is a feature, not a failure.*
2. **Preprocess** β€” OpenCV: CLAHE (contrast), denoise, optional low-light enhancement
   (gamma / Retinex-style), light deblur. Produces a normalised image; keeps the original
   untouched for evidence.
3. **Detect** β€” YOLO11/YOLOv8 (COCO) for car/truck/bus/motorcycle/bicycle/person/traffic
   light; a fine-tuned model adds helmet and plate classes. Output: boxes + class + conf.
4. **Scene Graph** β€” associate people to vehicles by geometry (rider↔motorcycle,
   driver↔car), attach traffic lights and plates. This graph is the single source of truth
   for all downstream reasoning (Β§5).
5. **Attribute classifiers** β€” helmet per rider (YOLO/classifier); seatbelt per driver
   (windshield crop classifier or VLM, Tier B); traffic-light state (HSV + classifier);
   plate text (fast-alpr OCR + Indian-plate regex correction).
6. **Rule Engine** β€” pure, deterministic functions, one per violation, consuming the graph
   + optional per-camera calibration. Emits **candidates** with a rule-evidence score and
   the tier from Β§3. No ML, no I/O β€” fully unit-testable.
7. **Fuse + Route** β€” combine detection / attribute / rule scores into a fused confidence;
   apply tier-aware routing (Β§7): auto-confirm, VLM-verify, human-review, or abstain.
8. **VLM Verify** β€” only for routed candidates: send the cropped evidence + a strict prompt
   to **Gemini Flash (free tier)**; get `{verified, confidence, reason}` or
   `insufficient_evidence`. Sparingly used to respect the 1,500/day, ~10/min free quota.
9. **Legal map** β€” deterministic lookup table: violation β†’ Motor Vehicles Act section β†’
   fine. (Correct, instant, free. Optional RAG is a bonus feature, not load-bearing β€” Β§10.)
10. **Evidence Composer** β€” annotated image (boxes, labels, plate, reason) + structured
    JSON + SHA-256 hash + audit trail; persist to Postgres (JSONB) + MinIO.

---

## 5. The Evidence Graph (data model)

Detections become a typed graph, not a flat list. This is what makes the system queryable,
explainable, and auditable.

```
Nodes:
  Vehicle  { id, type, bbox, conf, attrs:{orientation?, in_zones:[...]} }
  Person   { id, role: rider|driver|pedestrian, bbox, conf, attrs:{helmet?, seatbelt?} }
  Light    { id, state: red|amber|green|unknown, bbox, conf }
  Plate    { id, text, regex_ok, conf, bbox }
  Zone     { id, kind: stop_line|no_parking|lane, polygon }      # from camera calibration

Edges:
  rides(Person β†’ Vehicle:motorcycle)      drives(Person β†’ Vehicle:car)
  has_plate(Vehicle β†’ Plate)              located_in(Vehicle β†’ Zone)
  governed_by(Vehicle β†’ Light)
```

A violation is then a small, explainable subgraph, e.g. *triple riding* =
`motorcycle_1` with three `rides` edges. "Why?" is answerable directly from the graph
(*"3 riders linked to 1 motorcycle"*). Example output payload:

```json
{
  "violation": "TRIPLE_RIDING",
  "tier": "A",
  "evidence_sufficiency": "sufficient",
  "subjects": ["motorcycle_1", "rider_1", "rider_2", "rider_3"],
  "scores": { "detection": 0.94, "rule": 0.90, "vlm": 0.93, "fused": 0.92 },
  "route": "auto_confirmed",
  "reason": "Three distinct riders are linked to a single motorcycle.",
  "plate": { "text": "UP32AB1234", "regex_ok": true, "conf": 0.88 },
  "legal": { "act": "Motor Vehicles Act, 1988", "section": "128", "fine": "β‚Ή1000" },
  "evidence_hash": "sha256:…",
  "timestamp": "2026-06-20T10:31:00Z"
}
```

---

## 6. Confidence fusion & routing (with abstention)

Per candidate, fuse available signals with tier-specific weights:

```
fused = w_detΒ·detection_conf + w_attrΒ·attribute_conf + w_ruleΒ·rule_evidence (+ w_vlmΒ·vlm_conf)
```

Routing policy (thresholds are config, not hardcoded):

- **Tier A:** `fused β‰₯ 0.85` and unambiguous β†’ **auto-confirm**; `0.55–0.85` β†’ **VLM-verify**;
  `< 0.55` β†’ **human review**.
- **Tier B (seatbelt):** never auto-confirm β†’ **VLM-verify**, then human if VLM is unsure.
- **Tier C (stop-line / red-light / parking):** never auto-confirm. With calibration β†’
  **VLM-verify** and label `candidate`; without calibration β†’ **human review**.
- **Tier D (wrong-side):** emit **low-confidence candidate** + "recommend video
  confirmation"; never auto-confirm.
- **Quality gate failed or VLM returns `insufficient_evidence`** β†’ **abstain**
  (`undeterminable`), not a false positive.

This is what makes the free Gemini quota workable: only the genuinely ambiguous minority
hits the VLM, and clear Tier-A cases auto-confirm for free.

---

## 7. VLM verification layer (Gemini free tier)

- **Role:** auditor, never detector. It receives a *focused crop* + the proposed violation
  + the graph evidence, and answers a strict JSON contract:
  ```
  {"verified": bool, "confidence": 0.0-1.0, "reason": "<one sentence>",
   "insufficient_evidence": bool}
  ```
- **Model:** Gemini Flash (free tier) via a provider-agnostic client so it can swap to
  Groq / OpenRouter free models without touching call sites.
- **Budget discipline:** free tier = ~1,500 req/day, ~10/min. Therefore: route sparingly,
  batch where possible, **cache by `evidence_hash`** (identical crops never re-billed),
  and degrade gracefully to "human review" if the quota is exhausted. For the live demo,
  pre-cache responses for the demo images.

---

## 8. Legal grounding

For seven fixed violation types, the violation→section→fine mapping is a **small static
table** β€” correct, instant, and free. That table is the source of truth.

*Optional wow-factor (only if time allows):* a tiny RAG layer (ChromaDB +
sentence-transformers over the Motor Vehicles Act text) that lets an officer **ask
free-form questions** ("what's the penalty for X, and the appeal process?"). This is a
bonus feature layered on top β€” it must never be on the critical path of issuing a verdict.

---

## 9. Evidence package (court-ready output)

Each confirmed/escalated violation produces:
- **Annotated image** β€” original (untouched) + an overlay copy with boxes, labels,
  plate, and the natural-language reason.
- **Structured JSON** β€” the payload in Β§5, including scores, tier, route, legal ref.
- **Integrity** β€” SHA-256 hash of the original image + an append-only audit trail
  (who/what changed the verdict, when). This is the "admissible evidence" story.
- **e-challan-ready** β€” plate + violation + section + fine in one record, exportable.

---

## 10. Robustness to varying conditions (explicitly required)

- **Preprocessing** handles low light (CLAHE / gamma / Retinex), rain & noise (denoise),
  motion blur (mild deblur / sharpening).
- **Quality gate** quantifies blur/exposure and *abstains* on images too poor to judge β€”
  preventing confident-but-wrong outputs, which is the real failure mode under bad
  conditions.
- **VLM second opinion** adds robustness on ambiguous/degraded crops and can abstain.
- **Honest confidence** means degraded inputs surface as lower confidence β†’ review, not as
  silent false positives.

---

## 11. Why this wins (innovation summary)

1. **Evidence-sufficiency–aware detection** β€” explicit tiers + abstention. Rigorous and
   rare; directly answers "robust and accurate" without over-claiming.
2. **Evidence Graph** β€” relationship-first, explainable single source of truth.
3. **VLM-as-verifier-and-abstainer** β€” cuts false positives, writes human-readable reasons,
   stays inside a free quota by only judging the hard cases.
4. **Confidence routing with human-in-the-loop** β€” quantifiable "human-review reduction %".
5. **Court-ready evidence** β€” legal grounding + tamper-evident hash + audit trail.
6. **Calibration-optional** β€” works on arbitrary images for Tier A; uses optional per-camera
   zones for Tier C; clearly degrades elsewhere.

---

## 12. Evaluation plan (mapped to the required metrics)

- **Component metrics (P / R / F1 / mAP):** on public datasets β€” helmet & triple-riding
  (Roboflow), plate detection + OCR character/whole-plate accuracy. Report per class.
- **End-to-end:** a hand-curated ~60–100 image set spanning day/night/rain/density, with a
  per-violation confusion matrix and false-positive rate **per image** (the metric that
  actually governs reviewer workload).
- **Operational:** mean latency/image on CPU, throughput (images/min), and the
  **auto-confirm vs. VLM vs. human** split β†’ human-review-reduction %.
- **Ablation (great for judges):** rule-only vs. rule+VLM, to quantify the false-positive
  reduction the VLM buys.
- **Honesty:** report Tier A separately from C/D; never average them into one inflated
  "accuracy" number.

---

## 13. Tech stack (free) & scalability path

**MVP (build this) β€” a production-credible monolith, all free/self-hosted:**
- Python 3.11 Β· FastAPI Β· Uvicorn Β· **Redis-backed worker queue** (Celery or RQ/ARQ) so the
  API stays responsive and scales by adding workers; in-process fallback for quick local dev.
- **Detection:** Ultralytics **YOLO11/YOLOv8** (pip, pretrained COCO + fine-tuned
  helmet/plate from Roboflow). *(Open-vocab Grounding DINO / YOLO-World is an optional
  enhancement, not a dependency β€” it's slow and painful to install.)*
- **Plates:** **fast-alpr** (detection + `fast-plate-ocr`, ONNX, CPU-fast) + Indian-plate
  regex correction.
- **Seatbelt/ambiguous:** Gemini VLM (Tier B/C).
- **VLM:** **Gemini Flash free tier** via provider-agnostic `core/llm/` client β€” a one-line
  swap to a paid model (Claude / GPT-4o) later, with zero code rework. Free is a demo-time
  choice, not a capability ceiling.
- **Storage:** **Postgres (JSONB)** via SQLModel for metadata *and* the evidence graph
  (relational integrity + native JSON querying); **MinIO** (free, S3-compatible) for images.
  Both behind a `core/storage/` abstraction.
- **Legal:** static lookup table (+ optional ChromaDB RAG bonus).
- **Frontend:** React + Vite + Chart.js.
- **Packaging:** one `docker compose up` brings up api + worker + Postgres + Redis + MinIO.

**Why not full microservices?** This monolith already delivers the real scalability win β€” a
true task queue and horizontally-scalable workers β€” in one maintainable codebase. Splitting
the stages into separately-deployed services adds ops burden and demo risk for little extra
scalability during a hackathon.

**Scale-out path (a deployment change, not a rewrite):** add more workers; swap MinIO β†’ AWS
S3 and Postgres β†’ a managed instance; run the detection stage as dedicated GPU workers; and,
if ever truly needed, peel stages into separate services behind the same `storage/` and
`llm/` interfaces. Because everything already talks through those interfaces, none of this
touches business logic.

---

## 14. Pitch & demo flow (3–4 min)

1. **Hook (20s):** "Every team says they catch all 7 violations from a photo. They can't β€”
   and neither can a real red-light camera from one frame. We built the system that knows
   the difference."
2. **Idea (40s):** hybrid CV + VLM, Evidence Graph, evidence-sufficiency tiers, routing.
3. **Live demo (90s):** upload a triple-riding image β†’ annotated result, graph panel
   ("motorcycle β†’ 3 riders, 1 helmet"), VLM reason, plate OCR, legal section + fine; then a
   degraded night image β†’ system **abstains / routes to review** (show this on purpose);
   then the analytics + review-queue dashboard.
4. **Proof (30s):** metrics table (Tier A P/R/F1, ablation showing VLM cuts false positives,
   human-review-reduction %).
5. **Close (20s):** free/self-hostable, calibration-optional, scale-out path, e-challan-ready.

---

## 15. Risks & honest limitations

| Risk | Mitigation |
|---|---|
| Seatbelt detection unreliable | Tier B: best-effort + VLM/human; report honest confidence |
| Tier C/D need calibration we won't have for random images | Position as candidate generation; demo with one calibrated camera config |
| Free Gemini quota exhausted | Cache by hash, route sparingly, pre-cache demo, degrade to human review |
| Grounding DINO install eats hackathon time | Default to Ultralytics YOLO; open-vocab is optional |
| Public model accuracy varies on Indian scenes | Pick India-trained Roboflow models; report measured, not claimed, numbers |

---

## References
- Gemini API free tier limits (2026): https://tokenmix.ai/blog/gemini-api-free-tier-limits
- fast-alpr / fast-plate-ocr: https://github.com/ankandrew/fast-alpr
- Triple-riding + helmet via YOLOv8 (CCTV): https://www.atlantis-press.com/proceedings/computatia-25/126010076
- Helmet violation detection, Indian smart-city (YOLOv8/TAO): https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2025.1582257/full
- Red-light running needs video (false positives from braking): https://www.researchgate.net/publication/3154744_An_effective_video_analysis_method_for_detecting_red_light_runners
- VLM on complex traffic events (GPT-4V study): https://arxiv.org/pdf/2402.02205