File size: 20,695 Bytes
1c0c94d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 | # AVIS β Automated Violation Intelligence System
### Design & Concept Note Β· codename *Gridlock*
> **Thesis.** Most teams will claim they detect all seven violations at high accuracy from a
> single photo. That claim is false, and judges who know vision will see through it. AVIS
> wins by being the system that is *rigorous about evidence*: it detects what a photograph
> can actually prove, attaches an explicit **evidence-sufficiency level** and confidence to
> every finding, uses a vision-language model to verify ambiguous cases and to **abstain**
> when the image doesn't support a verdict, and emits a court-ready, explainable evidence
> package. Honesty + explainability + human-in-the-loop is the differentiator.
This document is the single source of truth for the project's objective and design. It
replaces the earlier draft (`context.md`), which over-engineered the infrastructure and
over-claimed what is detectable from one image.
---
## 1. Problem & what makes it genuinely hard
Traffic cameras produce huge volumes of images; manual review is slow, inconsistent, and
expensive. The task: automatically detect road users, identify and classify violations,
recognise plates, and generate annotated evidence β robust to low light, rain, shadows,
motion blur, density, and image quality.
The non-obvious difficulty β and the thing the original plan glossed over β is that
**violations differ enormously in how provable they are from a single still image:**
- Some are **appearance facts** visible in one frame (no helmet, three people on a bike).
- Some are **spatial facts** that need to know the scene geometry (a vehicle past the
stop line) β solvable from one image *only with per-camera calibration*.
- Some are **temporal/behavioural facts** (running a red light, driving the wrong way,
*parking* = staying put) that fundamentally require motion or duration. Real red-light
cameras use video buffers and multiple angles and *still* misfire on cars that brake
hard at the line. A single photo can produce a *candidate*, never proof.
A credible system must encode this distinction instead of pretending it doesn't exist.
---
## 2. Design philosophy (the objective, reframed)
1. **Evidence sufficiency first.** Every violation type is tagged with what evidence a
single image can provide. The system never auto-confirms a violation the photo can't
prove β it produces a *candidate* and routes it for verification or human review.
2. **Deterministic CV for detection; AI for judgement.** Fast, reliable open models do the
detecting and measuring. A vision-language model (VLM) is used **only** to verify
ambiguous candidates, to abstain when evidence is weak, and to write the human-readable
justification. This keeps the pipeline cheap, fast, and explainable.
3. **Explainable & accountable by construction.** Every output carries: the reason, the
confidence, the evidence-sufficiency level, the legal reference, and a tamper-evident
hash. No black-box verdicts; a human can always be put in the loop.
4. **Right-sized engineering.** A clean modular monolith that runs on a laptop/CPU and free
APIs β with a clearly documented path to horizontal scale. We spend our time on the CV
and the demo, not on orchestration plumbing.
5. **Free and self-hostable.** Open-source models + one free API tier (Gemini). No billing.
---
## 3. The Violation Detectability Matrix *(the core intellectual contribution)*
Each violation is assigned a **tier** that determines its evidence-sufficiency level and
how the pipeline routes it. This table drives the Rule Engine and the routing policy.
| Violation | Tier | What a single image proves | Inputs needed | Default route |
|---|---|---|---|---|
| **Helmet non-compliance** | **A β Appearance** | Reliable: rider head visible, helmet/no-helmet | Detector + helmet classifier | Auto-confirm if high conf, else VLM |
| **Triple riding** | **A β Appearance** | Reliable: count riders linked to one motorcycle | Detector + riderβbike association | Auto-confirm if clean count, else VLM |
| **Seatbelt non-compliance** | **B β Hard appearance** | Weak: windshield glare/resolution make this unreliable | Windshield crop classifier *or* VLM | VLM-verify or human; mark low-confidence |
| **Stop-line violation** | **C β Spatial (needs calibration)** | Strong *candidate*: vehicle body past the stop-line zone | Detector + per-camera stop-line polygon | VLM-verify; human if no calibration |
| **Red-light violation** | **C/D β Spatial + temporal** | Candidate only: light=red (from image) AND vehicle past stop line; true running is temporal | Detector + light-state + stop-line polygon | Always VLM-verify + flag "confirm with sequence" |
| **Illegal parking** | **C/D β Spatial + temporal** | Candidate only: vehicle inside no-parking zone; "parked" = duration, unprovable from one frame | Detector + no-parking polygon | Candidate; flag "needs dwell-time confirmation" |
| **Wrong-side driving** | **D β Temporal** | Not provable: direction of travel needs motion; orientation-vs-lane is a weak proxy | Detector + lane direction (+ orientation) | Low-confidence candidate; recommend video |
| **License-plate recognition** | *(supporting)* | Reliable when plate is legible | Plate detector + OCR + regex | Always attempted; attach to any violation |
**Consequence for the build:** Tier A is our headline, demo-grade capability. Tier B is
best-effort with honest confidence. Tiers C/D are positioned as **candidate generation +
calibration-assisted evidence**, never as "we solved red-light running from a JPEG." This
framing is defensible and impressive; over-claiming is not.
---
## 4. System overview
A **production-credible modular monolith**: one FastAPI codebase running a linear pipeline of
pure-ish stages, with the heavy CV work handed to a **worker via a Redis queue** so the API
stays responsive. This is production-grade *and* demo-safe β real horizontal scalability (add
workers) without splitting into many microservices. Everything is free/self-hosted and comes
up with one `docker compose up`. Storage and the queue sit behind interfaces, so a developer
can run a lightweight local mode (in-process tasks + filesystem) when they don't want to
start the full stack.
```
upload β FastAPI api ββenqueueβββΆ Redis βββΆ Worker(s): pipeline stages 1β10
β
1 Ingest+QualityGate β 2 Preprocess β 3 Detect β 4 SceneGraph
β 5 Attributes(helmet/seatbelt/plate) β 6 RuleEngine β 7 Fuse+Route
β 8 VLM Verify (Gemini, only when routed) β 9 Legal map β 10 Evidence Composer
β
ββββββββββββββββββββββββββββββββββββββββ΄ββββββββββββββββββββββββββββββ
Postgres (JSONB: metadata + evidence graph) MinIO (S3: images)
ββββββββββββββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββ
βΌ
React + Vite dashboard
(upload Β· review queue Β· analytics Β· search)
scale: add workers Β· swap MinIOβS3, Postgresβmanaged Β· GPU detection workers
```
### Stage responsibilities
1. **Ingest + Quality Gate** β store the image + metadata (camera_id, timestamp, geo if
given); run a blur/exposure/resolution check. If the image is unusable, mark
`undeterminable` and stop β *abstaining is a feature, not a failure.*
2. **Preprocess** β OpenCV: CLAHE (contrast), denoise, optional low-light enhancement
(gamma / Retinex-style), light deblur. Produces a normalised image; keeps the original
untouched for evidence.
3. **Detect** β YOLO11/YOLOv8 (COCO) for car/truck/bus/motorcycle/bicycle/person/traffic
light; a fine-tuned model adds helmet and plate classes. Output: boxes + class + conf.
4. **Scene Graph** β associate people to vehicles by geometry (riderβmotorcycle,
driverβcar), attach traffic lights and plates. This graph is the single source of truth
for all downstream reasoning (Β§5).
5. **Attribute classifiers** β helmet per rider (YOLO/classifier); seatbelt per driver
(windshield crop classifier or VLM, Tier B); traffic-light state (HSV + classifier);
plate text (fast-alpr OCR + Indian-plate regex correction).
6. **Rule Engine** β pure, deterministic functions, one per violation, consuming the graph
+ optional per-camera calibration. Emits **candidates** with a rule-evidence score and
the tier from Β§3. No ML, no I/O β fully unit-testable.
7. **Fuse + Route** β combine detection / attribute / rule scores into a fused confidence;
apply tier-aware routing (Β§7): auto-confirm, VLM-verify, human-review, or abstain.
8. **VLM Verify** β only for routed candidates: send the cropped evidence + a strict prompt
to **Gemini Flash (free tier)**; get `{verified, confidence, reason}` or
`insufficient_evidence`. Sparingly used to respect the 1,500/day, ~10/min free quota.
9. **Legal map** β deterministic lookup table: violation β Motor Vehicles Act section β
fine. (Correct, instant, free. Optional RAG is a bonus feature, not load-bearing β Β§10.)
10. **Evidence Composer** β annotated image (boxes, labels, plate, reason) + structured
JSON + SHA-256 hash + audit trail; persist to Postgres (JSONB) + MinIO.
---
## 5. The Evidence Graph (data model)
Detections become a typed graph, not a flat list. This is what makes the system queryable,
explainable, and auditable.
```
Nodes:
Vehicle { id, type, bbox, conf, attrs:{orientation?, in_zones:[...]} }
Person { id, role: rider|driver|pedestrian, bbox, conf, attrs:{helmet?, seatbelt?} }
Light { id, state: red|amber|green|unknown, bbox, conf }
Plate { id, text, regex_ok, conf, bbox }
Zone { id, kind: stop_line|no_parking|lane, polygon } # from camera calibration
Edges:
rides(Person β Vehicle:motorcycle) drives(Person β Vehicle:car)
has_plate(Vehicle β Plate) located_in(Vehicle β Zone)
governed_by(Vehicle β Light)
```
A violation is then a small, explainable subgraph, e.g. *triple riding* =
`motorcycle_1` with three `rides` edges. "Why?" is answerable directly from the graph
(*"3 riders linked to 1 motorcycle"*). Example output payload:
```json
{
"violation": "TRIPLE_RIDING",
"tier": "A",
"evidence_sufficiency": "sufficient",
"subjects": ["motorcycle_1", "rider_1", "rider_2", "rider_3"],
"scores": { "detection": 0.94, "rule": 0.90, "vlm": 0.93, "fused": 0.92 },
"route": "auto_confirmed",
"reason": "Three distinct riders are linked to a single motorcycle.",
"plate": { "text": "UP32AB1234", "regex_ok": true, "conf": 0.88 },
"legal": { "act": "Motor Vehicles Act, 1988", "section": "128", "fine": "βΉ1000" },
"evidence_hash": "sha256:β¦",
"timestamp": "2026-06-20T10:31:00Z"
}
```
---
## 6. Confidence fusion & routing (with abstention)
Per candidate, fuse available signals with tier-specific weights:
```
fused = w_detΒ·detection_conf + w_attrΒ·attribute_conf + w_ruleΒ·rule_evidence (+ w_vlmΒ·vlm_conf)
```
Routing policy (thresholds are config, not hardcoded):
- **Tier A:** `fused β₯ 0.85` and unambiguous β **auto-confirm**; `0.55β0.85` β **VLM-verify**;
`< 0.55` β **human review**.
- **Tier B (seatbelt):** never auto-confirm β **VLM-verify**, then human if VLM is unsure.
- **Tier C (stop-line / red-light / parking):** never auto-confirm. With calibration β
**VLM-verify** and label `candidate`; without calibration β **human review**.
- **Tier D (wrong-side):** emit **low-confidence candidate** + "recommend video
confirmation"; never auto-confirm.
- **Quality gate failed or VLM returns `insufficient_evidence`** β **abstain**
(`undeterminable`), not a false positive.
This is what makes the free Gemini quota workable: only the genuinely ambiguous minority
hits the VLM, and clear Tier-A cases auto-confirm for free.
---
## 7. VLM verification layer (Gemini free tier)
- **Role:** auditor, never detector. It receives a *focused crop* + the proposed violation
+ the graph evidence, and answers a strict JSON contract:
```
{"verified": bool, "confidence": 0.0-1.0, "reason": "<one sentence>",
"insufficient_evidence": bool}
```
- **Model:** Gemini Flash (free tier) via a provider-agnostic client so it can swap to
Groq / OpenRouter free models without touching call sites.
- **Budget discipline:** free tier = ~1,500 req/day, ~10/min. Therefore: route sparingly,
batch where possible, **cache by `evidence_hash`** (identical crops never re-billed),
and degrade gracefully to "human review" if the quota is exhausted. For the live demo,
pre-cache responses for the demo images.
---
## 8. Legal grounding
For seven fixed violation types, the violationβsectionβfine mapping is a **small static
table** β correct, instant, and free. That table is the source of truth.
*Optional wow-factor (only if time allows):* a tiny RAG layer (ChromaDB +
sentence-transformers over the Motor Vehicles Act text) that lets an officer **ask
free-form questions** ("what's the penalty for X, and the appeal process?"). This is a
bonus feature layered on top β it must never be on the critical path of issuing a verdict.
---
## 9. Evidence package (court-ready output)
Each confirmed/escalated violation produces:
- **Annotated image** β original (untouched) + an overlay copy with boxes, labels,
plate, and the natural-language reason.
- **Structured JSON** β the payload in Β§5, including scores, tier, route, legal ref.
- **Integrity** β SHA-256 hash of the original image + an append-only audit trail
(who/what changed the verdict, when). This is the "admissible evidence" story.
- **e-challan-ready** β plate + violation + section + fine in one record, exportable.
---
## 10. Robustness to varying conditions (explicitly required)
- **Preprocessing** handles low light (CLAHE / gamma / Retinex), rain & noise (denoise),
motion blur (mild deblur / sharpening).
- **Quality gate** quantifies blur/exposure and *abstains* on images too poor to judge β
preventing confident-but-wrong outputs, which is the real failure mode under bad
conditions.
- **VLM second opinion** adds robustness on ambiguous/degraded crops and can abstain.
- **Honest confidence** means degraded inputs surface as lower confidence β review, not as
silent false positives.
---
## 11. Why this wins (innovation summary)
1. **Evidence-sufficiencyβaware detection** β explicit tiers + abstention. Rigorous and
rare; directly answers "robust and accurate" without over-claiming.
2. **Evidence Graph** β relationship-first, explainable single source of truth.
3. **VLM-as-verifier-and-abstainer** β cuts false positives, writes human-readable reasons,
stays inside a free quota by only judging the hard cases.
4. **Confidence routing with human-in-the-loop** β quantifiable "human-review reduction %".
5. **Court-ready evidence** β legal grounding + tamper-evident hash + audit trail.
6. **Calibration-optional** β works on arbitrary images for Tier A; uses optional per-camera
zones for Tier C; clearly degrades elsewhere.
---
## 12. Evaluation plan (mapped to the required metrics)
- **Component metrics (P / R / F1 / mAP):** on public datasets β helmet & triple-riding
(Roboflow), plate detection + OCR character/whole-plate accuracy. Report per class.
- **End-to-end:** a hand-curated ~60β100 image set spanning day/night/rain/density, with a
per-violation confusion matrix and false-positive rate **per image** (the metric that
actually governs reviewer workload).
- **Operational:** mean latency/image on CPU, throughput (images/min), and the
**auto-confirm vs. VLM vs. human** split β human-review-reduction %.
- **Ablation (great for judges):** rule-only vs. rule+VLM, to quantify the false-positive
reduction the VLM buys.
- **Honesty:** report Tier A separately from C/D; never average them into one inflated
"accuracy" number.
---
## 13. Tech stack (free) & scalability path
**MVP (build this) β a production-credible monolith, all free/self-hosted:**
- Python 3.11 Β· FastAPI Β· Uvicorn Β· **Redis-backed worker queue** (Celery or RQ/ARQ) so the
API stays responsive and scales by adding workers; in-process fallback for quick local dev.
- **Detection:** Ultralytics **YOLO11/YOLOv8** (pip, pretrained COCO + fine-tuned
helmet/plate from Roboflow). *(Open-vocab Grounding DINO / YOLO-World is an optional
enhancement, not a dependency β it's slow and painful to install.)*
- **Plates:** **fast-alpr** (detection + `fast-plate-ocr`, ONNX, CPU-fast) + Indian-plate
regex correction.
- **Seatbelt/ambiguous:** Gemini VLM (Tier B/C).
- **VLM:** **Gemini Flash free tier** via provider-agnostic `core/llm/` client β a one-line
swap to a paid model (Claude / GPT-4o) later, with zero code rework. Free is a demo-time
choice, not a capability ceiling.
- **Storage:** **Postgres (JSONB)** via SQLModel for metadata *and* the evidence graph
(relational integrity + native JSON querying); **MinIO** (free, S3-compatible) for images.
Both behind a `core/storage/` abstraction.
- **Legal:** static lookup table (+ optional ChromaDB RAG bonus).
- **Frontend:** React + Vite + Chart.js.
- **Packaging:** one `docker compose up` brings up api + worker + Postgres + Redis + MinIO.
**Why not full microservices?** This monolith already delivers the real scalability win β a
true task queue and horizontally-scalable workers β in one maintainable codebase. Splitting
the stages into separately-deployed services adds ops burden and demo risk for little extra
scalability during a hackathon.
**Scale-out path (a deployment change, not a rewrite):** add more workers; swap MinIO β AWS
S3 and Postgres β a managed instance; run the detection stage as dedicated GPU workers; and,
if ever truly needed, peel stages into separate services behind the same `storage/` and
`llm/` interfaces. Because everything already talks through those interfaces, none of this
touches business logic.
---
## 14. Pitch & demo flow (3β4 min)
1. **Hook (20s):** "Every team says they catch all 7 violations from a photo. They can't β
and neither can a real red-light camera from one frame. We built the system that knows
the difference."
2. **Idea (40s):** hybrid CV + VLM, Evidence Graph, evidence-sufficiency tiers, routing.
3. **Live demo (90s):** upload a triple-riding image β annotated result, graph panel
("motorcycle β 3 riders, 1 helmet"), VLM reason, plate OCR, legal section + fine; then a
degraded night image β system **abstains / routes to review** (show this on purpose);
then the analytics + review-queue dashboard.
4. **Proof (30s):** metrics table (Tier A P/R/F1, ablation showing VLM cuts false positives,
human-review-reduction %).
5. **Close (20s):** free/self-hostable, calibration-optional, scale-out path, e-challan-ready.
---
## 15. Risks & honest limitations
| Risk | Mitigation |
|---|---|
| Seatbelt detection unreliable | Tier B: best-effort + VLM/human; report honest confidence |
| Tier C/D need calibration we won't have for random images | Position as candidate generation; demo with one calibrated camera config |
| Free Gemini quota exhausted | Cache by hash, route sparingly, pre-cache demo, degrade to human review |
| Grounding DINO install eats hackathon time | Default to Ultralytics YOLO; open-vocab is optional |
| Public model accuracy varies on Indian scenes | Pick India-trained Roboflow models; report measured, not claimed, numbers |
---
## References
- Gemini API free tier limits (2026): https://tokenmix.ai/blog/gemini-api-free-tier-limits
- fast-alpr / fast-plate-ocr: https://github.com/ankandrew/fast-alpr
- Triple-riding + helmet via YOLOv8 (CCTV): https://www.atlantis-press.com/proceedings/computatia-25/126010076
- Helmet violation detection, Indian smart-city (YOLOv8/TAO): https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2025.1582257/full
- Red-light running needs video (false positives from braking): https://www.researchgate.net/publication/3154744_An_effective_video_analysis_method_for_detecting_red_light_runners
- VLM on complex traffic events (GPT-4V study): https://arxiv.org/pdf/2402.02205
|