| # AVIS β Automated Violation Intelligence System |
| ### Design & Concept Note Β· codename *Gridlock* |
|
|
| > **Thesis.** Most teams will claim they detect all seven violations at high accuracy from a |
| > single photo. That claim is false, and judges who know vision will see through it. AVIS |
| > wins by being the system that is *rigorous about evidence*: it detects what a photograph |
| > can actually prove, attaches an explicit **evidence-sufficiency level** and confidence to |
| > every finding, uses a vision-language model to verify ambiguous cases and to **abstain** |
| > when the image doesn't support a verdict, and emits a court-ready, explainable evidence |
| > package. Honesty + explainability + human-in-the-loop is the differentiator. |
|
|
| This document is the single source of truth for the project's objective and design. It |
| replaces the earlier draft (`context.md`), which over-engineered the infrastructure and |
| over-claimed what is detectable from one image. |
|
|
| --- |
|
|
| ## 1. Problem & what makes it genuinely hard |
|
|
| Traffic cameras produce huge volumes of images; manual review is slow, inconsistent, and |
| expensive. The task: automatically detect road users, identify and classify violations, |
| recognise plates, and generate annotated evidence β robust to low light, rain, shadows, |
| motion blur, density, and image quality. |
|
|
| The non-obvious difficulty β and the thing the original plan glossed over β is that |
| **violations differ enormously in how provable they are from a single still image:** |
|
|
| - Some are **appearance facts** visible in one frame (no helmet, three people on a bike). |
| - Some are **spatial facts** that need to know the scene geometry (a vehicle past the |
| stop line) β solvable from one image *only with per-camera calibration*. |
| - Some are **temporal/behavioural facts** (running a red light, driving the wrong way, |
| *parking* = staying put) that fundamentally require motion or duration. Real red-light |
| cameras use video buffers and multiple angles and *still* misfire on cars that brake |
| hard at the line. A single photo can produce a *candidate*, never proof. |
|
|
| A credible system must encode this distinction instead of pretending it doesn't exist. |
|
|
| --- |
|
|
| ## 2. Design philosophy (the objective, reframed) |
|
|
| 1. **Evidence sufficiency first.** Every violation type is tagged with what evidence a |
| single image can provide. The system never auto-confirms a violation the photo can't |
| prove β it produces a *candidate* and routes it for verification or human review. |
| 2. **Deterministic CV for detection; AI for judgement.** Fast, reliable open models do the |
| detecting and measuring. A vision-language model (VLM) is used **only** to verify |
| ambiguous candidates, to abstain when evidence is weak, and to write the human-readable |
| justification. This keeps the pipeline cheap, fast, and explainable. |
| 3. **Explainable & accountable by construction.** Every output carries: the reason, the |
| confidence, the evidence-sufficiency level, the legal reference, and a tamper-evident |
| hash. No black-box verdicts; a human can always be put in the loop. |
| 4. **Right-sized engineering.** A clean modular monolith that runs on a laptop/CPU and free |
| APIs β with a clearly documented path to horizontal scale. We spend our time on the CV |
| and the demo, not on orchestration plumbing. |
| 5. **Free and self-hostable.** Open-source models + one free API tier (Gemini). No billing. |
|
|
| --- |
|
|
| ## 3. The Violation Detectability Matrix *(the core intellectual contribution)* |
|
|
| Each violation is assigned a **tier** that determines its evidence-sufficiency level and |
| how the pipeline routes it. This table drives the Rule Engine and the routing policy. |
|
|
| | Violation | Tier | What a single image proves | Inputs needed | Default route | |
| |---|---|---|---|---| |
| | **Helmet non-compliance** | **A β Appearance** | Reliable: rider head visible, helmet/no-helmet | Detector + helmet classifier | Auto-confirm if high conf, else VLM | |
| | **Triple riding** | **A β Appearance** | Reliable: count riders linked to one motorcycle | Detector + riderβbike association | Auto-confirm if clean count, else VLM | |
| | **Seatbelt non-compliance** | **B β Hard appearance** | Weak: windshield glare/resolution make this unreliable | Windshield crop classifier *or* VLM | VLM-verify or human; mark low-confidence | |
| | **Stop-line violation** | **C β Spatial (needs calibration)** | Strong *candidate*: vehicle body past the stop-line zone | Detector + per-camera stop-line polygon | VLM-verify; human if no calibration | |
| | **Red-light violation** | **C/D β Spatial + temporal** | Candidate only: light=red (from image) AND vehicle past stop line; true running is temporal | Detector + light-state + stop-line polygon | Always VLM-verify + flag "confirm with sequence" | |
| | **Illegal parking** | **C/D β Spatial + temporal** | Candidate only: vehicle inside no-parking zone; "parked" = duration, unprovable from one frame | Detector + no-parking polygon | Candidate; flag "needs dwell-time confirmation" | |
| | **Wrong-side driving** | **D β Temporal** | Not provable: direction of travel needs motion; orientation-vs-lane is a weak proxy | Detector + lane direction (+ orientation) | Low-confidence candidate; recommend video | |
| | **License-plate recognition** | *(supporting)* | Reliable when plate is legible | Plate detector + OCR + regex | Always attempted; attach to any violation | |
|
|
| **Consequence for the build:** Tier A is our headline, demo-grade capability. Tier B is |
| best-effort with honest confidence. Tiers C/D are positioned as **candidate generation + |
| calibration-assisted evidence**, never as "we solved red-light running from a JPEG." This |
| framing is defensible and impressive; over-claiming is not. |
|
|
| --- |
|
|
| ## 4. System overview |
|
|
| A **production-credible modular monolith**: one FastAPI codebase running a linear pipeline of |
| pure-ish stages, with the heavy CV work handed to a **worker via a Redis queue** so the API |
| stays responsive. This is production-grade *and* demo-safe β real horizontal scalability (add |
| workers) without splitting into many microservices. Everything is free/self-hosted and comes |
| up with one `docker compose up`. Storage and the queue sit behind interfaces, so a developer |
| can run a lightweight local mode (in-process tasks + filesystem) when they don't want to |
| start the full stack. |
|
|
| ``` |
| upload β FastAPI api ββenqueueβββΆ Redis βββΆ Worker(s): pipeline stages 1β10 |
| β |
| 1 Ingest+QualityGate β 2 Preprocess β 3 Detect β 4 SceneGraph |
| β 5 Attributes(helmet/seatbelt/plate) β 6 RuleEngine β 7 Fuse+Route |
| β 8 VLM Verify (Gemini, only when routed) β 9 Legal map β 10 Evidence Composer |
| β |
| ββββββββββββββββββββββββββββββββββββββββ΄ββββββββββββββββββββββββββββββ |
| Postgres (JSONB: metadata + evidence graph) MinIO (S3: images) |
| ββββββββββββββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββ |
| βΌ |
| React + Vite dashboard |
| (upload Β· review queue Β· analytics Β· search) |
| scale: add workers Β· swap MinIOβS3, Postgresβmanaged Β· GPU detection workers |
| ``` |
|
|
| ### Stage responsibilities |
| 1. **Ingest + Quality Gate** β store the image + metadata (camera_id, timestamp, geo if |
| given); run a blur/exposure/resolution check. If the image is unusable, mark |
| `undeterminable` and stop β *abstaining is a feature, not a failure.* |
| 2. **Preprocess** β OpenCV: CLAHE (contrast), denoise, optional low-light enhancement |
| (gamma / Retinex-style), light deblur. Produces a normalised image; keeps the original |
| untouched for evidence. |
| 3. **Detect** β YOLO11/YOLOv8 (COCO) for car/truck/bus/motorcycle/bicycle/person/traffic |
| light; a fine-tuned model adds helmet and plate classes. Output: boxes + class + conf. |
| 4. **Scene Graph** β associate people to vehicles by geometry (riderβmotorcycle, |
| driverβcar), attach traffic lights and plates. This graph is the single source of truth |
| for all downstream reasoning (Β§5). |
| 5. **Attribute classifiers** β helmet per rider (YOLO/classifier); seatbelt per driver |
| (windshield crop classifier or VLM, Tier B); traffic-light state (HSV + classifier); |
| plate text (fast-alpr OCR + Indian-plate regex correction). |
| 6. **Rule Engine** β pure, deterministic functions, one per violation, consuming the graph |
| + optional per-camera calibration. Emits **candidates** with a rule-evidence score and |
| the tier from Β§3. No ML, no I/O β fully unit-testable. |
| 7. **Fuse + Route** β combine detection / attribute / rule scores into a fused confidence; |
| apply tier-aware routing (Β§7): auto-confirm, VLM-verify, human-review, or abstain. |
| 8. **VLM Verify** β only for routed candidates: send the cropped evidence + a strict prompt |
| to **Gemini Flash (free tier)**; get `{verified, confidence, reason}` or |
| `insufficient_evidence`. Sparingly used to respect the 1,500/day, ~10/min free quota. |
| 9. **Legal map** β deterministic lookup table: violation β Motor Vehicles Act section β |
| fine. (Correct, instant, free. Optional RAG is a bonus feature, not load-bearing β Β§10.) |
| 10. **Evidence Composer** β annotated image (boxes, labels, plate, reason) + structured |
| JSON + SHA-256 hash + audit trail; persist to Postgres (JSONB) + MinIO. |
| |
| --- |
|
|
| ## 5. The Evidence Graph (data model) |
|
|
| Detections become a typed graph, not a flat list. This is what makes the system queryable, |
| explainable, and auditable. |
|
|
| ``` |
| Nodes: |
| Vehicle { id, type, bbox, conf, attrs:{orientation?, in_zones:[...]} } |
| Person { id, role: rider|driver|pedestrian, bbox, conf, attrs:{helmet?, seatbelt?} } |
| Light { id, state: red|amber|green|unknown, bbox, conf } |
| Plate { id, text, regex_ok, conf, bbox } |
| Zone { id, kind: stop_line|no_parking|lane, polygon } # from camera calibration |
| |
| Edges: |
| rides(Person β Vehicle:motorcycle) drives(Person β Vehicle:car) |
| has_plate(Vehicle β Plate) located_in(Vehicle β Zone) |
| governed_by(Vehicle β Light) |
| ``` |
|
|
| A violation is then a small, explainable subgraph, e.g. *triple riding* = |
| `motorcycle_1` with three `rides` edges. "Why?" is answerable directly from the graph |
| (*"3 riders linked to 1 motorcycle"*). Example output payload: |
|
|
| ```json |
| { |
| "violation": "TRIPLE_RIDING", |
| "tier": "A", |
| "evidence_sufficiency": "sufficient", |
| "subjects": ["motorcycle_1", "rider_1", "rider_2", "rider_3"], |
| "scores": { "detection": 0.94, "rule": 0.90, "vlm": 0.93, "fused": 0.92 }, |
| "route": "auto_confirmed", |
| "reason": "Three distinct riders are linked to a single motorcycle.", |
| "plate": { "text": "UP32AB1234", "regex_ok": true, "conf": 0.88 }, |
| "legal": { "act": "Motor Vehicles Act, 1988", "section": "128", "fine": "βΉ1000" }, |
| "evidence_hash": "sha256:β¦", |
| "timestamp": "2026-06-20T10:31:00Z" |
| } |
| ``` |
|
|
| --- |
|
|
| ## 6. Confidence fusion & routing (with abstention) |
|
|
| Per candidate, fuse available signals with tier-specific weights: |
|
|
| ``` |
| fused = w_detΒ·detection_conf + w_attrΒ·attribute_conf + w_ruleΒ·rule_evidence (+ w_vlmΒ·vlm_conf) |
| ``` |
|
|
| Routing policy (thresholds are config, not hardcoded): |
|
|
| - **Tier A:** `fused β₯ 0.85` and unambiguous β **auto-confirm**; `0.55β0.85` β **VLM-verify**; |
| `< 0.55` β **human review**. |
| - **Tier B (seatbelt):** never auto-confirm β **VLM-verify**, then human if VLM is unsure. |
| - **Tier C (stop-line / red-light / parking):** never auto-confirm. With calibration β |
| **VLM-verify** and label `candidate`; without calibration β **human review**. |
| - **Tier D (wrong-side):** emit **low-confidence candidate** + "recommend video |
| confirmation"; never auto-confirm. |
| - **Quality gate failed or VLM returns `insufficient_evidence`** β **abstain** |
| (`undeterminable`), not a false positive. |
| |
| This is what makes the free Gemini quota workable: only the genuinely ambiguous minority |
| hits the VLM, and clear Tier-A cases auto-confirm for free. |
| |
| --- |
| |
| ## 7. VLM verification layer (Gemini free tier) |
| |
| - **Role:** auditor, never detector. It receives a *focused crop* + the proposed violation |
| + the graph evidence, and answers a strict JSON contract: |
| ``` |
| {"verified": bool, "confidence": 0.0-1.0, "reason": "<one sentence>", |
| "insufficient_evidence": bool} |
| ``` |
| - **Model:** Gemini Flash (free tier) via a provider-agnostic client so it can swap to |
| Groq / OpenRouter free models without touching call sites. |
| - **Budget discipline:** free tier = ~1,500 req/day, ~10/min. Therefore: route sparingly, |
| batch where possible, **cache by `evidence_hash`** (identical crops never re-billed), |
| and degrade gracefully to "human review" if the quota is exhausted. For the live demo, |
| pre-cache responses for the demo images. |
| |
| --- |
| |
| ## 8. Legal grounding |
| |
| For seven fixed violation types, the violationβsectionβfine mapping is a **small static |
| table** β correct, instant, and free. That table is the source of truth. |
| |
| *Optional wow-factor (only if time allows):* a tiny RAG layer (ChromaDB + |
| sentence-transformers over the Motor Vehicles Act text) that lets an officer **ask |
| free-form questions** ("what's the penalty for X, and the appeal process?"). This is a |
| bonus feature layered on top β it must never be on the critical path of issuing a verdict. |
| |
| --- |
| |
| ## 9. Evidence package (court-ready output) |
| |
| Each confirmed/escalated violation produces: |
| - **Annotated image** β original (untouched) + an overlay copy with boxes, labels, |
| plate, and the natural-language reason. |
| - **Structured JSON** β the payload in Β§5, including scores, tier, route, legal ref. |
| - **Integrity** β SHA-256 hash of the original image + an append-only audit trail |
| (who/what changed the verdict, when). This is the "admissible evidence" story. |
| - **e-challan-ready** β plate + violation + section + fine in one record, exportable. |
| |
| --- |
| |
| ## 10. Robustness to varying conditions (explicitly required) |
| |
| - **Preprocessing** handles low light (CLAHE / gamma / Retinex), rain & noise (denoise), |
| motion blur (mild deblur / sharpening). |
| - **Quality gate** quantifies blur/exposure and *abstains* on images too poor to judge β |
| preventing confident-but-wrong outputs, which is the real failure mode under bad |
| conditions. |
| - **VLM second opinion** adds robustness on ambiguous/degraded crops and can abstain. |
| - **Honest confidence** means degraded inputs surface as lower confidence β review, not as |
| silent false positives. |
| |
| --- |
| |
| ## 11. Why this wins (innovation summary) |
| |
| 1. **Evidence-sufficiencyβaware detection** β explicit tiers + abstention. Rigorous and |
| rare; directly answers "robust and accurate" without over-claiming. |
| 2. **Evidence Graph** β relationship-first, explainable single source of truth. |
| 3. **VLM-as-verifier-and-abstainer** β cuts false positives, writes human-readable reasons, |
| stays inside a free quota by only judging the hard cases. |
| 4. **Confidence routing with human-in-the-loop** β quantifiable "human-review reduction %". |
| 5. **Court-ready evidence** β legal grounding + tamper-evident hash + audit trail. |
| 6. **Calibration-optional** β works on arbitrary images for Tier A; uses optional per-camera |
| zones for Tier C; clearly degrades elsewhere. |
| |
| --- |
| |
| ## 12. Evaluation plan (mapped to the required metrics) |
| |
| - **Component metrics (P / R / F1 / mAP):** on public datasets β helmet & triple-riding |
| (Roboflow), plate detection + OCR character/whole-plate accuracy. Report per class. |
| - **End-to-end:** a hand-curated ~60β100 image set spanning day/night/rain/density, with a |
| per-violation confusion matrix and false-positive rate **per image** (the metric that |
| actually governs reviewer workload). |
| - **Operational:** mean latency/image on CPU, throughput (images/min), and the |
| **auto-confirm vs. VLM vs. human** split β human-review-reduction %. |
| - **Ablation (great for judges):** rule-only vs. rule+VLM, to quantify the false-positive |
| reduction the VLM buys. |
| - **Honesty:** report Tier A separately from C/D; never average them into one inflated |
| "accuracy" number. |
| |
| --- |
| |
| ## 13. Tech stack (free) & scalability path |
| |
| **MVP (build this) β a production-credible monolith, all free/self-hosted:** |
| - Python 3.11 Β· FastAPI Β· Uvicorn Β· **Redis-backed worker queue** (Celery or RQ/ARQ) so the |
| API stays responsive and scales by adding workers; in-process fallback for quick local dev. |
| - **Detection:** Ultralytics **YOLO11/YOLOv8** (pip, pretrained COCO + fine-tuned |
| helmet/plate from Roboflow). *(Open-vocab Grounding DINO / YOLO-World is an optional |
| enhancement, not a dependency β it's slow and painful to install.)* |
| - **Plates:** **fast-alpr** (detection + `fast-plate-ocr`, ONNX, CPU-fast) + Indian-plate |
| regex correction. |
| - **Seatbelt/ambiguous:** Gemini VLM (Tier B/C). |
| - **VLM:** **Gemini Flash free tier** via provider-agnostic `core/llm/` client β a one-line |
| swap to a paid model (Claude / GPT-4o) later, with zero code rework. Free is a demo-time |
| choice, not a capability ceiling. |
| - **Storage:** **Postgres (JSONB)** via SQLModel for metadata *and* the evidence graph |
| (relational integrity + native JSON querying); **MinIO** (free, S3-compatible) for images. |
| Both behind a `core/storage/` abstraction. |
| - **Legal:** static lookup table (+ optional ChromaDB RAG bonus). |
| - **Frontend:** React + Vite + Chart.js. |
| - **Packaging:** one `docker compose up` brings up api + worker + Postgres + Redis + MinIO. |
| |
| **Why not full microservices?** This monolith already delivers the real scalability win β a |
| true task queue and horizontally-scalable workers β in one maintainable codebase. Splitting |
| the stages into separately-deployed services adds ops burden and demo risk for little extra |
| scalability during a hackathon. |
| |
| **Scale-out path (a deployment change, not a rewrite):** add more workers; swap MinIO β AWS |
| S3 and Postgres β a managed instance; run the detection stage as dedicated GPU workers; and, |
| if ever truly needed, peel stages into separate services behind the same `storage/` and |
| `llm/` interfaces. Because everything already talks through those interfaces, none of this |
| touches business logic. |
| |
| --- |
| |
| ## 14. Pitch & demo flow (3β4 min) |
| |
| 1. **Hook (20s):** "Every team says they catch all 7 violations from a photo. They can't β |
| and neither can a real red-light camera from one frame. We built the system that knows |
| the difference." |
| 2. **Idea (40s):** hybrid CV + VLM, Evidence Graph, evidence-sufficiency tiers, routing. |
| 3. **Live demo (90s):** upload a triple-riding image β annotated result, graph panel |
| ("motorcycle β 3 riders, 1 helmet"), VLM reason, plate OCR, legal section + fine; then a |
| degraded night image β system **abstains / routes to review** (show this on purpose); |
| then the analytics + review-queue dashboard. |
| 4. **Proof (30s):** metrics table (Tier A P/R/F1, ablation showing VLM cuts false positives, |
| human-review-reduction %). |
| 5. **Close (20s):** free/self-hostable, calibration-optional, scale-out path, e-challan-ready. |
| |
| --- |
| |
| ## 15. Risks & honest limitations |
| |
| | Risk | Mitigation | |
| |---|---| |
| | Seatbelt detection unreliable | Tier B: best-effort + VLM/human; report honest confidence | |
| | Tier C/D need calibration we won't have for random images | Position as candidate generation; demo with one calibrated camera config | |
| | Free Gemini quota exhausted | Cache by hash, route sparingly, pre-cache demo, degrade to human review | |
| | Grounding DINO install eats hackathon time | Default to Ultralytics YOLO; open-vocab is optional | |
| | Public model accuracy varies on Indian scenes | Pick India-trained Roboflow models; report measured, not claimed, numbers | |
| |
| --- |
| |
| ## References |
| - Gemini API free tier limits (2026): https://tokenmix.ai/blog/gemini-api-free-tier-limits |
| - fast-alpr / fast-plate-ocr: https://github.com/ankandrew/fast-alpr |
| - Triple-riding + helmet via YOLOv8 (CCTV): https://www.atlantis-press.com/proceedings/computatia-25/126010076 |
| - Helmet violation detection, Indian smart-city (YOLOv8/TAO): https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2025.1582257/full |
| - Red-light running needs video (false positives from braking): https://www.researchgate.net/publication/3154744_An_effective_video_analysis_method_for_detecting_red_light_runners |
| - VLM on complex traffic events (GPT-4V study): https://arxiv.org/pdf/2402.02205 |
| |