kumaraguruvs's picture
Add Hugging Face Spaces YAML frontmatter to README
3a1d68c verified
|
Raw
History Blame Contribute Delete
3.95 kB
---
title: Forensic Touch Tracker
emoji: 🕵️
colorFrom: blue
colorTo: red
sdk: gradio
sdk_version: "4.44.0"
python_version: "3.10"
app_file: app.py
pinned: false
---
# 🕵️ Forensic Touch Tracker
A computer-vision pipeline for **tracking human movements** and **logging touch points** so forensic investigators know exactly where to seek fingerprints.
## How it works
| Stage | Model | What it does |
|-------|-------|--------------|
| **1. Person tracking** | SAM3 Video (`facebook/sam3`) | Detects & tracks every person with a **persistent ID** across all frames |
| **2. Surface segmentation** | SAM3 Video (`facebook/sam3`) | Segments forensic-relevant surfaces: *door handles, countertops, desks, walls, railings, keyboards, phones, cups, bottles, drawer handles* |
| **3. Body pose** | ViTPose (`usyd-community/vitpose-base-simple`) | Estimates 17 COCO keypoints per person (wrists used as coarse touch proxies) |
| **4. Hand landmarks** | MediaPipe Hands | Detects 21 landmarks per hand, extracting **5 fingertip locations** for precise touch localization |
| **5. Contact inference** | Proximity-based geometric check | When a fingertip or wrist is within `contact_threshold_px` of a segmented surface for ≥ `min_contact_frames` consecutive frames, a **TouchEvent** is logged |
## Forensic output format (JSONL)
Each line in `forensic_touch_log.jsonl` is a structured event:
```json
{
"timestamp_seconds_start": 12.47,
"timestamp_seconds_end": 14.12,
"frame_index_start": 374,
"frame_index_end": 423,
"person_id": 3,
"body_part": "hand_0_index_tip",
"touch_point": {"x": 412, "y": 298},
"surface": "countertop",
"confidence": 0.91,
"num_frames": 50
}
```
This gives forensic teams:
- **Who** touched it (`person_id`)
- **What** was touched (`surface`)
- **Where** on the frame (`touch_point.x`, `touch_point.y`)
- **When** (`timestamp_seconds_start`, `timestamp_seconds_end`, `frame_index_start`, `frame_index_end`)
- **How sure** (`confidence`)
## Running locally
```bash
pip install torch transformers opencv-python mediapipe gradio numpy Pillow
python app.py
```
Then open `http://localhost:7860` and upload a video.
> **Hardware note:** SAM3 is ~860M params. A GPU with ≥16 GB VRAM (e.g. T4, A10G) is strongly recommended.
## Configuration knobs
| Parameter | Default | Description |
|-----------|---------|-------------|
| `max_frames` | 0 | Limit frames for quick tests (0 = process whole video) |
| `contact_threshold_px` | 25 | Pixel distance from fingertip to surface mask to declare a "touch" |
| `min_contact_frames` | 3 | Consecutive frames required before logging (reduces false positives) |
| `confidence_threshold` | 0.5 | Minimum detection score for person / surface |
| `surface_prompts` | 12 prompts | Text prompts SAM3 uses to find forensic-relevant surfaces |
## Limitations & future upgrades
- **No direct "contact classifier"** is available on HuggingFace yet. The current pipeline uses geometric proximity, which works well but is not as precise as research models such as **InteractVLM**, **DECO**, or **HOT** (which can infer 3D body-scene contact from images). If you need research-grade contact detection, those repos can be integrated as an additional stage.
- **Depth / stereo** is not used; true 3D contact would benefit from a depth sensor or monocular depth estimator (e.g. Depth Anything V2).
- **Hand keypoint quality** depends on resolution; very small hands in wide surveillance shots may need an upsampling step.
## Citation / References
- SAM 3: *Segment Anything with Concepts* (Meta AI, 2025)
- ViTPose: *ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation* (arXiv:2204.12484)
- RT-DETR: *DETRs Beat YOLOs on Real-time Object Detection* (arXiv:2304.08069)
- MediaPipe Hands (Google)
- InteractVLM / DECO / HOT — research-grade contact detection (not yet on HF Hub)
---
Built for forensic evidence collection & verification. 🚔🔬