Spaces:
Running
Running
Satyam S
Fix QR enrollment on hosted deployments, remove auto gallery-growth on high-confidence matches
a086f3c | title: PRISM AI | |
| emoji: π | |
| colorFrom: blue | |
| colorTo: indigo | |
| sdk: docker | |
| app_port: 8080 | |
| pinned: false | |
| # PRISM AI β Classroom Monitoring System | |
| An AI-powered classroom attendance and engagement monitoring system with a browser-based dashboard, designed to run on a teacher's own laptop. | |
| --- | |
| ## What It Does | |
| | Tab | Function | | |
| |---|---| | |
| | **Attendance** | Mark attendance from one or more classroom photos using face recognition, per classroom (CSE 1β8) | | |
| | **Enroll Student** | Students self-enroll (upload, folder path, or guided camera recording) β includes a QR-code flow so a whole class can enroll from their own phones | | |
| | **Classroom Monitoring** | Upload a classroom video β detect and track every student β classify engagement per window β per-student timeline with clips | | |
| | **How it Works** | In-app explanation of both pipelines, kept in sync with the actual thresholds/behavior below | | |
| --- | |
| ## Attendance β Face Recognition System | |
| ### Detection & recognition model | |
| **InsightFace SCRFD** (1280Γ1280 detection) for face detection + landmarks, **AdaFace IR-101** (WebFace12M, 512-d L2-normalised embeddings) for recognition. AdaFace replaced the previous glintr100/antelopev2 backbone after an offline evaluation showed cleaner separation between genuine and impostor matches on real classroom photos. | |
| ### Per-classroom rosters | |
| Each of the 8 classrooms (`cse1`β`cse8`) has its own independent JSON store β enrolling into one classroom never affects another, and a teacher picks a classroom before enrolling or marking attendance. | |
| ### How enrollment works | |
| 1. Provide a face sample β upload photo(s)/video, point at a local folder of clips, or record directly in the browser (see below). | |
| 2. Up to **30 frames** are sampled evenly across the clip (sequential decode, never seeking β seeking is unreliable for live-recorded webm from a browser's `MediaRecorder`). | |
| 3. **Anchor-based tracking**: the first accepted frame is the identity anchor; later frames are kept only if cosine similarity β₯ **0.35** against it, so a multi-person video can't mix identities. | |
| 4. Each accepted frame contributes its full-quality embedding plus **2 degraded copies** (downscaled to 50% / 30% then upscaled back) so the gallery also matches small, distant classroom faces, not just close-up enrollment shots. | |
| 5. Embeddings are capped at **128 per student** (oldest evicted first) with a weighted-mean prototype recomputed on every change; outliers more than **0.50** cosine distance from the running centroid are dropped before the prototype is built. | |
| ### Guided camera recording | |
| The in-browser recorder walks a student through a ~37-second sequence with **spoken** (not just written) instructions β including holding the phone at arm's length so the whole face fits in the on-screen oval, since a face filling the whole frame is a common cause of the detector missing it entirely. It also requests a real capture resolution (1280Γ720 ideal) rather than trusting the browser's default, which on some phones can be as low as 480Γ640. | |
| ### Self-enrollment via QR code | |
| A teacher can generate a QR code (Attendance tab β **Show enrollment QR**) that points students straight at their classroom's enroll page β no manual classroom picking, no need to be on the same Wi-Fi. This works by pairing the local server with a `cloudflared` quick tunnel: | |
| ```bash | |
| python start_enrollment_session.py | |
| ``` | |
| This starts the Flask app (threaded, so a burst of students enrolling at once doesn't serialize into a queue) and the tunnel together, and prints the public URL once it's live. Requires `cloudflared` installed once (`brew install cloudflared` on macOS). The tunnel is ephemeral β a fresh random URL every time the script (re)starts, and it self-heals if the tunnel reconnects mid-session with a new hostname. | |
| On a hosted deployment (e.g. the Hugging Face Space) there's no tunnel to run β the QR code falls back to the Space's own public URL automatically, since it's already internet-reachable. No `start_enrollment_session.py` needed there; `localhost` is the only host this fallback refuses to use, since that's never reachable from a student's phone. | |
| ### How attendance marking works | |
| 1. Upload **one or more** classroom photos at once β a student only needs to be clearly caught in *any one* of them to count present, so someone missed or turned away in one shot can still be caught by another. | |
| 2. Every detected face across all photos is matched by cosine similarity against every enrolled student's prototype and individual stored embeddings; each student's *best* match across all photos wins. | |
| 3. Three-tier result: | |
| - **< 0.28** similarity β no name candidate at all β **Unknown** (red box). | |
| - **0.28β0.30** β a real candidate, not confident enough to auto-confirm β **Suspicious** (amber box, for a teacher to review). | |
| - **β₯ 0.30** β **Present** (green box). | |
| - Anyone enrolled but not matched at either tier β **Absent**. | |
| 4. Face crops (Present, Suspicious, and Unknown) stay hidden by default and reveal on demand via a **"Show face"** button, cropped from the pre-annotation photo so the reveal isn't obscured by a box/label. | |
| ### Reinforcement β teachers correcting the model | |
| Gallery growth only ever happens from an explicit teacher action β never automatically just because a photo scored a high confidence. (An earlier version auto-added any β₯0.60 Present match; that meant re-testing with the same handful of demo photos silently folded them into the gallery, so a second run compared a photo against an embedding derived from itself β inflated, unrealistic-looking confidence and ballooning embedding counts for no real reason.) | |
| - **Confirming** a Suspicious match adds that embedding to the student's gallery and marks them present. | |
| - **Rejecting** a Suspicious match doesn't just discard it β the same face gets re-matched against the roster *excluding* the rejected student, and lands wherever that turns up: a confident hit β straight to Present, a borderline one β a fresh Suspicious entry for the new candidate, nothing left β into the Unknown pool. The wrongly-suggested student drops to Absent unless already seen elsewhere in the same result. (This rematch itself doesn't auto-grow the gallery either.) | |
| - **Unknown faces** can be directly assigned to any enrolled student via a dropdown on each face card β reinforces that student's gallery and marks them present, since a teacher pointing at a photo and naming someone is direct evidence they were there. | |
| --- | |
| ## Classroom Monitoring β Pipeline | |
| The **Classroom pipeline** (`CLASSROOM PIPELINE/classroom_pipeline.py`) processes a lecture video in 30-second, 24-frame bursts rather than every frame: | |
| | Step | How | | |
| |---|---| | |
| | **Detection & tracking** | YOLOv8-pose for body/keypoints, InsightFace (SCRFD) for faces + 106-point landmarks, linked into per-student tracks by IoU overlap within a burst | | |
| | **Re-identification** | AdaFace IR-101 embeddings, resolved as a one-to-one assignment per burst (β₯ 0.35 similarity) β two different people in the same burst can't collapse into one identity | | |
| | **Roster recognition** | The same embedding is checked read-only against the enrolled Attendance roster (β₯ 0.35, with a margin over the runner-up) β a confident match shows the real name instead of an anonymous `student_00N` label | | |
| | **Signal extraction** | 106-point landmarks drive mouth open/closed + head yaw/pitch; YOLO for phone detection; optical flow for motion; DeepFace for dominant emotion | | |
| | **Action classification** | Priority ruleset: On Phone β Sleeping β Writing β Talking β Attentive β otherwise Distracted. "Attentive" is driven by mouth-closed percentage (β₯ 80%) β a calibration pass against a labeled reference dataset found eye-openness (EAR) had no measurable correlation with attentiveness, while mouth state separated attentive/non-attentive far more cleanly | | |
| | **Engagement rollup** | Fraction of attentive windows β High (β₯ 70%) / Medium (β₯ 40%) / Low, plus a short saved clip per window | | |
| **Known limitations**: phone detection is a generic, un-fine-tuned COCO model and currently finds close to none of the real phones in testing β needs replacing, not re-tuning. Attentive/not-attentive classification runs at roughly 67β70% accuracy against a 156-clip labeled reference set β a real improvement over the previous EAR-based approach, but not a solved problem. | |
| --- | |
| ## Running Locally | |
| ```bash | |
| python3 -m venv .venv && source .venv/bin/activate | |
| pip install -r requirements.txt | |
| # Download YOLO model weights (Git LFS pointers β run once after cloning) | |
| python download_models.py | |
| ``` | |
| > `deepface` (emotion detection) pulls in `tensorflow`, which only ships wheels for Python 3.9β3.12 β create the venv with one of those versions, not whatever newest Python happens to be installed. The Docker image (below) already uses `python:3.10-slim`. | |
| **Just the dashboard**, no QR/tunnel: | |
| ```bash | |
| OBJC_DISABLE_INITIALIZE_FORK_SAFETY=YES python -m activity_web.backend.app | |
| ``` | |
| **Dashboard + QR self-enrollment** (starts the tunnel too): | |
| ```bash | |
| python start_enrollment_session.py | |
| ``` | |
| Open **http://localhost:8080** | |
| --- | |
| ## Deployment | |
| See [`deployment/README.md`](deployment/README.md) for full Railway deployment instructions. | |
| > The Hugging Face Space build (this repo's `Dockerfile`) has no persistent storage β enrolled rosters don't survive a rebuild there. For a real classroom, run this locally on the teacher's own machine (see above) so attendance data persists between sessions; the Space is a demo/showcase deployment, not where you'd actually enroll students. | |
| --- | |
| ## Repository Structure | |
| ``` | |
| CLASSROOM PIPELINE/ Main classroom analysis pipeline (mouth/gaze/YOLO/emotion, AdaFace re-ID) | |
| ENGAGEMENT PIPELINE/ YOLOv8-pose engagement signals | |
| COGNITIVE PIPELINE/ EAR / gaze / emotion (dlib + DeepFace) | |
| COMBINED PIPELINE/ Merged engagement + cognitive | |
| activity_web/backend/ Flask web app β app.py, attendance_service.py, templates, static (incl. camera-recorder.js, enroll.js) | |
| utils/ adaface_backbone.py (recognition), roster_match.py, retinaface_detector.py (SAHI + MTCNN + buffalo_l) | |
| Activity monitoring/models/ Trained model weights | |
| deployment/ Self-contained Railway deployment build | |
| start_enrollment_session.py One-command launcher: threaded Flask server + cloudflared tunnel for QR self-enrollment | |
| ``` | |