Spaces:
Running
title: PRISM AI
emoji: π
colorFrom: blue
colorTo: indigo
sdk: docker
app_port: 8080
pinned: false
PRISM AI β Classroom Monitoring System
An AI-powered classroom attendance and engagement monitoring system with a browser-based dashboard, designed to run on a teacher's own laptop.
What It Does
| Tab | Function |
|---|---|
| Attendance | Mark attendance from one or more classroom photos using face recognition, per classroom (CSE 1β8) |
| Enroll Student | Students self-enroll (upload, folder path, or guided camera recording) β includes a QR-code flow so a whole class can enroll from their own phones |
| Classroom Monitoring | Upload a classroom video β detect and track every student β classify engagement per window β per-student timeline with clips |
| How it Works | In-app explanation of both pipelines, kept in sync with the actual thresholds/behavior below |
Attendance β Face Recognition System
Detection & recognition model
InsightFace SCRFD (1280Γ1280 detection) for face detection + landmarks, AdaFace IR-101 (WebFace12M, 512-d L2-normalised embeddings) for recognition. AdaFace replaced the previous glintr100/antelopev2 backbone after an offline evaluation showed cleaner separation between genuine and impostor matches on real classroom photos.
Per-classroom rosters
Each of the 8 classrooms (cse1βcse8) has its own independent JSON store β enrolling into one classroom never affects another, and a teacher picks a classroom before enrolling or marking attendance.
How enrollment works
- Provide a face sample β upload photo(s)/video, point at a local folder of clips, or record directly in the browser (see below).
- Up to 30 frames are sampled evenly across the clip (sequential decode, never seeking β seeking is unreliable for live-recorded webm from a browser's
MediaRecorder). - Anchor-based tracking: the first accepted frame is the identity anchor; later frames are kept only if cosine similarity β₯ 0.35 against it, so a multi-person video can't mix identities.
- Each accepted frame contributes its full-quality embedding plus 2 degraded copies (downscaled to 50% / 30% then upscaled back) so the gallery also matches small, distant classroom faces, not just close-up enrollment shots.
- Embeddings are capped at 128 per student (oldest evicted first) with a weighted-mean prototype recomputed on every change; outliers more than 0.50 cosine distance from the running centroid are dropped before the prototype is built.
Guided camera recording
The in-browser recorder walks a student through a ~37-second sequence with spoken (not just written) instructions β including holding the phone at arm's length so the whole face fits in the on-screen oval, since a face filling the whole frame is a common cause of the detector missing it entirely. It also requests a real capture resolution (1280Γ720 ideal) rather than trusting the browser's default, which on some phones can be as low as 480Γ640.
Self-enrollment via QR code
A teacher can generate a QR code (Attendance tab β Show enrollment QR) that points students straight at their classroom's enroll page β no manual classroom picking, no need to be on the same Wi-Fi. This works by pairing the local server with a cloudflared quick tunnel:
python start_enrollment_session.py
This starts the Flask app (threaded, so a burst of students enrolling at once doesn't serialize into a queue) and the tunnel together, and prints the public URL once it's live. Requires cloudflared installed once (brew install cloudflared on macOS). The tunnel is ephemeral β a fresh random URL every time the script (re)starts, and it self-heals if the tunnel reconnects mid-session with a new hostname.
On a hosted deployment (e.g. the Hugging Face Space) there's no tunnel to run β the QR code falls back to the Space's own public URL automatically, since it's already internet-reachable. No start_enrollment_session.py needed there; localhost is the only host this fallback refuses to use, since that's never reachable from a student's phone.
How attendance marking works
- Upload one or more classroom photos at once β a student only needs to be clearly caught in any one of them to count present, so someone missed or turned away in one shot can still be caught by another.
- Every detected face across all photos is matched by cosine similarity against every enrolled student's prototype and individual stored embeddings; each student's best match across all photos wins.
- Three-tier result:
- < 0.28 similarity β no name candidate at all β Unknown (red box).
- 0.28β0.30 β a real candidate, not confident enough to auto-confirm β Suspicious (amber box, for a teacher to review).
- β₯ 0.30 β Present (green box).
- Anyone enrolled but not matched at either tier β Absent.
- Face crops (Present, Suspicious, and Unknown) stay hidden by default and reveal on demand via a "Show face" button, cropped from the pre-annotation photo so the reveal isn't obscured by a box/label.
Reinforcement β teachers correcting the model
Gallery growth only ever happens from an explicit teacher action β never automatically just because a photo scored a high confidence. (An earlier version auto-added any β₯0.60 Present match; that meant re-testing with the same handful of demo photos silently folded them into the gallery, so a second run compared a photo against an embedding derived from itself β inflated, unrealistic-looking confidence and ballooning embedding counts for no real reason.)
- Confirming a Suspicious match adds that embedding to the student's gallery and marks them present.
- Rejecting a Suspicious match doesn't just discard it β the same face gets re-matched against the roster excluding the rejected student, and lands wherever that turns up: a confident hit β straight to Present, a borderline one β a fresh Suspicious entry for the new candidate, nothing left β into the Unknown pool. The wrongly-suggested student drops to Absent unless already seen elsewhere in the same result. (This rematch itself doesn't auto-grow the gallery either.)
- Unknown faces can be directly assigned to any enrolled student via a dropdown on each face card β reinforces that student's gallery and marks them present, since a teacher pointing at a photo and naming someone is direct evidence they were there.
Classroom Monitoring β Pipeline
The Classroom pipeline (CLASSROOM PIPELINE/classroom_pipeline.py) processes a lecture video in 30-second, 24-frame bursts rather than every frame:
| Step | How |
|---|---|
| Detection & tracking | YOLOv8-pose for body/keypoints, InsightFace (SCRFD) for faces + 106-point landmarks, linked into per-student tracks by IoU overlap within a burst |
| Re-identification | AdaFace IR-101 embeddings, resolved as a one-to-one assignment per burst (β₯ 0.35 similarity) β two different people in the same burst can't collapse into one identity |
| Roster recognition | The same embedding is checked read-only against the enrolled Attendance roster (β₯ 0.35, with a margin over the runner-up) β a confident match shows the real name instead of an anonymous student_00N label |
| Signal extraction | 106-point landmarks drive mouth open/closed + head yaw/pitch; YOLO for phone detection; optical flow for motion; DeepFace for dominant emotion |
| Action classification | Priority ruleset: On Phone β Sleeping β Writing β Talking β Attentive β otherwise Distracted. "Attentive" is driven by mouth-closed percentage (β₯ 80%) β a calibration pass against a labeled reference dataset found eye-openness (EAR) had no measurable correlation with attentiveness, while mouth state separated attentive/non-attentive far more cleanly |
| Engagement rollup | Fraction of attentive windows β High (β₯ 70%) / Medium (β₯ 40%) / Low, plus a short saved clip per window |
Known limitations: phone detection is a generic, un-fine-tuned COCO model and currently finds close to none of the real phones in testing β needs replacing, not re-tuning. Attentive/not-attentive classification runs at roughly 67β70% accuracy against a 156-clip labeled reference set β a real improvement over the previous EAR-based approach, but not a solved problem.
Running Locally
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# Download YOLO model weights (Git LFS pointers β run once after cloning)
python download_models.py
deepface(emotion detection) pulls intensorflow, which only ships wheels for Python 3.9β3.12 β create the venv with one of those versions, not whatever newest Python happens to be installed. The Docker image (below) already usespython:3.10-slim.
Just the dashboard, no QR/tunnel:
OBJC_DISABLE_INITIALIZE_FORK_SAFETY=YES python -m activity_web.backend.app
Dashboard + QR self-enrollment (starts the tunnel too):
python start_enrollment_session.py
Deployment
See deployment/README.md for full Railway deployment instructions.
The Hugging Face Space build (this repo's
Dockerfile) has no persistent storage β enrolled rosters don't survive a rebuild there. For a real classroom, run this locally on the teacher's own machine (see above) so attendance data persists between sessions; the Space is a demo/showcase deployment, not where you'd actually enroll students.
Repository Structure
CLASSROOM PIPELINE/ Main classroom analysis pipeline (mouth/gaze/YOLO/emotion, AdaFace re-ID)
ENGAGEMENT PIPELINE/ YOLOv8-pose engagement signals
COGNITIVE PIPELINE/ EAR / gaze / emotion (dlib + DeepFace)
COMBINED PIPELINE/ Merged engagement + cognitive
activity_web/backend/ Flask web app β app.py, attendance_service.py, templates, static (incl. camera-recorder.js, enroll.js)
utils/ adaface_backbone.py (recognition), roster_match.py, retinaface_detector.py (SAHI + MTCNN + buffalo_l)
Activity monitoring/models/ Trained model weights
deployment/ Self-contained Railway deployment build
start_enrollment_session.py One-command launcher: threaded Flask server + cloudflared tunnel for QR self-enrollment