File size: 7,986 Bytes
af32e72
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
# PRISM AI β€” Classroom Monitoring System

An AI-powered classroom engagement and attendance monitoring system. It processes classroom video to detect, track, and classify student engagement β€” and marks attendance from a single classroom photo β€” all through a browser-based dashboard.

---

## What It Does

### Classroom Monitoring tab
Upload a classroom video and the system:
- Detects every student's face every few seconds using SCRFD (InsightFace)
- Re-identifies the same student across windows using face embeddings + **seat-position tracking** (students stay in their seats)
- Classifies each student's action per window: Attentive / Writing / Talking / On Phone / Sleeping / Distracted
- Outputs a per-student timeline, action breakdown bar, and per-window face-crop clips
- Downloadable summary JSON + CSV

### Attendance tab
- **Enroll students** from a close-up selfie video or photo (upload files or paste a local folder path)
- **Mark attendance** from a classroom photo β€” detects all faces, matches against enrolled roster, returns a marked image with names and confidence scores

---

## Project Structure

```
deployment/
β”œβ”€β”€ activity_web/                  # Flask web application
β”‚   └── backend/
β”‚       β”œβ”€β”€ app.py                 # API routes
β”‚       β”œβ”€β”€ attendance_service.py  # Face enrollment & matching
β”‚       β”œβ”€β”€ *_loader.py            # Pipeline loaders
β”‚       β”œβ”€β”€ static/                # Frontend JS + CSS
β”‚       └── templates/             # HTML (2 tabs: Classroom + Attendance)
β”œβ”€β”€ ACTIVITY CLASSIFICATION PIPELINE/
β”‚   └── student_activity_pipeline.py   # 3D CNN engagement classifier
β”œβ”€β”€ ENGAGEMENT PIPELINE/
β”œβ”€β”€ COGNITIVE PIPELINE/
β”œβ”€β”€ COMBINED PIPELINE/
β”œβ”€β”€ CLASSROOM PIPELINE/
β”‚   └── classroom_pipeline.py          # Main classroom analysis pipeline
β”œβ”€β”€ utils/
β”‚   └── retinaface_detector.py         # SAHI face detection wrapper
β”œβ”€β”€ Activity monitoring/
β”‚   β”œβ”€β”€ models/best_model/
β”‚   β”‚   └── 3dcnn_r3d18_weighted.pt    # Trained activity classifier (127 MB)
β”‚   └── Training Pipelines/assets/
β”‚       β”œβ”€β”€ yolo11m.pt                 # Phone detector (39 MB)
β”‚       └── pose_landmarker_lite.task
β”œβ”€β”€ models/retinaface_finetune/        # Fine-tuned RetinaFace ONNX
β”œβ”€β”€ yolov8s-pose.pt
β”œβ”€β”€ yolov8n-pose.pt
β”œβ”€β”€ requirements.txt
β”œβ”€β”€ Procfile
└── Dockerfile.railway
```

---

## Prerequisites

- **Python 3.10–3.12** (recommended; Python 3.14 has some package compatibility issues)
- **Git LFS** β€” required to clone the model files
- **ffmpeg** β€” for video clip encoding (optional but recommended)

---

## Local Setup

### 1. Clone the repository

```bash
git lfs install          # ensure LFS is active before cloning
git clone <repo-url>
cd deployment
```

> If you cloned before running `git lfs install`, run `git lfs pull` to download the model files.

### 2. Create a virtual environment

```bash
python3 -m venv .venv
source .venv/bin/activate        # macOS / Linux
# .venv\Scripts\activate         # Windows
```

### 3. Install dependencies

```bash
pip install -r requirements.txt
```

> First run will auto-download InsightFace `antelopev2` models (~350 MB) into `~/.insightface/`.

### 4. Set environment variables (optional)

```bash
cp .env.example .env
# edit .env as needed
```

---

## Running the Website Locally

```bash
# Standard
PORT=8080 gunicorn --bind 0.0.0.0:8080 --timeout 600 --workers 1 activity_web.backend.app:app

# macOS (required β€” prevents fork-safety crash with PyTorch)
OBJC_DISABLE_INITIALIZE_FORK_SAFETY=YES PORT=8080 gunicorn \
  --bind 0.0.0.0:8080 --timeout 600 --workers 1 \
  activity_web.backend.app:app
```

Then open **http://localhost:8080**

> **macOS note:** Always include `OBJC_DISABLE_INITIALIZE_FORK_SAFETY=YES` on macOS, otherwise the worker will crash mid-processing when PyTorch and Objective-C libraries conflict during fork.

---

## Using the Website

### Classroom Monitoring

1. Go to the **Classroom Monitoring** tab
2. Upload a classroom video (MP4 / MOV / AVI)
3. Click **Analyse classroom**
4. Results show:
   - Class-wide action distribution bar
   - Per-student collapsible cards with timeline, action labels, and face-crop clips
   - Download links for summary JSON and CSV

### Enroll a student (Attendance tab)

1. Go to the **Attendance** tab β†’ **Enroll student**
2. Enter the student's name
3. Choose an upload method:
   - **Upload files** β€” select one or more close-up photos/videos
   - **From folder path** β€” paste a local folder path (e.g. `/Users/you/jai_clips`) β€” all video/image files in that folder are used automatically
4. Click **Enroll student**

> Re-enrolling the same name **adds** embeddings to the existing gallery β€” it does not overwrite.

### Mark attendance

1. Upload a classroom photo under **Mark attendance**
2. Click **Mark attendance**
3. Returns a marked image with recognised student names and confidence scores, plus an attendance log

---

## Deploying to Railway

### Steps

1. Push this folder to a GitHub repository (with Git LFS for `*.pt` and `*.onnx.data` files)
2. Create a new Railway project β†’ **Deploy from GitHub repo**
3. Set environment variables in Railway dashboard (copy from `.env.example`)
4. Railway picks up `Procfile` automatically:

```
web: gunicorn --bind 0.0.0.0:$PORT activity_web.backend.app:app
```

### Persistent storage

Railway's ephemeral filesystem loses uploaded files on redeploy. Set these variables to point to a Railway Volume:

```
ACTIVITY_WEB_RUNTIME_DIR=/data/runtime
ACTIVITY_WEB_UPLOAD_DIR=/data/runtime/uploads
ACTIVITY_WEB_OUTPUT_DIR=/data/runtime/outputs
ACTIVITY_WEB_ATTENDANCE_DIR=/data/runtime/attendance
```

---

## Face Recognition Details

**Model:** InsightFace `antelopev2` β€” ResNet-100, Glint360K, 512-d L2-normalised embeddings

**Enrollment:**
1. Sample 1 frame/second from the enrollment video (up to 30 frames)
2. Lock on first detected face as anchor; accept subsequent frames if cosine similarity β‰₯ 0.35
3. For each accepted frame, also store degraded variants at **28 / 36 / 44 px** absolute width (INTER_AREA β†’ Gaussian Οƒ=1 β†’ JPEG q50 β†’ bicubic up) β€” bridges the gap between close-up enrollment and distant classroom faces
4. Store all embeddings + weighted-mean prototype

**Matching at attendance time:**
- Threshold: **0.38** cosine similarity
- Compares against prototype AND all stored individual embeddings (takes max)
- Only the highest-scoring face per student gets the name label

---

## Activity Classification Model

| Model | Accuracy | High engagement F1 | Macro F1 |
|---|---|---|---|
| 3D CNN R3D-18 (class-weighted) | **93.2%** | **79.3%** | **87.6%** |

Evaluated on 176 labeled student clips β€” binary: **high engagement** (attentive) vs **low engagement** (talking, head down, distracted, on phone, sleeping).

---

## Student Re-Identification (Classroom Pipeline)

Students are tracked across burst windows using a **two-stage matching strategy**:

1. **Position-first** β€” find the bank entry whose last known seat position is closest to the current track. If distance < 12% of frame width AND cosine similarity β‰₯ 0.28 β†’ assign that student. Students almost never change seats during a lecture.
2. **Embedding fallback** β€” if no positional match, use pure cosine similarity β‰₯ 0.40.

This prevents the common mis-assignment where two students with similar faces swap IDs between windows.

---

## Requirements

Key dependencies (see `requirements.txt`):

```
Flask / gunicorn
numpy / opencv-python-headless
ultralytics        # YOLOv8 / YOLO11
insightface        # face detection + recognition
torch / torchvision
dlib               # cognitive pipeline landmarks
deepface           # emotion recognition
```

> `dlib` requires CMake. macOS: `brew install cmake`. Ubuntu: `apt install cmake build-essential`.