Spaces:
Runtime error
Cursor AI Skill: DeepFake Eye-Blink Detector
What This Project Is
An Enhanced Eye-Blinking LRCN (Long-term Recurrent ConvNet) for DeepFake video detection using Attentive Adversarial Training (AAT). The model detects deepfakes by analysing unnatural blink patterns and fine-grained ocular artifacts using a hybrid ViT + LSTM architecture hardened with PGD adversarial training.
Research by: Alina Chikwado Godsaves (BSc Computer Science) Supervisor: Mr. Akanji
Critical Constraints β Read Before Writing Any Code
| Constraint | Rule |
|---|---|
| Python version | Always use Python 3.11. Do NOT use 3.12 or 3.13. |
| Virtual environment | .venv311/ at project root. All scripts activate it first. |
| ViT input size | Frames are already 224Γ224 from the streaming script. Never resize again inside the model. |
| Face detection | Use MediaPipe only. Not Dlib. |
| No raw videos on disk | The streaming script processes videos in RAM and saves only .npz files. Never write .mp4 files permanently. |
| Dataset source | bitmind/FaceForensicsC23 on HuggingFace, streamed β NOT downloaded. |
| No localhost hardcoding | API base URL comes from VITE_API_URL env var in the frontend. |
No <form> tags in React |
Use onClick + FormData + axios.post() instead. |
Dataset: How Streaming Works
The project does NOT download the 17.9GB dataset zip. Instead:
HuggingFace β one video at a time (in RAM, ~5MB)
β MediaPipe extracts eye crops + EAR signal
β saves .npz file (~50KB) to data/processed/
β video discarded from memory
β repeat
Key file: src/data/stream_ff_dataset.py
- Run once to produce all
.npzfiles - Takes ~20β60 min depending on internet
- Requires
hf auth login(one-time setup) - Done - After it finishes, everything runs fully offline
HuggingFace login (one-time):
pip install huggingface_hub
hf auth login
# Paste token from https://huggingface.co/settings/tokens (Read access)
Architecture Reference
Model Pipeline
Input: eye sequence (T=16 frames, 224Γ224 RGB) + EAR signal (T floats)
β
Reshape: (B, T, 3, 224, 224) β (B*T, 3, 224, 224)
β
ViT Backbone (timm: vit_small_patch16_224, pretrained=True)
β Extract [CLS] token per frame β shape: (B*T, 384)
β Reshape back: (B, T, 384)
β
LSTM Temporal Encoder
β input: (B, T, 384), hidden_size=256, num_layers=2, dropout=0.3
β output: (B, T, 256) + final hidden
β
Blink Dynamics Head
β Concatenate LSTM output[:, -1, :] + EAR[:, -1:] β (B, 257)
β FC(257, 128) β ReLU
β Blink timing constraint (0.1β0.4s window)
β
Classifier Head
β FC(256, 128) β ReLU β Dropout(0.5) β FC(128, 2)
β Output: logits (B, 2)
Loss Function
L_total = L_ce(clean)
+ alpha * L_ce(adversarial)
+ lambda_attn * L_attn_consistency # KL div between adjacent frame attentions
+ lambda_blink * L_blink_regularizer # penalise uncertain preds during blinks
Default values: alpha=0.5, lambda_attn=0.1, lambda_blink=0.05
File Contracts (Do Not Break These)
src/data/stream_ff_dataset.py
Produces .npz files and data/metadata.csv. Each .npz contains:
{
'frames': np.uint8, # shape (16, 224, 224, 3)
'ear': np.float32, # shape (16,)
'label': np.int, # 0 = real, 1 = fake
'video_id': np.str_, # unique identifier
}
src/data/dataset.py β __getitem__ must return:
{
'frames': torch.Tensor, # shape (T, 3, 224, 224), float32, normalised [0,1]
'ear': torch.Tensor, # shape (T,), float32
'label': int, # 0 = real, 1 = fake
}
src/models/lrcn_vit.py β forward() must return:
{
'logits': torch.Tensor, # shape (B, 2)
'attn_maps': torch.Tensor, # shape (B, T, num_heads, num_patches, num_patches)
'blink_logits': torch.Tensor, # shape (B, T, 2)
}
api/inference.py β predict_video() must return:
{
'label': str, # 'REAL' or 'FAKE'
'confidence': float, # probability of fake (0.0β1.0)
'blink_rate': float, # detected blinks per second
'frame_scores': list[float], # per-frame fake probability
'attention_map_path': str | None, # path to Grad-CAM PNG or None
}
Key Source Files & Their Purpose
| File | Purpose |
|---|---|
src/data/stream_ff_dataset.py |
ONE-TIME script: streams FF++ from HuggingFace, saves .npz files |
src/data/dataset.py |
PyTorch Dataset reading .npz files + albumentations for train |
src/models/backbones.py |
build_backbone(config) β timm ViT |
src/models/lrcn_vit.py |
Full LRCN-ViT hybrid model |
src/train/train.py |
Training loop: AdamW + cosine LR + optional wandb |
src/train/adversarial.py |
FGSM and PGD attack implementations |
src/eval/evaluate.py |
Metrics: Accuracy, Precision, Recall, F1, AUC |
src/eval/ablation.py |
4 ablation configs compared |
src/viz/attention_maps.py |
Grad-CAM on eye-region patches |
api/main.py |
FastAPI: /predict + /health |
api/inference.py |
Stateless inference: video bytes β prediction dict |
frontend/src/App.jsx |
React UI: upload β result view |
Eye Aspect Ratio (EAR)
EAR = (||p2-p6|| + ||p3-p5||) / (2 * ||p1-p4||)
MediaPipe left eye landmark indices: [33, 160, 158, 133, 153, 144]
MediaPipe right eye landmark indices: [362, 385, 387, 263, 373, 380]
Use average of both eyes.
Blink threshold: EAR < 0.2 for β₯ 2 consecutive frames Normal blink duration: 0.1β0.4 seconds (3β12 frames at 30fps)
Config System
from src.utils import load_config
config = load_config('configs/train/aat_pgd.yaml')
# inherits from configs/base.yaml automatically
Never hardcode hyperparameters in model/training code β always read from config.
configs/base.yaml must include:
data:
metadata_csv: data/metadata.csv
processed_root: data/processed
seq_len: 16
img_size: 224
Albumentations (Training Only)
import albumentations as A
from albumentations.pytorch import ToTensorV2
train_transform = A.Compose([
A.HorizontalFlip(p=0.5),
A.ColorJitter(brightness=0.2, contrast=0.2, saturation=0.1, p=0.5),
A.GaussNoise(var_limit=(10.0, 50.0), p=0.3),
A.ImageCompression(quality_lower=70, quality_upper=100, p=0.3),
A.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
ToTensorV2(),
])
Do NOT apply to val or test splits.
FastAPI Setup
CORS (required β frontend runs on port 5173)
from fastapi.middleware.cors import CORSMiddleware
app.add_middleware(
CORSMiddleware,
allow_origins=["http://localhost:5173", "http://127.0.0.1:5173"],
allow_methods=["*"],
allow_headers=["*"],
)
Model loaded once at startup
from contextlib import asynccontextmanager
@asynccontextmanager
async def lifespan(app: FastAPI):
app.state.model = load_model(
checkpoint="outputs/best.pt",
config="configs/train/aat_pgd.yaml"
)
yield
app = FastAPI(lifespan=lifespan)
Start command
source .venv311/bin/activate
uvicorn api.main:app --reload --host 0.0.0.0 --port 8000
Frontend Environment Variables
frontend/.env:
VITE_API_URL=http://localhost:8000
In React:
const API_URL = import.meta.env.VITE_API_URL ?? 'http://localhost:8000'
Common Bugs to Avoid
Streaming item format: When streaming HuggingFace datasets,
item['video']may be a dict{'bytes': b'...', 'path': '...'}not a raw path. Always checkisinstance(item['video'], dict)and read.bytes.Temp file cleanup: The streaming script writes each video to a temp file for OpenCV. Always delete it with
os.unlink(tmp_path)after processing β even if an exception occurs. Use try/finally.EAR not passed to model:
dataset.pyreturnsearbut it's easy to forget to pass it through the training loop. The blink dynamics head silently fails with all-zero EAR.ViT reshape bug: Frames enter as
(B, T, 3, 224, 224). Before ViT, reshape to(B*T, 3, 224, 224). After ViT, reshape back to(B, T, embed_dim). Forgetting this causes a shape mismatch in the LSTM.PGD on full frame: Apply PGD only to eye-region frames tensor, not to any full-frame input.
Model loaded per-request: Use FastAPI lifespan to load model once. Loading per request causes 10β30s timeouts.
Split leakage: Splits in
stream_ff_dataset.pyare assigned byvideo_id, not by sequence. One video's sequences must all go to the same split.
Demo Day Checklist
-
python -m src.eval.evaluate --checkpoint outputs/best.ptruns cleanly -
/healthreturns{"status": "ok"} - Upload known fake video β returns FAKE with >70% confidence
- Upload real webcam video β returns REAL
- Frame chart renders
- App shows friendly error if API is not running
- Short (<3 second) video doesn't crash the app
Evaluation Targets
| Metric | Baseline (Gazi 2021 LRCN) | Target (AAT-ViT) |
|---|---|---|
| Accuracy | ~88% | >91% |
| AUC | ~0.90 | >0.93 |
| F1 | ~0.87 | >0.90 |
On FaceForensics++ c23, video-disjoint test split.