face-model / README.md
Banaxi-Tech's picture
Upload face detector (YOLO11n/s), ONNX exports, scripts, model card
d176ecd verified
|
Raw History Blame Contribute Delete
4.36 kB
---
license: agpl-3.0
library_name: ultralytics
pipeline_tag: object-detection
tags:
- face-detection
- yolo11
- object-detection
- onnx
- privacy
---
# face-model — single-class face detector for blurring faces in video
YOLO11 (nano and small) fine-tuned from the COCO-pretrained Ultralytics weights as a one-class `face` detector on WIDER FACE,
plus a video script that detects, tracks and pixelates (or blacks out) faces. Built for **recall** (privacy), not leaderboard accuracy.
## Important licence notes (read before using commercially)
- The code and weights are derived from **Ultralytics YOLO11, which is AGPL-3.0**. Using them in a product or network service means
complying with AGPL-3.0 or obtaining an Ultralytics enterprise licence.
- The **training data, WIDER FACE, is listed as non-commercial** (CC BY-NC-ND 4.0 on its Hugging Face card; the original
WIDER FACE terms are for non-commercial research). Do not assume these weights are cleared for commercial use.
- `code/face_id.py` (optional `--keep` mode) expects InsightFace `buffalo_l` ONNX models, which are **non-commercial research only**.
They are *not* included in this repo; download them yourself.
## Files
| File | Notes |
|---|---|
| `yolo11n/face_yolo11n.pt` | PyTorch weights, 5.4 MB |
| `yolo11n/face_yolo11n_fp32.onnx`, `_fp16.onnx` | ONNX, fixed 640x640, batch 1 |
| `yolo11n/face_yolo11n_int8.onnx` | INT8 (ONNX Runtime static QDQ). The Detect-head decode nodes are kept in float; plain INT8 returned zero detections |
| `yolo11s/face_yolo11s.pt`, `_fp32.onnx`, `_fp16.onnx` | Larger, more accurate |
| `code/` | `blur_video.py`, `train.py`, `convert_wider_to_yolo.py`, `eval_model.py`, `recall_at_conf.py`, `quantize_int8.py`, `face_id.py` |
## Results (WIDER FACE val, 3,222 images / 39,112 faces, 640 px, all difficulty levels together)
| Model | mAP50 | mAP50-95 |
|---|---|---|
| YOLO11n `.pt` | 0.684 | 0.362 |
| YOLO11n ONNX FP32 | 0.683 | 0.362 |
| YOLO11n ONNX FP16 | 0.682 | 0.361 |
| YOLO11n ONNX INT8 | 0.667 | 0.348 |
| YOLO11s `.pt` | 0.741 | 0.401 |
| YOLO11s ONNX FP16 | 0.739 | 0.399 |
**Recall at confidence 0.25 (IoU >= 0.5)**, i.e. the fraction of labelled faces that get a box:
| Face height (px, images stored at <=640 px) | n=39,112 | YOLO11n FP16 | YOLO11n INT8 | YOLO11s FP16 |
|---|---|---|---|---|
| < 10 | 16,305 | 33.1% | 31.6% | 42.1% |
| 10-20 | 10,521 | 72.2% | 70.7% | 79.1% |
| 20-40 | 7,478 | 86.6% | 85.4% | 90.4% |
| 40-80 | 3,167 | 93.7% | 92.9% | 95.3% |
| > 80 | 1,641 | 95.6% | 95.0% | 97.3% |
| **Overall** | | **61.4%** | **60.0%** | **67.9%** |
WIDER FACE is dominated by tiny faces (median face height about 12 px). Performance on tiny, far-away faces is the main weakness;
on faces >= 20 px the nano finds about 90%. **No detector guarantees every face is caught in every frame** — spot-check any
video where missing a face matters.
Not measured: speed on phones or other hardware. Any such numbers you may see elsewhere in this project were estimates.
## Training
- Start: COCO `yolo11n.pt` / `yolo11s.pt`, fine-tuned (not from scratch), single class `face`.
- Data: WIDER FACE train, converted to YOLO format (`code/convert_wider_to_yolo.py`): invalid boxes dropped, small faces kept, images downscaled to <= 640 px long side.
- 640 px, AMP. Nano: 40 epochs. Small: 40 epochs (resumed from a checkpoint at epoch 34 with the schedule shortened to end at 40).
- Extra augmentation: random motion blur and JPEG compression to imitate video frames (`train.py`).
## Use
```bash
pip install ultralytics onnxruntime-gpu lap
python code/blur_video.py in.mp4 out.mp4 --model yolo11n/face_yolo11n_fp16.onnx -n 1 # pixelate
python code/blur_video.py in.mp4 out.mp4 --model yolo11n/face_yolo11n_fp16.onnx --mode black # solid black
```
Defaults: detect every 3 frames (`-n`) with ByteTrack in between, confidence 0.25, 15% box padding, 8 pixel blocks across the
shorter side of each face. Use `-n 1` for the strictest coverage. Audio is preserved via ffmpeg. With an ONNX model use
`--device 0` (GPU); `--device cpu` makes Ultralytics try to pip-install `onnxruntime`.
`--keep ref.jpg [...]` leaves one person visible and blurs everyone else using face recognition; every other face stays blurred
unless confirmed over several frames. It needs the InsightFace models mentioned above.