--- license: agpl-3.0 library_name: ultralytics pipeline_tag: object-detection tags: - face-detection - yolo11 - object-detection - onnx - privacy --- # face-model — single-class face detector for blurring faces in video YOLO11 (nano and small) fine-tuned from the COCO-pretrained Ultralytics weights as a one-class `face` detector on WIDER FACE, plus a video script that detects, tracks and pixelates (or blacks out) faces. Built for **recall** (privacy), not leaderboard accuracy. ## Important licence notes (read before using commercially) - The code and weights are derived from **Ultralytics YOLO11, which is AGPL-3.0**. Using them in a product or network service means complying with AGPL-3.0 or obtaining an Ultralytics enterprise licence. - The **training data, WIDER FACE, is listed as non-commercial** (CC BY-NC-ND 4.0 on its Hugging Face card; the original WIDER FACE terms are for non-commercial research). Do not assume these weights are cleared for commercial use. - `code/face_id.py` (optional `--keep` mode) expects InsightFace `buffalo_l` ONNX models, which are **non-commercial research only**. They are *not* included in this repo; download them yourself. ## Files | File | Notes | |---|---| | `yolo11n/face_yolo11n.pt` | PyTorch weights, 5.4 MB | | `yolo11n/face_yolo11n_fp32.onnx`, `_fp16.onnx` | ONNX, fixed 640x640, batch 1 | | `yolo11n/face_yolo11n_int8.onnx` | INT8 (ONNX Runtime static QDQ). The Detect-head decode nodes are kept in float; plain INT8 returned zero detections | | `yolo11s/face_yolo11s.pt`, `_fp32.onnx`, `_fp16.onnx` | Larger, more accurate | | `code/` | `blur_video.py`, `train.py`, `convert_wider_to_yolo.py`, `eval_model.py`, `recall_at_conf.py`, `quantize_int8.py`, `face_id.py` | ## Results (WIDER FACE val, 3,222 images / 39,112 faces, 640 px, all difficulty levels together) | Model | mAP50 | mAP50-95 | |---|---|---| | YOLO11n `.pt` | 0.684 | 0.362 | | YOLO11n ONNX FP32 | 0.683 | 0.362 | | YOLO11n ONNX FP16 | 0.682 | 0.361 | | YOLO11n ONNX INT8 | 0.667 | 0.348 | | YOLO11s `.pt` | 0.741 | 0.401 | | YOLO11s ONNX FP16 | 0.739 | 0.399 | **Recall at confidence 0.25 (IoU >= 0.5)**, i.e. the fraction of labelled faces that get a box: | Face height (px, images stored at <=640 px) | n=39,112 | YOLO11n FP16 | YOLO11n INT8 | YOLO11s FP16 | |---|---|---|---|---| | < 10 | 16,305 | 33.1% | 31.6% | 42.1% | | 10-20 | 10,521 | 72.2% | 70.7% | 79.1% | | 20-40 | 7,478 | 86.6% | 85.4% | 90.4% | | 40-80 | 3,167 | 93.7% | 92.9% | 95.3% | | > 80 | 1,641 | 95.6% | 95.0% | 97.3% | | **Overall** | | **61.4%** | **60.0%** | **67.9%** | WIDER FACE is dominated by tiny faces (median face height about 12 px). Performance on tiny, far-away faces is the main weakness; on faces >= 20 px the nano finds about 90%. **No detector guarantees every face is caught in every frame** — spot-check any video where missing a face matters. Not measured: speed on phones or other hardware. Any such numbers you may see elsewhere in this project were estimates. ## Training - Start: COCO `yolo11n.pt` / `yolo11s.pt`, fine-tuned (not from scratch), single class `face`. - Data: WIDER FACE train, converted to YOLO format (`code/convert_wider_to_yolo.py`): invalid boxes dropped, small faces kept, images downscaled to <= 640 px long side. - 640 px, AMP. Nano: 40 epochs. Small: 40 epochs (resumed from a checkpoint at epoch 34 with the schedule shortened to end at 40). - Extra augmentation: random motion blur and JPEG compression to imitate video frames (`train.py`). ## Use ```bash pip install ultralytics onnxruntime-gpu lap python code/blur_video.py in.mp4 out.mp4 --model yolo11n/face_yolo11n_fp16.onnx -n 1 # pixelate python code/blur_video.py in.mp4 out.mp4 --model yolo11n/face_yolo11n_fp16.onnx --mode black # solid black ``` Defaults: detect every 3 frames (`-n`) with ByteTrack in between, confidence 0.25, 15% box padding, 8 pixel blocks across the shorter side of each face. Use `-n 1` for the strictest coverage. Audio is preserved via ffmpeg. With an ONNX model use `--device 0` (GPU); `--device cpu` makes Ultralytics try to pip-install `onnxruntime`. `--keep ref.jpg [...]` leaves one person visible and blurs everyone else using face recognition; every other face stays blurred unless confirmed over several frames. It needs the InsightFace models mentioned above.