TinyFaceNet-1M: a 1.0M-param face embedding net distilled from ArcFace, for edge devices
A MobileFaceNet-style CNN (112x112 RGB in, 128-d L2-normalised embedding out) trained by knowledge distillation from insightface buffalo_l (ArcFace, 65M params) on a small private photo archive, then quantised to INT8 with ONNX Runtime. Built as the recognition core of an offline family-photo organiser that runs on a Raspberry Pi 5.
This repo is an end-to-end recipe, not just weights: data pipeline, training, quantisation study, webcam test and the lessons learned are all here.
β What this model is for
Closed-set identification of a small, enrolled group of people on an edge device. Enroll N people (family, household, a small team: ~5-20 faces each with
code/enroll.py), then for every detected face answer "which of my N people is this, or nobody?"
Task Fit Evidence Tag family members in a personal photo/video archive (offline, on a Pi) β Primary use case 94-98 % on held-out dates, 9 identities Reject strangers as "unknown" β Works with THR 0.55 97.2 % of 12k unseen faces rejected Smart-doorbell / home-cam "is this a household member?" β Same problem low-light v2 + 7-frame vote in live_test.pyRun anywhere with ~1 MB and no GPU β INT8 ONNX, 1.03 M params, 3 ms/face desktop CPU Recipe to distil your own tiny face net from ArcFace β Everything is here code/train.py,quantize_eval.pyGeneral face recognition / verification of unknown people β Not this model LFW 62.6 % (chance 50 %) 1:N search over large galleries, surveillance, identity verification β Do not use cannot separate strangers from each other Rule of thumb: if every person you care about can be enrolled in advance, this works. If the model has to compare two people it has never seen, it does not.
π₯οΈ Where it runs
Embedder cost: 171 M MACs per face, 1.1 MB INT8 weights, 312 KB peak activation (INT8). Only the desktop row is measured; the rest are estimates from MACs and public per-core throughput, so expect Β±2x. The face detector, not this embedder, is the bottleneck on small boards: pair it with a tiny detector (BlazeFace ~0.1 M, SCRFD-500M, Ultra-Light-Fast ~1 MB) rather than the buffalo_l
det_10gused for training.
Device Runs? Est. embed time / face (INT8) Notes Raspberry Pi 5 (Cortex-A76 x4) β Target device ~10-20 ms onnxruntime-arm64, 4 threads; whole pipeline (tiny detector + align + embed) ~5-10 fps Raspberry Pi 4 / 400 / CM4 (Cortex-A72 x4) β ~30-60 ms same stack, ~2-3x slower than Pi 5 Raspberry Pi 3 / 3B+ (Cortex-A53 x4) β ~80-150 ms fine for photo-archive batch jobs, slow for live video Raspberry Pi Zero 2 W (Cortex-A53 x4, 512 MB) β ~100-200 ms 512 MB is enough for ORT + INT8 model; use a 160 px detector Raspberry Pi Zero / Zero W (ARMv6 single core) β οΈ Barely ~2-5 s works, but only for offline batch tagging NVIDIA Jetson Orin Nano / Nano β Overkill <2 ms TensorRT or ORT-CUDA; run the full buffalo_l detector here Google Coral Edge TPU β ~1-3 ms needs ONNX β TFLite INT8 conversion, then the edgetpu compiler Rockchip RK3566/RK3588 (Orange Pi, Radxa), Intel N100 mini-PC β ~5-40 ms ORT CPU; RK3588 NPU via RKNN also possible Android / iOS phone β ~5-15 ms ORT Mobile / NNAPI / CoreML from the ONNX file ESP32-S3 (240 MHz, 8 MB PSRAM) β οΈ Experimental ~3-8 s fits in PSRAM via esp-dl / TFLite-Micro; no room for a detector, only crop-and-embed Raspberry Pi Pico 2 / Pico 2 W (RP2350, 520 KB SRAM, 4 MB flash) β Not practical ~5-15 s weights fit flash, activations barely fit SRAM, no SIMD; and no detector Raspberry Pi Pico / Pico W (RP2040, 264 KB SRAM, 2 MB flash) β No β 312 KB activations do not fit; use the Pico only as a camera/IO companion to a Pi Zero 2 Arduino (AVR, Cortex-M0/M4 boards) β No β orders of magnitude short on RAM and compute Minimum sensible target: Pi Zero 2 W. Comfortable target: Pi 4. Sweet spot: Pi 5. Microcontrollers (Pico, ESP32) should capture frames and hand them to a Linux board.
What is published, what is not
| Published | Not published |
|---|---|
Backbone weights (.pt, .safetensors, ONNX fp32 + INT8) |
Any photo, video, or face crop |
All code (code/) |
Identity centroids / classifier head (biometric templates) |
| Analysis & quantisation reports | Names of people: everywhere they are referred to as PERSON_A ... PERSON_I + other |
| Training logs (identities as variables) | The labels.json cluster-to-name map (an example template is included) |
The checkpoints contain only the embedding backbone. The training-time distillation head and the 10-way classifier head were removed before upload, so the model cannot by itself name anyone. You enroll your own people with code/enroll.py.
Results
Public benchmark: LFW (6000 pairs, standard 10-fold protocol)
| Model | LFW acc | AUC | TAR@FAR 1e-2 | mean cos: same / different person |
|---|---|---|---|---|
| buffalo_l teacher (reference) | 99.85 % | 0.999 | 0.998 | 0.67 / 0.00 |
| TinyFaceNet v1 | 62.6 % | 0.672 | 0.072 | 0.76 / 0.64 |
| TinyFaceNet v2 | 63.6 % | 0.685 | 0.078 | 0.79 / 0.69 |
Chance is 50 %. This model is not a general face recogniser. Distilled on ~5k faces of 10 identities, the embedding only learned to separate those identities; two strangers land at cosine 0.64 on average (teacher: 0.00). The teacher scoring 99.85 % through the same pipeline confirms the evaluation itself is sound. Reproduce with code/eval_lfw.py (LFW via scikit-learn, re-aligned with buffalo_l; numbers in analysis/lfw_model_v*.json).
The fix is data, not architecture: run the same train.py recipe on 50-100k public unlabeled faces (only teacher embeddings are needed, no labels). MobileFaceNet-class nets reach 99.5 % LFW when trained on MS1M.
Stranger test (what the closed-set use case actually needs)
All 12,000 LFW faces treated as people outside the enrolled set, scored against the 9 enrolled centroids (v1):
| Open-set THR | Stranger accepted as an enrolled person | Enrolled faces still recognised (49 held-out) |
|---|---|---|
| 0.45 | 6.2 % | 93.9 % |
| 0.55 (default) | 2.8 % | 93.9 % |
| 0.60 | 1.7 % | 93.9 % |
| 0.65 | 1.0 % | 87.8 % |
So for "is this one of my N people, or nobody" the model works; it just cannot tell strangers apart from each other. Strangers' max similarity to any enrolled centroid averages 0.17. Default threshold moved from 0.45 to 0.55 after this test (--strangers flag in eval_lfw.py).
Private test split (date-held-out, 10 classes = 9 identities + other)
The test set is small (59 faces, 49 of them family), so treat accuracy as +-2 faces, +-4 %. Agreement metrics are over 1500 faces and are more stable.
| Variant | Family acc | Open-set acc (THR 0.45, pre-retune) | Cosine vs fp32 | Argmax agree vs fp32 | Size | Latency (1 face, CPU) |
|---|---|---|---|---|---|---|
| v1 fp32 (PyTorch) | 0.980 | 0.898 | 1.000 | 1.000 | 3.9 MB | 22 ms |
| v1 fp32 ONNX | 0.980 | 0.898 | 1.000 | 1.000 | 3.9 MB | 59 ms |
| v1 INT8 ONNX (static, per-channel) | 0.959 | 0.881 | 0.996 | 0.991 | 1.1 MB | 82 ms |
| v1 INT4 (simulated, weight-only) | 0.939 | 0.881 | 0.976 | 0.989 | 0.5 MB | n/a |
| v2 fp32 (PyTorch) | 0.959 | 0.898 | 1.000 | 1.000 | 3.9 MB | 12 ms |
| v2 INT8 ONNX | 0.959 | 0.915 | 0.998 | 0.997 | 1.1 MB | 91 ms |
| v2 INT4 (simulated) | 0.939 | 0.864 | 0.978 | 0.992 | 0.5 MB | n/a |
Latency above was measured on a desktop CPU with ORT's default thread settings and is not a Pi number; the fp32 PyTorch numbers vs ORT differ mostly from thread configuration. Per-class breakdown is in analysis/quant_v1.json and analysis/quant_v2.json; a rendered report is analysis/quant_report.html.
- v1: clean crops, standard augmentation. Best on well-lit photos.
- v2: v1 recipe + low-light augmentation (random gamma darkening, contrast crush, sensor noise). Meant for indoor webcam use. Slightly lower on the clean test set.
- INT4 gives no kernel speed-up on CPU/Pi and costs 2-4 points, so ship INT8.
Architecture
TinyFaceNet(emb=128, width=1.6)
stem: 3x3/2 conv -> dw 3x3 (112 -> 56)
blocks: 10 depthwise-separable bottlenecks (expansion 2, PReLU, residual when stride 1)
channels 51 -> 76 -> 153 -> 204, strides at 56->28->14->7
head: 1x1 conv 409 -> dw 7x7 (global) -> Linear 409->128 -> BatchNorm1d
deploy params: 1,034,621
Training-only heads (removed from the published weights): distill: Linear(128, 512) mapping to teacher space and cls: Linear(128, n_classes).
Loss: CE(cls(e), y, label_smoothing=0.1) + 2.0 * (1 - cos(distill(e), teacher_emb)), AdamW 2e-3, OneCycle, 40 epochs, batch 128, sqrt-inverse-frequency class sampling, other capped at 600 train faces. Trained on CPU in under an hour (~5.3k crops).
Usage
PyTorch
import torch, torch.nn.functional as F, numpy as np
from PIL import Image
from model import TinyFaceNet # code/model.py
m = TinyFaceNet(); m.load_state_dict(torch.load("tinyfacenet_v1.pt", weights_only=True)["model"], strict=False); m.eval()
def prep(img): # img: 112x112 RGB, 5-point aligned crop (insightface norm_crop)
x = torch.from_numpy(np.asarray(img, np.float32)).permute(2,0,1)[None]
return (x/255. - 0.5)/0.5
with torch.no_grad():
e = F.normalize(m(prep(Image.open("face.jpg").convert("RGB").resize((112,112)))), dim=1) # [1,128]
ONNX Runtime (what runs on the Pi)
import onnxruntime as ort, numpy as np
s = ort.InferenceSession("onnx/tinyfacenet_v1_int8.onnx", providers=["CPUExecutionProvider"])
x = (face_rgb_112.transpose(2,0,1)[None].astype(np.float32)/255. - 0.5)/0.5
e = s.run(None, {"x": x})[0]; e /= np.linalg.norm(e, axis=1, keepdims=True)
Identification = nearest centroid + open-set threshold
# centroids.npz from code/enroll.py: names [N], centroids [N,128]
sim = centroids @ e[0]; j = sim.argmax()
label = names[j] if sim[j] >= 0.55 else "unknown" # 0.55: 2.8 % stranger accept, see stranger test
Input must be an aligned crop. The reference pipeline uses the insightface buffalo_l detector + norm_crop for 5-point alignment; any detector that gives 5 landmarks works. Unaligned crops will degrade accuracy sharply.
Files
tinyfacenet_v1.pt / .safetensors backbone state_dict (fp32), v1
tinyfacenet_v2.pt / .safetensors backbone state_dict (fp32), v2 low-light
onnx/tinyfacenet_{v1,v2}_{fp32,int8}.onnx opset 17, input "x" [B,3,112,112], output "emb" [B,128]
config.json architecture / preprocessing summary
code/model.py TinyFaceNet definition
code/build_dataset.py, rebuild_jpg.py aligned-crop dataset + teacher embeddings from a flat photo/video folder ($FACE_DATA)
code/train.py distillation training (DATA=... OUT=... THREADS=... python train.py 40)
code/quantize_eval.py ONNX export, ORT static INT8, simulated INT4, metrics -> analysis/quant_<tag>.json
code/make_report.py quant json -> html
code/enroll.py build centroids for your own people
code/eval_lfw.py LFW 10-fold benchmark + stranger/open-set threshold test (--teacher, --strangers)
code/live_test.py, live_ab.py webcam test (auto-gamma + CLAHE, 7-frame vote), INT8 vs INT4 side by side
code/labels.example.json cluster id -> identity variable template
analysis/REPORT.md data analysis + feasibility study (identities as categories only)
analysis/quant_*.json, lfw_model_v*.json, quant_report.html, train_*.log
Lessons learned (the expensive ones)
- Crop selection bug cost 20 points. v0 picked the largest detected face inside a padded sub-crop instead of the one matching the clustered box. Label noise dropped family accuracy to 77 %. Selecting by IoU with the original box fixed it (98 %).
- Split by date, not at random. Random split showed 99 %+ and was lying. Holding out whole months is the honest number.
- Distillation is the whole game at this data size. From-scratch on ~2k faces overfits to 85-92 %.
- "other" was the weak class, and LFW explained why. Only 10 negative faces in the test split hid it; 12k LFW strangers showed 6 % of them were accepted at THR 0.45. Retuned to 0.55 (2.8 %) at no recall cost. Always calibrate open-set thresholds on thousands of negatives, not tens.
- Run a public benchmark before believing a private number. 98 % on the family split and 62.6 % on LFW are both true; the first is a 10-way classifier, the second is what a face recogniser means.
- INT8 static (ORT, per-channel, 200 calibration images) is nearly free: ~1 point, 99 %+ argmax agreement, 3.5x smaller. INT4 is not worth it on CPU.
- Never lower process priority/affinity on an onnxruntime process. Spin-wait threads thrash and it stalls. Cap threads instead.
Limitations and intended use
- LFW 62.6 %. Trained on ~10 identities from one family's photos; the embedding does not generalise to unseen people. Use it for small closed-set / personal-archive identification with an open-set threshold (see stranger test), not as a general face recogniser or for 1:N search over strangers.
- Faces drift over time (age, glasses, hair); re-enroll (recompute centroids) periodically. No retraining needed.
- Intended for on-device, consent-based personal photo organisation. Not for surveillance or identifying people without their consent.
Privacy note
No images, face crops, embeddings of specific people, centroids, or names are included. The published weights are a generic backbone; identity-specific state lives only in files you generate locally with enroll.py.
Citation
@misc{tinyfacenet1m2026,
title = {TinyFaceNet-1M: ArcFace-distilled 1M-param face embedding for edge devices},
author = {sraivante},
year = {2026},
url = {https://huggingface.co/sraivante/TinyFaceNet-1M}
}
Teacher: insightface buffalo_l (Deng et al., ArcFace). Architecture family: MobileFaceNet (Chen et al., 2018).
- Downloads last month
- -