TinyFaceNet-1M: a 1.0M-param face embedding net distilled from ArcFace, for edge devices

downloads (30d) downloads (all time) params int8

A MobileFaceNet-style CNN (112x112 RGB in, 128-d L2-normalised embedding out) trained by knowledge distillation from insightface buffalo_l (ArcFace, 65M params) on a small private photo archive, then quantised to INT8 with ONNX Runtime. Built as the recognition core of an offline family-photo organiser that runs on a Raspberry Pi 5.

This repo is an end-to-end recipe, not just weights: data pipeline, training, quantisation study, webcam test and the lessons learned are all here.

βœ… What this model is for

Closed-set identification of a small, enrolled group of people on an edge device. Enroll N people (family, household, a small team: ~5-20 faces each with code/enroll.py), then for every detected face answer "which of my N people is this, or nobody?"

Task Fit Evidence
Tag family members in a personal photo/video archive (offline, on a Pi) βœ… Primary use case 94-98 % on held-out dates, 9 identities
Reject strangers as "unknown" βœ… Works with THR 0.55 97.2 % of 12k unseen faces rejected
Smart-doorbell / home-cam "is this a household member?" βœ… Same problem low-light v2 + 7-frame vote in live_test.py
Run anywhere with ~1 MB and no GPU βœ… INT8 ONNX, 1.03 M params, 3 ms/face desktop CPU
Recipe to distil your own tiny face net from ArcFace βœ… Everything is here code/train.py, quantize_eval.py
General face recognition / verification of unknown people ❌ Not this model LFW 62.6 % (chance 50 %)
1:N search over large galleries, surveillance, identity verification ❌ Do not use cannot separate strangers from each other

Rule of thumb: if every person you care about can be enrolled in advance, this works. If the model has to compare two people it has never seen, it does not.

πŸ–₯️ Where it runs

Embedder cost: 171 M MACs per face, 1.1 MB INT8 weights, 312 KB peak activation (INT8). Only the desktop row is measured; the rest are estimates from MACs and public per-core throughput, so expect Β±2x. The face detector, not this embedder, is the bottleneck on small boards: pair it with a tiny detector (BlazeFace ~0.1 M, SCRFD-500M, Ultra-Light-Fast ~1 MB) rather than the buffalo_l det_10g used for training.

Device Runs? Est. embed time / face (INT8) Notes
Raspberry Pi 5 (Cortex-A76 x4) βœ… Target device ~10-20 ms onnxruntime-arm64, 4 threads; whole pipeline (tiny detector + align + embed) ~5-10 fps
Raspberry Pi 4 / 400 / CM4 (Cortex-A72 x4) βœ… ~30-60 ms same stack, ~2-3x slower than Pi 5
Raspberry Pi 3 / 3B+ (Cortex-A53 x4) βœ… ~80-150 ms fine for photo-archive batch jobs, slow for live video
Raspberry Pi Zero 2 W (Cortex-A53 x4, 512 MB) βœ… ~100-200 ms 512 MB is enough for ORT + INT8 model; use a 160 px detector
Raspberry Pi Zero / Zero W (ARMv6 single core) ⚠️ Barely ~2-5 s works, but only for offline batch tagging
NVIDIA Jetson Orin Nano / Nano βœ… Overkill <2 ms TensorRT or ORT-CUDA; run the full buffalo_l detector here
Google Coral Edge TPU βœ… ~1-3 ms needs ONNX β†’ TFLite INT8 conversion, then the edgetpu compiler
Rockchip RK3566/RK3588 (Orange Pi, Radxa), Intel N100 mini-PC βœ… ~5-40 ms ORT CPU; RK3588 NPU via RKNN also possible
Android / iOS phone βœ… ~5-15 ms ORT Mobile / NNAPI / CoreML from the ONNX file
ESP32-S3 (240 MHz, 8 MB PSRAM) ⚠️ Experimental ~3-8 s fits in PSRAM via esp-dl / TFLite-Micro; no room for a detector, only crop-and-embed
Raspberry Pi Pico 2 / Pico 2 W (RP2350, 520 KB SRAM, 4 MB flash) ❌ Not practical ~5-15 s weights fit flash, activations barely fit SRAM, no SIMD; and no detector
Raspberry Pi Pico / Pico W (RP2040, 264 KB SRAM, 2 MB flash) ❌ No β€” 312 KB activations do not fit; use the Pico only as a camera/IO companion to a Pi Zero 2
Arduino (AVR, Cortex-M0/M4 boards) ❌ No β€” orders of magnitude short on RAM and compute

Minimum sensible target: Pi Zero 2 W. Comfortable target: Pi 4. Sweet spot: Pi 5. Microcontrollers (Pico, ESP32) should capture frames and hand them to a Linux board.

What is published, what is not

Published Not published
Backbone weights (.pt, .safetensors, ONNX fp32 + INT8) Any photo, video, or face crop
All code (code/) Identity centroids / classifier head (biometric templates)
Analysis & quantisation reports Names of people: everywhere they are referred to as PERSON_A ... PERSON_I + other
Training logs (identities as variables) The labels.json cluster-to-name map (an example template is included)

The checkpoints contain only the embedding backbone. The training-time distillation head and the 10-way classifier head were removed before upload, so the model cannot by itself name anyone. You enroll your own people with code/enroll.py.

Results

Public benchmark: LFW (6000 pairs, standard 10-fold protocol)

Model LFW acc AUC TAR@FAR 1e-2 mean cos: same / different person
buffalo_l teacher (reference) 99.85 % 0.999 0.998 0.67 / 0.00
TinyFaceNet v1 62.6 % 0.672 0.072 0.76 / 0.64
TinyFaceNet v2 63.6 % 0.685 0.078 0.79 / 0.69

Chance is 50 %. This model is not a general face recogniser. Distilled on ~5k faces of 10 identities, the embedding only learned to separate those identities; two strangers land at cosine 0.64 on average (teacher: 0.00). The teacher scoring 99.85 % through the same pipeline confirms the evaluation itself is sound. Reproduce with code/eval_lfw.py (LFW via scikit-learn, re-aligned with buffalo_l; numbers in analysis/lfw_model_v*.json).

The fix is data, not architecture: run the same train.py recipe on 50-100k public unlabeled faces (only teacher embeddings are needed, no labels). MobileFaceNet-class nets reach 99.5 % LFW when trained on MS1M.

Stranger test (what the closed-set use case actually needs)

All 12,000 LFW faces treated as people outside the enrolled set, scored against the 9 enrolled centroids (v1):

Open-set THR Stranger accepted as an enrolled person Enrolled faces still recognised (49 held-out)
0.45 6.2 % 93.9 %
0.55 (default) 2.8 % 93.9 %
0.60 1.7 % 93.9 %
0.65 1.0 % 87.8 %

So for "is this one of my N people, or nobody" the model works; it just cannot tell strangers apart from each other. Strangers' max similarity to any enrolled centroid averages 0.17. Default threshold moved from 0.45 to 0.55 after this test (--strangers flag in eval_lfw.py).

Private test split (date-held-out, 10 classes = 9 identities + other)

The test set is small (59 faces, 49 of them family), so treat accuracy as +-2 faces, +-4 %. Agreement metrics are over 1500 faces and are more stable.

Variant Family acc Open-set acc (THR 0.45, pre-retune) Cosine vs fp32 Argmax agree vs fp32 Size Latency (1 face, CPU)
v1 fp32 (PyTorch) 0.980 0.898 1.000 1.000 3.9 MB 22 ms
v1 fp32 ONNX 0.980 0.898 1.000 1.000 3.9 MB 59 ms
v1 INT8 ONNX (static, per-channel) 0.959 0.881 0.996 0.991 1.1 MB 82 ms
v1 INT4 (simulated, weight-only) 0.939 0.881 0.976 0.989 0.5 MB n/a
v2 fp32 (PyTorch) 0.959 0.898 1.000 1.000 3.9 MB 12 ms
v2 INT8 ONNX 0.959 0.915 0.998 0.997 1.1 MB 91 ms
v2 INT4 (simulated) 0.939 0.864 0.978 0.992 0.5 MB n/a

Latency above was measured on a desktop CPU with ORT's default thread settings and is not a Pi number; the fp32 PyTorch numbers vs ORT differ mostly from thread configuration. Per-class breakdown is in analysis/quant_v1.json and analysis/quant_v2.json; a rendered report is analysis/quant_report.html.

  • v1: clean crops, standard augmentation. Best on well-lit photos.
  • v2: v1 recipe + low-light augmentation (random gamma darkening, contrast crush, sensor noise). Meant for indoor webcam use. Slightly lower on the clean test set.
  • INT4 gives no kernel speed-up on CPU/Pi and costs 2-4 points, so ship INT8.

Architecture

TinyFaceNet(emb=128, width=1.6)
  stem:   3x3/2 conv -> dw 3x3           (112 -> 56)
  blocks: 10 depthwise-separable bottlenecks (expansion 2, PReLU, residual when stride 1)
          channels 51 -> 76 -> 153 -> 204, strides at 56->28->14->7
  head:   1x1 conv 409 -> dw 7x7 (global) -> Linear 409->128 -> BatchNorm1d
  deploy params: 1,034,621

Training-only heads (removed from the published weights): distill: Linear(128, 512) mapping to teacher space and cls: Linear(128, n_classes).

Loss: CE(cls(e), y, label_smoothing=0.1) + 2.0 * (1 - cos(distill(e), teacher_emb)), AdamW 2e-3, OneCycle, 40 epochs, batch 128, sqrt-inverse-frequency class sampling, other capped at 600 train faces. Trained on CPU in under an hour (~5.3k crops).

Usage

PyTorch

import torch, torch.nn.functional as F, numpy as np
from PIL import Image
from model import TinyFaceNet           # code/model.py

m = TinyFaceNet(); m.load_state_dict(torch.load("tinyfacenet_v1.pt", weights_only=True)["model"], strict=False); m.eval()
def prep(img):                          # img: 112x112 RGB, 5-point aligned crop (insightface norm_crop)
    x = torch.from_numpy(np.asarray(img, np.float32)).permute(2,0,1)[None]
    return (x/255. - 0.5)/0.5
with torch.no_grad():
    e = F.normalize(m(prep(Image.open("face.jpg").convert("RGB").resize((112,112)))), dim=1)   # [1,128]

ONNX Runtime (what runs on the Pi)

import onnxruntime as ort, numpy as np
s = ort.InferenceSession("onnx/tinyfacenet_v1_int8.onnx", providers=["CPUExecutionProvider"])
x = (face_rgb_112.transpose(2,0,1)[None].astype(np.float32)/255. - 0.5)/0.5
e = s.run(None, {"x": x})[0]; e /= np.linalg.norm(e, axis=1, keepdims=True)

Identification = nearest centroid + open-set threshold

# centroids.npz from code/enroll.py: names [N], centroids [N,128]
sim = centroids @ e[0]; j = sim.argmax()
label = names[j] if sim[j] >= 0.55 else "unknown"   # 0.55: 2.8 % stranger accept, see stranger test

Input must be an aligned crop. The reference pipeline uses the insightface buffalo_l detector + norm_crop for 5-point alignment; any detector that gives 5 landmarks works. Unaligned crops will degrade accuracy sharply.

Files

tinyfacenet_v1.pt / .safetensors      backbone state_dict (fp32), v1
tinyfacenet_v2.pt / .safetensors      backbone state_dict (fp32), v2 low-light
onnx/tinyfacenet_{v1,v2}_{fp32,int8}.onnx   opset 17, input "x" [B,3,112,112], output "emb" [B,128]
config.json                           architecture / preprocessing summary
code/model.py                         TinyFaceNet definition
code/build_dataset.py, rebuild_jpg.py aligned-crop dataset + teacher embeddings from a flat photo/video folder ($FACE_DATA)
code/train.py                         distillation training (DATA=... OUT=... THREADS=... python train.py 40)
code/quantize_eval.py                 ONNX export, ORT static INT8, simulated INT4, metrics -> analysis/quant_<tag>.json
code/make_report.py                   quant json -> html
code/enroll.py                        build centroids for your own people
code/eval_lfw.py                      LFW 10-fold benchmark + stranger/open-set threshold test (--teacher, --strangers)
code/live_test.py, live_ab.py         webcam test (auto-gamma + CLAHE, 7-frame vote), INT8 vs INT4 side by side
code/labels.example.json              cluster id -> identity variable template
analysis/REPORT.md                    data analysis + feasibility study (identities as categories only)
analysis/quant_*.json, lfw_model_v*.json, quant_report.html, train_*.log

Lessons learned (the expensive ones)

  1. Crop selection bug cost 20 points. v0 picked the largest detected face inside a padded sub-crop instead of the one matching the clustered box. Label noise dropped family accuracy to 77 %. Selecting by IoU with the original box fixed it (98 %).
  2. Split by date, not at random. Random split showed 99 %+ and was lying. Holding out whole months is the honest number.
  3. Distillation is the whole game at this data size. From-scratch on ~2k faces overfits to 85-92 %.
  4. "other" was the weak class, and LFW explained why. Only 10 negative faces in the test split hid it; 12k LFW strangers showed 6 % of them were accepted at THR 0.45. Retuned to 0.55 (2.8 %) at no recall cost. Always calibrate open-set thresholds on thousands of negatives, not tens.
  5. Run a public benchmark before believing a private number. 98 % on the family split and 62.6 % on LFW are both true; the first is a 10-way classifier, the second is what a face recogniser means.
  6. INT8 static (ORT, per-channel, 200 calibration images) is nearly free: ~1 point, 99 %+ argmax agreement, 3.5x smaller. INT4 is not worth it on CPU.
  7. Never lower process priority/affinity on an onnxruntime process. Spin-wait threads thrash and it stalls. Cap threads instead.

Limitations and intended use

  • LFW 62.6 %. Trained on ~10 identities from one family's photos; the embedding does not generalise to unseen people. Use it for small closed-set / personal-archive identification with an open-set threshold (see stranger test), not as a general face recogniser or for 1:N search over strangers.
  • Faces drift over time (age, glasses, hair); re-enroll (recompute centroids) periodically. No retraining needed.
  • Intended for on-device, consent-based personal photo organisation. Not for surveillance or identifying people without their consent.

Privacy note

No images, face crops, embeddings of specific people, centroids, or names are included. The published weights are a generic backbone; identity-specific state lives only in files you generate locally with enroll.py.

Citation

@misc{tinyfacenet1m2026,
  title  = {TinyFaceNet-1M: ArcFace-distilled 1M-param face embedding for edge devices},
  author = {sraivante},
  year   = {2026},
  url    = {https://huggingface.co/sraivante/TinyFaceNet-1M}
}

Teacher: insightface buffalo_l (Deng et al., ArcFace). Architecture family: MobileFaceNet (Chen et al., 2018).

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support