Image Classification
timm
PyTorch
facial-expression-recognition
driver-monitoring
vision-transformer
parameter-efficient-fine-tuning
lora
adaptformer
ssf
Instructions to use headless-start/parameter-efficient-dfer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- timm
How to use headless-start/parameter-efficient-dfer with timm:
import timm model = timm.create_model("hf_hub:headless-start/parameter-efficient-dfer", pretrained=True) - Notebooks
- Google Colab
- Kaggle
File size: 5,638 Bytes
2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a 55e26a0 2f6b87a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 | ---
license: cc-by-nc-4.0
library_name: timm
pipeline_tag: image-classification
base_model: timm/vit_base_patch16_224.augreg_in21k_ft_in1k
tags:
- facial-expression-recognition
- driver-monitoring
- vision-transformer
- parameter-efficient-fine-tuning
- lora
- adaptformer
- ssf
- pytorch
metrics:
- accuracy
---
# Parameter-Efficient Adaptation of Facial Expression Recognition Models for Driver Monitoring
Trained weights for the MSc thesis by **Ayush Tiwari**, BTU
Cottbus-Senftenberg, 2026.
[Code and training instructions](https://github.com/headless-start/parameter-efficient-dfer)
· [Weight files](https://huggingface.co/headless-start/parameter-efficient-dfer/tree/main)
· [Author profile](https://huggingface.co/headless-start)
The code repository is currently a private preview; its link becomes accessible
to everyone when the code is published. The weights here are public.
A public ImageNet-pretrained ViT-B/16 is fine-tuned on six-class FER+ (Stage 1),
then its encoder is adapted to KMU-FED with a fresh classifier (Stage 2).
Five strategies are provided: full fine-tuning, linear probe, AdaptFormer,
LoRA and SSF.
## Contents
There are **51 checkpoints**: one selected FER+ model and ten held-out-fold
models for each selected KMU-FED strategy. Per-fold results, predictions and
search-selection summaries accompany the weights. No image data is included.
| Path | Contents |
|---|---|
| `stage1/best.ckpt` | Complete FER+-adapted model |
| `stage2/<strategy>/fold_00/` … `fold_09/` | Fold weights, results and predictions |
| `stage2/<strategy>/selection.json` | Search scores and selected configuration |
| `load_weights.py` | Loader for the accompanying code release |
Full fine-tuning stores complete networks. The other Stage-2 strategies store
trained parameters and their classifier; the loader combines them with the
Stage-1 encoder. There is no separately trained all-data deployment model.
## Results
FER+ six-class test: **93.21% accuracy**, **81.40% UAR**, **93.13% weighted F1**.
The learning rate (1e-5) and epoch (28) were selected on the official validation
split; the selected model was tested once.
KMU-FED values below are **fold means**; accuracy ± is sample standard deviation
across ten folds. The thesis and code README instead report **pooled weighted
F1**, which differs from averaging fold F1 scores.
| Strategy | Selected setting | Trainable parameters | Accuracy (%) | UAR (%) | Fold-mean weighted F1 (%) |
|---|---|---:|---:|---:|---:|
| Full fine-tuning | lr 3e-5 | 85,803,270 | 97.55 ± 4.11 | 96.87 | 97.25 |
| Linear probe | lr 3e-3 | 4,614 | 93.27 ± 6.83 | 92.76 | 92.42 |
| AdaptFormer | lr 1e-3, width 64 (reduction 12) | 1,194,246 | 98.45 ± 2.43 | 98.21 | 98.40 |
| LoRA | lr 1e-3, rank 4 | 152,070 | 99.00 ± 2.59 | 98.56 | 98.83 |
| SSF | lr 3e-3 | 210,438 | 97.91 ± 2.88 | 97.48 | 97.75 |
Both the retained epoch and configuration were selected on the reported folds,
so these are optimistic, best-observed estimates. Folds share drivers and
sequences; they measure familiar-driver performance. Fold spread is descriptive,
with one base seed (42) and deterministic fold-specific initialisation. The study
does not establish equivalence, condition-specific robustness, the benefit of
the intermediate FER+ stage, or inference speed.
## Use the weights
Install the accompanying code release's requirements and run from its root
directory. The loader needs its `methods.py`, `data.py` and `experiment.py`;
the checkpoint files alone are not a Transformers pipeline.
Download the helper at the verified weight revision:
```bash
python - <<'PY'
from huggingface_hub import hf_hub_download
hf_hub_download(
"headless-start/parameter-efficient-dfer", "load_weights.py",
revision="92a70a21dc4810e346d7daf7b4e3c04b7dc721bb",
local_dir="runs/released",
)
PY
```
```python
import sys
import torch
from PIL import Image
from data import build_transform
from experiment import CLASS_NAMES
sys.path.insert(0, "runs/released")
from load_weights import load_stage2
model = load_stage2("lora", fold=0)
with Image.open("face_crop.jpg") as image:
x = build_transform(False, kmufed=True)(image.convert("RGB"))[None]
with torch.inference_mode():
predicted = model(x).argmax(1).item()
print(CLASS_NAMES[predicted])
```
The loader downloads the required weights and returns an evaluation-mode model.
Use `load_stage1()` for FER+ or `load_stage2(method, fold)` for KMU-FED;
methods are `full_ft`, `linear_probe`, `adaptformer`, `lora`, `ssf`, and folds
are zero-based (`0`–`9`). Class order is `angry, disgust, fear, happy, sad,
surprise`. Inputs must be face crops; preprocessing makes three-channel grayscale,
resizes to 224 × 224 with bicubic interpolation and normalises with mean/std 0.5.
These original-format weights are for inference through the helper, not
`--resume` inputs for the standalone training scripts. Hosted result JSONs also
use the original format; the code release's plotter reads its own training
outputs. The supplied loader retains adapter modules rather than merging them.
## Licence and citation
Weights: **CC BY-NC 4.0**. The code release is MIT-licensed. Dataset and
pretrained-weight terms remain with their providers. These are research models
evaluated on a small driver corpus, without validation for operational use.
```bibtex
@mastersthesis{tiwari2026peft_dfer,
author = {Ayush Tiwari},
title = {Parameter-Efficient Adaptation of Facial Expression Recognition Models for Driver Monitoring},
school = {Brandenburgische Technische Universit{\"a}t Cottbus-Senftenberg},
year = {2026}
}
```
|