headless-start's picture
Update model card
55e26a0 verified
|
Raw
History Blame Contribute Delete
5.64 kB
---
license: cc-by-nc-4.0
library_name: timm
pipeline_tag: image-classification
base_model: timm/vit_base_patch16_224.augreg_in21k_ft_in1k
tags:
- facial-expression-recognition
- driver-monitoring
- vision-transformer
- parameter-efficient-fine-tuning
- lora
- adaptformer
- ssf
- pytorch
metrics:
- accuracy
---
# Parameter-Efficient Adaptation of Facial Expression Recognition Models for Driver Monitoring
Trained weights for the MSc thesis by **Ayush Tiwari**, BTU
Cottbus-Senftenberg, 2026.
[Code and training instructions](https://github.com/headless-start/parameter-efficient-dfer)
· [Weight files](https://huggingface.co/headless-start/parameter-efficient-dfer/tree/main)
· [Author profile](https://huggingface.co/headless-start)
The code repository is currently a private preview; its link becomes accessible
to everyone when the code is published. The weights here are public.
A public ImageNet-pretrained ViT-B/16 is fine-tuned on six-class FER+ (Stage 1),
then its encoder is adapted to KMU-FED with a fresh classifier (Stage 2).
Five strategies are provided: full fine-tuning, linear probe, AdaptFormer,
LoRA and SSF.
## Contents
There are **51 checkpoints**: one selected FER+ model and ten held-out-fold
models for each selected KMU-FED strategy. Per-fold results, predictions and
search-selection summaries accompany the weights. No image data is included.
| Path | Contents |
|---|---|
| `stage1/best.ckpt` | Complete FER+-adapted model |
| `stage2/<strategy>/fold_00/` … `fold_09/` | Fold weights, results and predictions |
| `stage2/<strategy>/selection.json` | Search scores and selected configuration |
| `load_weights.py` | Loader for the accompanying code release |
Full fine-tuning stores complete networks. The other Stage-2 strategies store
trained parameters and their classifier; the loader combines them with the
Stage-1 encoder. There is no separately trained all-data deployment model.
## Results
FER+ six-class test: **93.21% accuracy**, **81.40% UAR**, **93.13% weighted F1**.
The learning rate (1e-5) and epoch (28) were selected on the official validation
split; the selected model was tested once.
KMU-FED values below are **fold means**; accuracy ± is sample standard deviation
across ten folds. The thesis and code README instead report **pooled weighted
F1**, which differs from averaging fold F1 scores.
| Strategy | Selected setting | Trainable parameters | Accuracy (%) | UAR (%) | Fold-mean weighted F1 (%) |
|---|---|---:|---:|---:|---:|
| Full fine-tuning | lr 3e-5 | 85,803,270 | 97.55 ± 4.11 | 96.87 | 97.25 |
| Linear probe | lr 3e-3 | 4,614 | 93.27 ± 6.83 | 92.76 | 92.42 |
| AdaptFormer | lr 1e-3, width 64 (reduction 12) | 1,194,246 | 98.45 ± 2.43 | 98.21 | 98.40 |
| LoRA | lr 1e-3, rank 4 | 152,070 | 99.00 ± 2.59 | 98.56 | 98.83 |
| SSF | lr 3e-3 | 210,438 | 97.91 ± 2.88 | 97.48 | 97.75 |
Both the retained epoch and configuration were selected on the reported folds,
so these are optimistic, best-observed estimates. Folds share drivers and
sequences; they measure familiar-driver performance. Fold spread is descriptive,
with one base seed (42) and deterministic fold-specific initialisation. The study
does not establish equivalence, condition-specific robustness, the benefit of
the intermediate FER+ stage, or inference speed.
## Use the weights
Install the accompanying code release's requirements and run from its root
directory. The loader needs its `methods.py`, `data.py` and `experiment.py`;
the checkpoint files alone are not a Transformers pipeline.
Download the helper at the verified weight revision:
```bash
python - <<'PY'
from huggingface_hub import hf_hub_download
hf_hub_download(
"headless-start/parameter-efficient-dfer", "load_weights.py",
revision="92a70a21dc4810e346d7daf7b4e3c04b7dc721bb",
local_dir="runs/released",
)
PY
```
```python
import sys
import torch
from PIL import Image
from data import build_transform
from experiment import CLASS_NAMES
sys.path.insert(0, "runs/released")
from load_weights import load_stage2
model = load_stage2("lora", fold=0)
with Image.open("face_crop.jpg") as image:
x = build_transform(False, kmufed=True)(image.convert("RGB"))[None]
with torch.inference_mode():
predicted = model(x).argmax(1).item()
print(CLASS_NAMES[predicted])
```
The loader downloads the required weights and returns an evaluation-mode model.
Use `load_stage1()` for FER+ or `load_stage2(method, fold)` for KMU-FED;
methods are `full_ft`, `linear_probe`, `adaptformer`, `lora`, `ssf`, and folds
are zero-based (`0`–`9`). Class order is `angry, disgust, fear, happy, sad,
surprise`. Inputs must be face crops; preprocessing makes three-channel grayscale,
resizes to 224 × 224 with bicubic interpolation and normalises with mean/std 0.5.
These original-format weights are for inference through the helper, not
`--resume` inputs for the standalone training scripts. Hosted result JSONs also
use the original format; the code release's plotter reads its own training
outputs. The supplied loader retains adapter modules rather than merging them.
## Licence and citation
Weights: **CC BY-NC 4.0**. The code release is MIT-licensed. Dataset and
pretrained-weight terms remain with their providers. These are research models
evaluated on a small driver corpus, without validation for operational use.
```bibtex
@mastersthesis{tiwari2026peft_dfer,
author = {Ayush Tiwari},
title = {Parameter-Efficient Adaptation of Facial Expression Recognition Models for Driver Monitoring},
school = {Brandenburgische Technische Universit{\"a}t Cottbus-Senftenberg},
year = {2026}
}
```