File size: 5,638 Bytes
2f6b87a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
55e26a0
 
2f6b87a
55e26a0
 
 
2f6b87a
55e26a0
 
2f6b87a
55e26a0
 
 
 
2f6b87a
55e26a0
2f6b87a
55e26a0
 
 
2f6b87a
55e26a0
 
 
 
 
 
2f6b87a
55e26a0
 
 
2f6b87a
55e26a0
2f6b87a
55e26a0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2f6b87a
 
55e26a0
 
 
 
 
 
2f6b87a
55e26a0
 
2f6b87a
55e26a0
 
 
 
 
 
 
2f6b87a
55e26a0
 
 
 
 
 
2f6b87a
55e26a0
 
 
 
2f6b87a
55e26a0
2f6b87a
55e26a0
 
 
2f6b87a
 
 
 
55e26a0
2f6b87a
55e26a0
2f6b87a
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
---
license: cc-by-nc-4.0
library_name: timm
pipeline_tag: image-classification
base_model: timm/vit_base_patch16_224.augreg_in21k_ft_in1k
tags:
  - facial-expression-recognition
  - driver-monitoring
  - vision-transformer
  - parameter-efficient-fine-tuning
  - lora
  - adaptformer
  - ssf
  - pytorch
metrics:
  - accuracy
---

# Parameter-Efficient Adaptation of Facial Expression Recognition Models for Driver Monitoring

Trained weights for the MSc thesis by **Ayush Tiwari**, BTU
Cottbus-Senftenberg, 2026.

[Code and training instructions](https://github.com/headless-start/parameter-efficient-dfer)
· [Weight files](https://huggingface.co/headless-start/parameter-efficient-dfer/tree/main)
· [Author profile](https://huggingface.co/headless-start)

The code repository is currently a private preview; its link becomes accessible
to everyone when the code is published. The weights here are public.

A public ImageNet-pretrained ViT-B/16 is fine-tuned on six-class FER+ (Stage 1),
then its encoder is adapted to KMU-FED with a fresh classifier (Stage 2).
Five strategies are provided: full fine-tuning, linear probe, AdaptFormer,
LoRA and SSF.

## Contents

There are **51 checkpoints**: one selected FER+ model and ten held-out-fold
models for each selected KMU-FED strategy. Per-fold results, predictions and
search-selection summaries accompany the weights. No image data is included.

| Path | Contents |
|---|---|
| `stage1/best.ckpt` | Complete FER+-adapted model |
| `stage2/<strategy>/fold_00/` … `fold_09/` | Fold weights, results and predictions |
| `stage2/<strategy>/selection.json` | Search scores and selected configuration |
| `load_weights.py` | Loader for the accompanying code release |

Full fine-tuning stores complete networks. The other Stage-2 strategies store
trained parameters and their classifier; the loader combines them with the
Stage-1 encoder. There is no separately trained all-data deployment model.

## Results

FER+ six-class test: **93.21% accuracy**, **81.40% UAR**, **93.13% weighted F1**.
The learning rate (1e-5) and epoch (28) were selected on the official validation
split; the selected model was tested once.

KMU-FED values below are **fold means**; accuracy ± is sample standard deviation
across ten folds. The thesis and code README instead report **pooled weighted
F1**, which differs from averaging fold F1 scores.

| Strategy | Selected setting | Trainable parameters | Accuracy (%) | UAR (%) | Fold-mean weighted F1 (%) |
|---|---|---:|---:|---:|---:|
| Full fine-tuning | lr 3e-5 | 85,803,270 | 97.55 ± 4.11 | 96.87 | 97.25 |
| Linear probe | lr 3e-3 | 4,614 | 93.27 ± 6.83 | 92.76 | 92.42 |
| AdaptFormer | lr 1e-3, width 64 (reduction 12) | 1,194,246 | 98.45 ± 2.43 | 98.21 | 98.40 |
| LoRA | lr 1e-3, rank 4 | 152,070 | 99.00 ± 2.59 | 98.56 | 98.83 |
| SSF | lr 3e-3 | 210,438 | 97.91 ± 2.88 | 97.48 | 97.75 |

Both the retained epoch and configuration were selected on the reported folds,
so these are optimistic, best-observed estimates. Folds share drivers and
sequences; they measure familiar-driver performance. Fold spread is descriptive,
with one base seed (42) and deterministic fold-specific initialisation. The study
does not establish equivalence, condition-specific robustness, the benefit of
the intermediate FER+ stage, or inference speed.

## Use the weights

Install the accompanying code release's requirements and run from its root
directory. The loader needs its `methods.py`, `data.py` and `experiment.py`;
the checkpoint files alone are not a Transformers pipeline.

Download the helper at the verified weight revision:

```bash
python - <<'PY'
from huggingface_hub import hf_hub_download
hf_hub_download(
    "headless-start/parameter-efficient-dfer", "load_weights.py",
    revision="92a70a21dc4810e346d7daf7b4e3c04b7dc721bb",
    local_dir="runs/released",
)
PY
```

```python
import sys
import torch
from PIL import Image
from data import build_transform
from experiment import CLASS_NAMES

sys.path.insert(0, "runs/released")
from load_weights import load_stage2

model = load_stage2("lora", fold=0)
with Image.open("face_crop.jpg") as image:
    x = build_transform(False, kmufed=True)(image.convert("RGB"))[None]
with torch.inference_mode():
    predicted = model(x).argmax(1).item()
print(CLASS_NAMES[predicted])
```

The loader downloads the required weights and returns an evaluation-mode model.
Use `load_stage1()` for FER+ or `load_stage2(method, fold)` for KMU-FED;
methods are `full_ft`, `linear_probe`, `adaptformer`, `lora`, `ssf`, and folds
are zero-based (`0`–`9`). Class order is `angry, disgust, fear, happy, sad,
surprise`. Inputs must be face crops; preprocessing makes three-channel grayscale,
resizes to 224 × 224 with bicubic interpolation and normalises with mean/std 0.5.

These original-format weights are for inference through the helper, not
`--resume` inputs for the standalone training scripts. Hosted result JSONs also
use the original format; the code release's plotter reads its own training
outputs. The supplied loader retains adapter modules rather than merging them.

## Licence and citation

Weights: **CC BY-NC 4.0**. The code release is MIT-licensed. Dataset and
pretrained-weight terms remain with their providers. These are research models
evaluated on a small driver corpus, without validation for operational use.

```bibtex
@mastersthesis{tiwari2026peft_dfer,
  author = {Ayush Tiwari},
  title = {Parameter-Efficient Adaptation of Facial Expression Recognition Models for Driver Monitoring},
  school = {Brandenburgische Technische Universit{\"a}t Cottbus-Senftenberg},
  year = {2026}
}
```