--- license: cc-by-nc-4.0 library_name: timm pipeline_tag: image-classification base_model: timm/vit_base_patch16_224.augreg_in21k_ft_in1k tags: - facial-expression-recognition - driver-monitoring - vision-transformer - parameter-efficient-fine-tuning - lora - adaptformer - ssf - pytorch metrics: - accuracy --- # Parameter-Efficient Adaptation of Facial Expression Recognition Models for Driver Monitoring Trained weights for the MSc thesis by **Ayush Tiwari**, BTU Cottbus-Senftenberg, 2026. [Code and training instructions](https://github.com/headless-start/parameter-efficient-dfer) · [Weight files](https://huggingface.co/headless-start/parameter-efficient-dfer/tree/main) · [Author profile](https://huggingface.co/headless-start) The code repository is currently a private preview; its link becomes accessible to everyone when the code is published. The weights here are public. A public ImageNet-pretrained ViT-B/16 is fine-tuned on six-class FER+ (Stage 1), then its encoder is adapted to KMU-FED with a fresh classifier (Stage 2). Five strategies are provided: full fine-tuning, linear probe, AdaptFormer, LoRA and SSF. ## Contents There are **51 checkpoints**: one selected FER+ model and ten held-out-fold models for each selected KMU-FED strategy. Per-fold results, predictions and search-selection summaries accompany the weights. No image data is included. | Path | Contents | |---|---| | `stage1/best.ckpt` | Complete FER+-adapted model | | `stage2//fold_00/` … `fold_09/` | Fold weights, results and predictions | | `stage2//selection.json` | Search scores and selected configuration | | `load_weights.py` | Loader for the accompanying code release | Full fine-tuning stores complete networks. The other Stage-2 strategies store trained parameters and their classifier; the loader combines them with the Stage-1 encoder. There is no separately trained all-data deployment model. ## Results FER+ six-class test: **93.21% accuracy**, **81.40% UAR**, **93.13% weighted F1**. The learning rate (1e-5) and epoch (28) were selected on the official validation split; the selected model was tested once. KMU-FED values below are **fold means**; accuracy ± is sample standard deviation across ten folds. The thesis and code README instead report **pooled weighted F1**, which differs from averaging fold F1 scores. | Strategy | Selected setting | Trainable parameters | Accuracy (%) | UAR (%) | Fold-mean weighted F1 (%) | |---|---|---:|---:|---:|---:| | Full fine-tuning | lr 3e-5 | 85,803,270 | 97.55 ± 4.11 | 96.87 | 97.25 | | Linear probe | lr 3e-3 | 4,614 | 93.27 ± 6.83 | 92.76 | 92.42 | | AdaptFormer | lr 1e-3, width 64 (reduction 12) | 1,194,246 | 98.45 ± 2.43 | 98.21 | 98.40 | | LoRA | lr 1e-3, rank 4 | 152,070 | 99.00 ± 2.59 | 98.56 | 98.83 | | SSF | lr 3e-3 | 210,438 | 97.91 ± 2.88 | 97.48 | 97.75 | Both the retained epoch and configuration were selected on the reported folds, so these are optimistic, best-observed estimates. Folds share drivers and sequences; they measure familiar-driver performance. Fold spread is descriptive, with one base seed (42) and deterministic fold-specific initialisation. The study does not establish equivalence, condition-specific robustness, the benefit of the intermediate FER+ stage, or inference speed. ## Use the weights Install the accompanying code release's requirements and run from its root directory. The loader needs its `methods.py`, `data.py` and `experiment.py`; the checkpoint files alone are not a Transformers pipeline. Download the helper at the verified weight revision: ```bash python - <<'PY' from huggingface_hub import hf_hub_download hf_hub_download( "headless-start/parameter-efficient-dfer", "load_weights.py", revision="92a70a21dc4810e346d7daf7b4e3c04b7dc721bb", local_dir="runs/released", ) PY ``` ```python import sys import torch from PIL import Image from data import build_transform from experiment import CLASS_NAMES sys.path.insert(0, "runs/released") from load_weights import load_stage2 model = load_stage2("lora", fold=0) with Image.open("face_crop.jpg") as image: x = build_transform(False, kmufed=True)(image.convert("RGB"))[None] with torch.inference_mode(): predicted = model(x).argmax(1).item() print(CLASS_NAMES[predicted]) ``` The loader downloads the required weights and returns an evaluation-mode model. Use `load_stage1()` for FER+ or `load_stage2(method, fold)` for KMU-FED; methods are `full_ft`, `linear_probe`, `adaptformer`, `lora`, `ssf`, and folds are zero-based (`0`–`9`). Class order is `angry, disgust, fear, happy, sad, surprise`. Inputs must be face crops; preprocessing makes three-channel grayscale, resizes to 224 × 224 with bicubic interpolation and normalises with mean/std 0.5. These original-format weights are for inference through the helper, not `--resume` inputs for the standalone training scripts. Hosted result JSONs also use the original format; the code release's plotter reads its own training outputs. The supplied loader retains adapter modules rather than merging them. ## Licence and citation Weights: **CC BY-NC 4.0**. The code release is MIT-licensed. Dataset and pretrained-weight terms remain with their providers. These are research models evaluated on a small driver corpus, without validation for operational use. ```bibtex @mastersthesis{tiwari2026peft_dfer, author = {Ayush Tiwari}, title = {Parameter-Efficient Adaptation of Facial Expression Recognition Models for Driver Monitoring}, school = {Brandenburgische Technische Universit{\"a}t Cottbus-Senftenberg}, year = {2026} } ```