headless-start's picture
Update model card
55e26a0 verified
|
Raw
History Blame Contribute Delete
5.64 kB
metadata
license: cc-by-nc-4.0
library_name: timm
pipeline_tag: image-classification
base_model: timm/vit_base_patch16_224.augreg_in21k_ft_in1k
tags:
  - facial-expression-recognition
  - driver-monitoring
  - vision-transformer
  - parameter-efficient-fine-tuning
  - lora
  - adaptformer
  - ssf
  - pytorch
metrics:
  - accuracy

Parameter-Efficient Adaptation of Facial Expression Recognition Models for Driver Monitoring

Trained weights for the MSc thesis by Ayush Tiwari, BTU Cottbus-Senftenberg, 2026.

Code and training instructions · Weight files · Author profile

The code repository is currently a private preview; its link becomes accessible to everyone when the code is published. The weights here are public.

A public ImageNet-pretrained ViT-B/16 is fine-tuned on six-class FER+ (Stage 1), then its encoder is adapted to KMU-FED with a fresh classifier (Stage 2). Five strategies are provided: full fine-tuning, linear probe, AdaptFormer, LoRA and SSF.

Contents

There are 51 checkpoints: one selected FER+ model and ten held-out-fold models for each selected KMU-FED strategy. Per-fold results, predictions and search-selection summaries accompany the weights. No image data is included.

Path Contents
stage1/best.ckpt Complete FER+-adapted model
stage2/<strategy>/fold_00/ … fold_09/ Fold weights, results and predictions
stage2/<strategy>/selection.json Search scores and selected configuration
load_weights.py Loader for the accompanying code release

Full fine-tuning stores complete networks. The other Stage-2 strategies store trained parameters and their classifier; the loader combines them with the Stage-1 encoder. There is no separately trained all-data deployment model.

Results

FER+ six-class test: 93.21% accuracy, 81.40% UAR, 93.13% weighted F1. The learning rate (1e-5) and epoch (28) were selected on the official validation split; the selected model was tested once.

KMU-FED values below are fold means; accuracy ± is sample standard deviation across ten folds. The thesis and code README instead report pooled weighted F1, which differs from averaging fold F1 scores.

Strategy Selected setting Trainable parameters Accuracy (%) UAR (%) Fold-mean weighted F1 (%)
Full fine-tuning lr 3e-5 85,803,270 97.55 ± 4.11 96.87 97.25
Linear probe lr 3e-3 4,614 93.27 ± 6.83 92.76 92.42
AdaptFormer lr 1e-3, width 64 (reduction 12) 1,194,246 98.45 ± 2.43 98.21 98.40
LoRA lr 1e-3, rank 4 152,070 99.00 ± 2.59 98.56 98.83
SSF lr 3e-3 210,438 97.91 ± 2.88 97.48 97.75

Both the retained epoch and configuration were selected on the reported folds, so these are optimistic, best-observed estimates. Folds share drivers and sequences; they measure familiar-driver performance. Fold spread is descriptive, with one base seed (42) and deterministic fold-specific initialisation. The study does not establish equivalence, condition-specific robustness, the benefit of the intermediate FER+ stage, or inference speed.

Use the weights

Install the accompanying code release's requirements and run from its root directory. The loader needs its methods.py, data.py and experiment.py; the checkpoint files alone are not a Transformers pipeline.

Download the helper at the verified weight revision:

python - <<'PY'
from huggingface_hub import hf_hub_download
hf_hub_download(
    "headless-start/parameter-efficient-dfer", "load_weights.py",
    revision="92a70a21dc4810e346d7daf7b4e3c04b7dc721bb",
    local_dir="runs/released",
)
PY
import sys
import torch
from PIL import Image
from data import build_transform
from experiment import CLASS_NAMES

sys.path.insert(0, "runs/released")
from load_weights import load_stage2

model = load_stage2("lora", fold=0)
with Image.open("face_crop.jpg") as image:
    x = build_transform(False, kmufed=True)(image.convert("RGB"))[None]
with torch.inference_mode():
    predicted = model(x).argmax(1).item()
print(CLASS_NAMES[predicted])

The loader downloads the required weights and returns an evaluation-mode model. Use load_stage1() for FER+ or load_stage2(method, fold) for KMU-FED; methods are full_ft, linear_probe, adaptformer, lora, ssf, and folds are zero-based (0–9). Class order is angry, disgust, fear, happy, sad, surprise. Inputs must be face crops; preprocessing makes three-channel grayscale, resizes to 224 × 224 with bicubic interpolation and normalises with mean/std 0.5.

These original-format weights are for inference through the helper, not --resume inputs for the standalone training scripts. Hosted result JSONs also use the original format; the code release's plotter reads its own training outputs. The supplied loader retains adapter modules rather than merging them.

Licence and citation

Weights: CC BY-NC 4.0. The code release is MIT-licensed. Dataset and pretrained-weight terms remain with their providers. These are research models evaluated on a small driver corpus, without validation for operational use.

@mastersthesis{tiwari2026peft_dfer,
  author = {Ayush Tiwari},
  title = {Parameter-Efficient Adaptation of Facial Expression Recognition Models for Driver Monitoring},
  school = {Brandenburgische Technische Universit{\"a}t Cottbus-Senftenberg},
  year = {2026}
}