File size: 4,189 Bytes
a6e22df 248459a a6e22df 248459a a6e22df | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 | ---
license: apache-2.0
tags:
- audio-deepfake-detection
- anti-spoofing
- self-supervised-learning
- speech
- benchmark
library_name: pytorch
pretty_name: Spoof-SUPERB Detector Checkpoints
---
# Spoof-SUPERB: Detector Checkpoints
Trained weights for every detector in **Spoof-SUPERB**, a benchmark of
self-supervised speech representations for audio deepfake detection.
- Code: https://github.com/issflab/spoof_SUPERB
- Score files: https://huggingface.co/datasets/issf/spoof-superb-scores
- Leaderboard: https://huggingface.co/spaces/issf/Spoof-SUPERB
## These are downstream weights only
Spoof-SUPERB evaluates SSL front-ends **frozen**. Only three things are trained:
the learnable layer-weighting over the upstream's hidden states, a projector, and
a linear classifier. Each checkpoint here holds exactly those:
```
ssl_model.featurizer.weights (n_layers + 1,) learnable layer weighting
projector.weight (256, upstream_dim)
projector.bias (256,)
post_net.linear.weight (2, 256)
post_net.linear.bias (2,)
```
The frozen upstream is **not** included. It is loaded from s3prl at run time
under its own name, so nothing here redistributes wav2vec 2.0, HuBERT, WavLM,
XLS-R or any other published encoder.
This is why the whole set is 19 MB rather than 12.2 GB. The original training
checkpoints embedded a full copy of the frozen encoder, 766x larger in total and
byte-identical to what s3prl already serves.
**Verified, not assumed.** The encoder tensors in the original checkpoints were
compared against freshly instantiated s3prl upstreams: identical for every
comparable tensor across `apc` (14/14), `wav2vec2_base_960` (218/218) and `tera`
(57/58). The single TERA difference is `_pseudo_wavs`, a random scratch buffer
s3prl regenerates on load, not a learned parameter.
## Contents
19 SSL detectors, one per row of the paper's main results table:
`apc`, `vq_apc`, `npc`, `mockingjay_960hr`, `tera`, `decoar2`, `wav2vec`,
`wav2vec2_base_960`, `wav2vec2_large_ll60k`, `hubert_base`, `hubert_large_ll60k`,
`multires_hubert_multilingual_large600k`, `xls_r_300m`, `unispeech_sat_large`,
`data2vec_large_ll60k`, `wavlablm_ek_40k`, `wavlm_large`, `ssast_frame_base`,
`mae_ast_frame`
Plus the two non-SSL reference systems under `non_ssl/`:
- `aasist_raw_best_epoch44.pth` — AASIST trained end-to-end on raw waveform.
Full weights, since nothing about it is frozen. This is the best-dev-EER epoch
(1.178%), which is what the paper scores from.
- `lfcc_gmm/gmm_bonafide.pkl`, `lfcc_gmm/gmm_spoof.pkl` — the two 512-component
diagonal-covariance Gaussian mixtures. Scored as the difference of mean
per-frame log-likelihoods.
## Training setup
All SSL detectors share one recipe: ASVspoof 2019 LA train, frozen upstream,
weighted layer aggregation, a lightweight fully connected classifier, binary
cross-entropy with class weights [0.1, 0.9], 50 epochs, batch size 64, SWA. The
checkpoint is the SWA average.
## Usage
Load through the benchmark's own scoring driver, which builds the model, fetches
the matching s3prl upstream and applies these weights.
**Requires the repository at commit `02bf38e` or later.** Earlier revisions
loaded checkpoints with `strict=True` and will reject these files for missing the
upstream tensors. The current loader accepts both shapes and still rejects a
checkpoint that is genuinely incomplete.
```bash
git clone https://github.com/issflab/spoof_SUPERB
python -m spoof_superb.scoring.driver --model linear_head \
--ssl_model xls_r_300m --model_path xls_r_300m.pth \
--source protocol_csv --output_file out.txt
```
## Integrity
`SHA256SUMS` covers every file in this repository:
```bash
sha256sum -c SHA256SUMS
```
## Licence
Apache-2.0, matching the benchmark code. The upstream encoders these detectors
sit on top of remain under their own licences and are not redistributed here.
## Citation
```bibtex
@article{ali2026spoofsuperb,
title = {Spoof-SUPERB: A Comprehensive Benchmark of Self-Supervised Speech
Representations for Audio Deepfake Detection},
author = {Ali, Hashim and Adupa, Nithin Sai and Malik, Hafiz},
year = {2026}
}
```
|