File size: 4,497 Bytes
dbdbc56 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 | ---
license: mit
tags:
- audio
- audio-classification
- deepfake-detection
- source-tracing
- open-set-recognition
- wavlm
- encodec
---
# sourcetrace β fitted heads for open-set audio deepfake source tracing
Fitted heads for [`sourcetrace`](https://github.com/pujariaditya/sourcetrace), which
names **which generator produced a synthetic speech clip** β or reports that the
generator is not one it has seen.
**These are not standalone models.** Each file is a `torch.save` dict with
`format: "sourcetrace-method-checkpoint"`, loaded by `sourcetrace.method.Method.load`.
They hold the small trained head plus the fitted scoring stack (class anchors,
relative-Mahalanobis density, z-norm constants, conformal calibration, per-block
whitening). The front-ends β `microsoft/wavlm-large` and `facebook/encodec_24khz` β are
frozen, are **not** included here, and are fetched separately by
`scripts/download_models.py`.
## Input
Not audio. A **2133-d** feature vector per clip, laid out as
`[ SSL 0:2048 | signature 2048:2100 | codec residual 2100:2133 ]`, produced by
`scripts/extract_features.py`. There is no way to run these weights without the
repository and an extracted feature cache.
## `mlaad_v5.pt`
MLAAD v5, family-level open-set protocol: 65 known generator families, 9,620 evaluation
trials, split seed 42, fit seed 0.
| Metric | Value | Published SOTA |
|---|---|---|
| FPR95 (lower better) | **1.14 %** | 3.36 % |
| OOD-EER (lower better) | 3.66 % | β |
| Closed-set accuracy | 99.34 % | β |
Conformal abstention: measured coverage 96.11 % against a 95 % nominal level.
**Reproducible, not just reported.** Refitting from the public code at these seeds
reproduces `results/ablation/full.json` bit-for-bit β FPR95 1.1428571428571428, OOD-EER
3.6571428571428575, closed-set 99.33714285714285, under exact equality rather than a
tolerance. This file is that fit. Downloading it saves the ~35-40 min fit and nothing
else: evaluation still reads the feature cache, so MLAAD v5 must be downloaded and
extracted first.
```bash
huggingface-cli download RootAccess4Life/ood-source-tracing mlaad_v5.pt \
--local-dir checkpoints
python scripts/evaluate.py --task mlaad_v5 --checkpoint checkpoints/mlaad_v5.pt
```
**Single-seed.** One split seed, one fit seed. These are point estimates with no
variance attached; do not read the margin over 3.36 % as a measured effect size.
## `stopa.pt`
STOPA cross-corpus open-set protocol: 3 EET attacks fitted, 5 known enrolled, 5 unknown
held out; 33,200 enrolment and 629,800 trial clips, conditions pooled, split seed 42 /
fit seed 0.
| Metric | Value | Published SOTA |
|---|---|---|
| Unknown-attack EER (lower better) | **9.33 %** | 16.43 % |
| Known-attack EER | 10.79 % | β |
Measured by the released code on 2026-08-18; raw numbers in `results/stopa_measured.json`.
A lost earlier implementation had recorded 9.38 % and the paper declines to print it,
because no identified run stood behind it. The figure above is **not** a recovery of that
run β different implementation, and the reference head's seeded initialisation is
unrecoverable. It is an independent measurement that happens to land 0.05 points away,
with a protocol, seeds and a results file behind it. Cite this one, not 9.38.
**Single-seed**, like the MLAAD numbers above.
```bash
huggingface-cli download RootAccess4Life/ood-source-tracing stopa.pt --local-dir checkpoints
python scripts/evaluate.py --task stopa --checkpoint checkpoints/stopa.pt
```
## Limitations
- Trained on MLAAD v5 only. Attribution across other corpora, languages, codecs or
recording conditions is untested.
- The head in the public code is a reconstruction from `docs/METHOD.md`. It does not
reproduce an earlier reference implementation's FPR95 = 0.7543 % / OOD-EER = 2.81 %;
that head's seeded initialisation is unrecoverable and those figures are retired.
- For research on open-set attribution. Not validated for forensic, legal or moderation
use; the abstention rule is calibrated on this protocol and its coverage guarantee
does not transfer off it.
## Licence
MIT, matching the code repository. MLAAD and STOPA carry their own terms; no audio is
redistributed here.
## Citation
```bibtex
@inproceedings{pujari2027frontend,
title = {Front-End Evidence and Calibrated Abstention for
Open-Set Audio Deepfake Attribution},
author = {Pujari, Aditya},
booktitle = {Proc. IEEE ICASSP},
year = {2027}
}
```
|