license: cc-by-4.0
tags:
- audio
- audio-classification
- multi-label
- sound-event-detection
- fsd50k
- pytorch
library_name: apcen-multihead-tagger
datasets:
- Fhrozen/FSD50k
metrics:
- lwlrap
- mean_average_precision
apcen-multihead-tagger β FSD50K (200 labels)
Efficient multi-label audio tagging with a learned, physically-grounded front-end (ERB SuperGaussian filter bank + adaptive PCEN / APCEN) and a three-head model on a shared encoder (peak-pool / modulation-spectrum / transformer-SED) fused by a per-class gate. 5.06M parameters, trained from scratch on FSD50K β no AudioSet pretraining.
Code: https://github.com/Thomas-Durand-Texte/apcen-multihead-tagger
Quick start
pip install git+https://github.com/Thomas-Durand-Texte/apcen-multihead-tagger
from apcen_multihead_tagger.model import ApcenMultiheadTagger
# No argument β downloads this checkpoint from the Hub (cached afterwards).
model = ApcenMultiheadTagger.from_pretrained()
probs = model.predict(waveform) # [B, 200] sigmoid probabilities, mono @ 44.1 kHz
print(model.labels[:3]) # ['Accelerating_and_revving_and_vroom', 'Accordion', ...]
Or from the command line β --weights defaults to this repo:
python -m apcen_multihead_tagger.infer --input clip.wav --topk 5
Files
| file | what it is |
|---|---|
best_model.pt |
the released checkpoint β EMA weights + the 200-label vocabulary, 21.5 MB |
configs/waveform_cnn-19.yaml, configs/shared.yaml |
the exact architecture config (also bundled in the package) |
The checkpoint is a plain torch.save payload: model_state_dict (EMA weights),
vocabulary (200 class names, index-aligned to the logits), and the selection metadata
(epoch, val_lwlrap, val_mAP). Optimizer/scheduler state was stripped for release.
Model description
- Front-end: STFT β analytical ERB SuperGaussian filter bank (72 bands, 50 Hzβ20 kHz) β APCEN (arXiv:2510.18206), a learned closed-loop per-band automatic gain control β learned 2D front-end.
- Encoder: factorized time-frequency (freq-first, frequency-dynamic conv, sub-spectral norm).
- Heads: three reducers with different temporal inductive biases (pool 86 fps / modulation-spectrum 43 fps / transformer-SED), fused by a per-class gate; each head adds an auxiliary loss that co-supervises the shared encoder.
- Output:
[B, 200]logits βsigmoidβ per-class probabilities.
Intended use
- Inference / tagging: predict FSD50K sound-event tags for audio clips.
- Transfer / fine-tuning: a compact backbone for your own labelled audio (see the repo's
fine_tune.py; full fine-tune or--freeze-backbonefeature-extraction). - Feature extraction for downstream audio tasks.
Not intended for speaker ID, transcription, or any safety-critical use.
Metrics (FSD50K validation)
| metric | value |
|---|---|
| lwlrap | 0.735 |
| mAP (macro) | 0.601 |
| params | 5.06M |
Checkpoint selection: best macro-mAP epoch (65), EMA weights. Macro mAP weights every class equally, so this selection favours the long tail; the lwlrap-best epoch (70) scores 0.7350 / 0.5992 β a 0.0002 lwlrap trade for +0.002 mAP.
This is the wc19 line. A larger successor (11.46M params, analytical SuperGaussian filter bank, transformer-only head) is still training and will be published separately β it does not supersede this checkpoint yet.
Per-head (single-head readouts on the same encoder): transformer 0.734 lwlrap / 0.598 mAP, pool 0.664 / 0.497, modulation 0.638 / 0.454. Honest note: at maturity the per-class gate largely concentrates on the transformer head β the pool/modulation heads' fusion contribution is small (~+0.003 mAP); their main value is the encoder co-supervision during training.
Training data
FSD50K (Fonseca et al., 2022) β 200 labels, ~41k clips, CC-BY 4.0. Weak (clip-level) labels, 44.1 kHz. No label corrections applied to this pretrain.
Limitations & biases
- Weak clip-level labels; no temporal event boundaries.
- Long-tailed label distribution β rare classes are harder.
- Trained on FSD50K's Freesound-sourced audio; domain shift (recording conditions, languages, unseen event types) will degrade performance.
- The FSD50K vocabulary is hierarchical (
PurrandAnimalcan both fire) β treat the scores as a multi-label set, not a single argmax.
License & attribution
- Code: MIT.
- Weights / dataset: derived from FSD50K, released under CC-BY 4.0 β attribute FSD50K (Fonseca, Favory, Pons, Font, Serra, "FSD50K: an Open Dataset of Human-Labeled Sound Events", IEEE/ACM TASLP 2022) when redistributing.
Citation
If APCEN is central to your use, cite the adaptive-PCEN work (arXiv:2510.18206) and FSD50K.