Instructions to use schism-audio/efficient-at-mn40-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use schism-audio/efficient-at-mn40-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir efficient-at-mn40-mlx schism-audio/efficient-at-mn40-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
EfficientAT mn40_as_ext (AudioSet) β MLX
EfficientAT mn40_as_ext (MobileNetV3, width 4.0, 68.43M
params, AudioSet mAP 0.487,
Schmid et al. 2022) converted to
MLX for Apple Silicon. Source
checkpoint: mn40_as_ext_mAP_487.pt from the
fschmid56/EfficientAT GitHub
Releases (v0.0.1). The models were distilled from PaSST transformers on
AudioSet (527 classes).
This build uses the mn40_as_ext weights (extended training, 300
epochs) β the best single static MobileNet in the release table
(mAP 0.487 vs 0.484 for the standard mn40_as run); same
architecture, so it loads into the same config-driven model class.
Runs with the open schism-mlx
package:
from schism_mlx.classify.efficient_at import EfficientAT
model = EfficientAT.from_pretrained("schism-audio/efficient-at-mn40-mlx")
for label, score in model.classify("song.wav", top_k=5):
print(f"{score:.3f} {label}")
Handles arbitrary-length audio in a single pass β the network is fully convolutional up to the global average pool, so no windowing or truncation is needed.
Conversion & verification
- Weights are the original float32 checkpoint, tensor-renamed and with convs permuted to MLX's channels-last layout. No quantization, no finetuning.
- Verified end to end against the genuine reference implementation
(fschmid56/EfficientAT,
models/mn+AugmentMelSTFT) on synthetic signals and real drum recordings from schism-audio/e-gmd: max logit difference 4.8e-6 with identical features (2.5e-5 end to end through the numpy frontend), identical top-5 everywhere. ~5.4 ms per 10-second clip on an M5 Max (RTF 0.0005). - The mel frontend (pre-emphasis, 32 kHz, n_fft 1024 / win 800 / hop 320,
128 Kaldi-style mel bins, 0β15000 Hz,
log(mel + 1e-5),(x + 4.5) / 5normalization) is reimplemented in pure numpy from the reference's ownAugmentMelSTFTcode. No torch/torchaudio needed at inference time.
Precision note: on M5-class GPUs, MLX silently runs float32 matmuls at
TF32-class precision (ml-explore/mlx#3860).
schism_mlx disables this on import; if you load these weights with your own
MLX code, set MLX_ENABLE_TF32=0 to match the reference implementation.
License
MIT, inherited from the source repository (code and weights). Model by Florian Schmid, Khaled Koutini, and Gerhard Widmer (CP JKU Linz); MLX conversion by schism-audio.
Citation
@inproceedings{schmid2023efficient,
title={Efficient Large-Scale Audio Tagging Via Transformer-to-CNN Knowledge Distillation},
author={Schmid, Florian and Koutini, Khaled and Widmer, Gerhard},
booktitle={ICASSP 2023},
year={2023}
}
- Downloads last month
- 19
Quantized