EfficientAT mn40 (AudioSet, mn40_as_ext) β€” Core ML

EfficientAT mn40_as_ext (width-4.0 MobileNetV3 distilled from AudioSet transformers, 527-class tagging, mAP .487 β€” the strongest released EfficientAT tagger; Schmid et al. 2022, fschmid56/EfficientAT) for Core ML on Apple devices. 68.4M params.

Converted from the same verified reference used by the schism-mlx MLX ports. Two variants per model:

File Precision Compute units max logit diff
EfficientAT_mn40_fp16.mlpackage FLOAT16 ALL (ANE) 3.5e-2, top-5 identical on tested clips
EfficientAT_mn40_fp32.mlpackage FLOAT32 CPU+GPU 5.7e-6

Verified on-device-equivalently via coremltools on an M5 Max, against the reference implementation on real audio. fp16 is ANE-eligible and recommended for iPhone / iPad; fp32 is the tight-parity fallback.

Download

.mlpackage bundles must be materialized as real files β€” the Core ML compiler rejects the symlinks that a default snapshot_download creates in the Hugging Face cache:

from huggingface_hub import snapshot_download
path = snapshot_download("schism-audio/efficient-at-mn40-coreml", local_dir="./efficient-at-mn40-coreml")

(or hf download schism-audio/efficient-at-mn40-coreml --local-dir ./efficient-at-mn40-coreml). Swift hosts downloading files directly are unaffected.

I/O contract

  • input logmel: (1, 1, 128, 1000) float32 β€” EfficientAT mel frontend (32 kHz, n_fft 1024, win 800 symmetric hann zero-padded, hop 320, pre-emphasis 0.97, Kaldi mel 128 bins 0–15000 Hz, ln(x + 1e-5), (x + 4.5) / 5) of a 10 s window, transposed to (mel, time) β€” the frontend emits (frames, mels)
  • 10 s at 32 kHz is exactly 1000 frames (pre-emphasis drops one sample: 1 + 319999 // 320)
  • output logits: (1, 527) float32 β€” apply sigmoid; multi-label
  • longer audio: 1000-frame windows, aggregate scores; shorter: zero-pad the waveform to 10 s before the frontend

DSP frontend (host-side)

The Core ML graph contains the network only. The host implements the audio frontend and must match schism_mlx.audio numerically β€” test_vectors_*.npz in this repo holds deterministic input/output pairs (float32; match within ~1e-4 relative to be interchangeable with what these models were verified against). The architecture is fully convolutional up to the global average pool, but this fixed-shape export takes exactly 10 s windows β€” window the full-file mel and aggregate scores. A validated Swift implementation (Accelerate; modules SchismDSP and SchismPipeline) is available at schism-audio/schism-dsp, tested against these exact vectors.

License

MIT, inherited from fschmid56/EfficientAT (Schmid, Koutini, Widmer β€” CP JKU; arXiv:2211.04772). Core ML conversion by schism-audio.

Downloads last month
14
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for schism-audio/efficient-at-mn40-coreml