--- license: apache-2.0 tags: - coreml - apple-silicon - ios - basic-pitch - audio-to-midi - music-transcription - multipitch-estimation - schism-coreml --- # Basic Pitch — Core ML [Spotify's Basic Pitch](https://github.com/spotify/basic-pitch) (Bittner et al., ICASSP 2022 — a ~17k-parameter instrument-agnostic model for polyphonic note transcription and multipitch estimation, audio → MIDI) for Core ML on Apple devices. `BasicPitch_fp32.mlpackage` **is Spotify's official Core ML artifact** (repo commit `fa5997af`, sha256-pinned, redistributed bit-identical); the fp16 variant is derived from it by recasting the CNN body while keeping the CQT frontend in FLOAT32. Verified against the official `nmp.onnx` reference — the same oracle the [schism-mlx](https://github.com/schism-audio/schism-mlx) MLX port is verified against (posteriorgram parity and note-event decode agreement on synthetic melodies and an E-GMD drum clip). Two variants per model: | File | Precision | Compute units | max posteriorgram diff vs nmp.onnx | |---|---|---|---| | `BasicPitch_fp16.mlpackage` | FLOAT16 body, FLOAT32 CQT frontend | ALL (ANE) | 1.3e-3, decoded note events identical on tested clips | | `BasicPitch_fp32.mlpackage` | FLOAT32 (official artifact, iOS 15+) | CPU+GPU | 1.3e-6 | Verified on-device-equivalently via coremltools on an M5 Max, against the reference implementation on real audio. fp16 is ANE-eligible and recommended for iPhone / iPad; fp32 is the tight-parity fallback. ## Download `.mlpackage` bundles must be materialized as real files — the Core ML compiler rejects the symlinks that a default `snapshot_download` creates in the Hugging Face cache: ```python from huggingface_hub import snapshot_download path = snapshot_download("schism-audio/basic-pitch-coreml", local_dir="./basic-pitch-coreml") ``` (or `hf download schism-audio/basic-pitch-coreml --local-dir ./basic-pitch-coreml`). Swift hosts downloading files directly are unaffected. ## I/O contract - input `input_2`: `(1, 43844, 1)` float32 — one raw ~2 s window of 22050 Hz mono audio (43844 = 2·22050 − 256) - output `Identity`: `(1, 172, 264)` float32 — pitch-contour posteriorgram (3 bins/semitone from A0) - output `Identity_1`: `(1, 172, 88)` float32 — note posteriorgram (semitones A0–C8) - output `Identity_2`: `(1, 172, 88)` float32 — onset posteriorgram - frames are ~86.13 fps (22050/256); all outputs are sigmoid probabilities - full tracks: front-pad 3840 zeros, windows every 36164 samples, drop 15 frames per side of each window's 172 output frames, concatenate, trim to `int(n_samples/36164*142)`; frame→seconds includes a per-window offset with the reference's 0.0018 s magic alignment constant — the exact recipe (plus note-event decoding defaults) is in `test_vectors_cqt.npz` params and implemented in schism-mlx `transcribe/basic_pitch/model.py` - **caveat**: skip the model for windows that are *entirely* zero and emit zero posteriorgram frames instead — the official artifact's normalization diverges from `nmp.onnx` on exact digital silence (≥ 2 s of zero samples) and can hallucinate notes; any nonzero sample in the window is fine (verified to 7e-7 on a mostly-zero window) ## DSP frontend (host-side) Unlike the other models in this org, the Core ML graph **embeds the whole DSP frontend** (multirate CQT, log normalization, harmonic stacking) — the host feeds raw audio windows and implements only windowing, output unwrapping, frame-time mapping and note-event decoding, all documented in the `test_vectors_cqt.npz` params json. The vectors also pin every internal frontend stage (CQT magnitude, normalized log-CQT, harmonic stack — float32, computed by the verified schism-mlx port, 1e-6 parity vs `nmp.onnx`) alongside the posteriorgrams the shipped fp32 package produces for the same window, for hosts that want a native frontend or to validate the pipeline end to end. Note-event decoding is pure post-processing (peak-picking + note segmentation); the schism-mlx numpy transliteration is verified exactly equal to `basic_pitch.note_creation`. ## License Apache-2.0, same as the upstream code and weights. `BasicPitch_fp32.mlpackage` is the unmodified official artifact from [spotify/basic-pitch](https://github.com/spotify/basic-pitch); `BasicPitch_fp16.mlpackage` is derived from it. Cite the [ICASSP 2022 paper](https://arxiv.org/abs/2203.09893) if you use this model in research. Core ML packaging and verification by [schism-audio](https://huggingface.co/schism-audio).