Karaoke voice models (Vitrine)
A mirror of Darkkos/spoti-sing, kept so Vitrine's Sing can download it. The model files and NOTICE are unchanged.
The vocal separator Sing runs on the iPhone to turn a song's vocals down while it plays. It is
Mel-Band RoFormer with KimberleyJensen's vocal checkpoint, its spectral core exported for Core ML with
two-second windows: a float32 spectrum of shape [1, 2050, 201, 2] in, the vocals_spectrum of the
same shape out, the STFT around it done by the app. The files are a compiled separator.mlmodelc. The
app downloads each one and keeps it only if its size and SHA-256 match the table. The model declares
iOS 18 (Core ML specification 9, the ios18 op set in model.mil). No audio leaves the
phone.
| File | Bytes | SHA-256 |
|---|---|---|
weights/weight.bin |
488986336 | 970a99fb4b15724bf76d2918ceb177df592c69265d3e2fabaab6e5ba72738e62 |
model.mil |
669061 | 966560ed5125174a98f19b94f5de04450a7112ade0e731f2236c202c0280a623 |
metadata.json |
2431 | 52a8d5e3f09e33236d495dbed5bbce1c75bac6f2a6b6097637cf214f37c1de53 |
coremldata.bin |
507 | 2090acaf7a6df72ec83857cb88d101654023a6baad222a25d0173827d3347e28 |
analytics/coremldata.bin |
243 | f7ee4ec9b5cc1c97171bd5aad93af61e183aa0db1451ef3b775bd4e77f0b7cfd |
The Neural Engine model: separator-ane.mlmodelc/
The same checkpoint, exported again so the whole model runs on the iPhone's Neural Engine, which keeps the GPU free and the phone cool. Vitrine uses it for Karaoke on devices with a Neural Engine and keeps the model above for the CPU and GPU.
- Layout: Apple's Neural Engine layout (tensors as
(B, C, 1, S), 1ร1 convolutions for the linear layers, rotary embedding folded into the projections, attention split per head), all fp16, with each RMSNorm dividing a row by its own maximum first so fp16 does not overflow (normalisation ignores scale, so the result is the same). Weights palettized to 6 bits, grouped by 16 channels. - Same contract:
spectrum[1, 2050, 201, 2]in,vocals_spectrumout (Float16; Float32 input is accepted). Declares iOS 18. - Quality: 14.58 dB SDR against the true vocals on four MUSDB18 test excerpts, against 14.61 dB for the model above; 34 dB against that model's own output.
- Speed: every operation on the Neural Engine, none on the CPU or GPU. On an iPhone 15 Pro: 226โ359 ms per 1.5 s of song, against 800โ960 ms for the model above on the GPU; the first load compiles for about 30 s, later loads take under half a second.
| File | Bytes | SHA-256 |
|---|---|---|
separator-ane.mlmodelc/weights/weight.bin |
209077208 | bb4a0effafb5121b9aaa5ea96371fa14d0d8dc30a5ae44a93e0c15a2f4077635 |
separator-ane.mlmodelc/model.mil |
1346642 | 957024c1075fe331f7ca3410b8bf5e2f83a58c61763caf154729829ce0a92d08 |
separator-ane.mlmodelc/metadata.json |
2333 | 101714d1a37ad29c70bf1705ec25ec99584c3cbce27f2245858ace06dacbfb44 |
separator-ane.mlmodelc/coremldata.bin |
388 | bcfd38a8b121681dfb9fdfe6f9d88ffb29068e0dd6d24d4012f4c3a7bdcbc81b |
separator-ane.mlmodelc/analytics/coremldata.bin |
243 | 55f78aab64f5b1250ce4d6016dff23aee43fb97feb77bbceab4e36b3d9ccffb3 |
Converted by Vitrine from KimberleyJensen's checkpoint (below) with coremltools 9; the published checkpoint was checked to match the model above weight for weight.
Provenance
- Checkpoint: KimberleyJSN/melbandroformer, MIT.
- Conversion: john-rocky/coreai-model-zoo.
- Reference implementation: Mel-Band-Roformer-Vocal-Model at revision
25f44ffb55ee3c301281bba21b2d6d311cb69ae2. - Exported by
harness/sing/export_coreml.pyin the spoti.pw repository.
Mel-Band RoFormer by Ju-Chiang Wang, Wei-Tsung Lu and Minz Won; the vocal checkpoint by KimberleyJensen;
lucidrains' BS-RoFormer implementation; ZFTurbo's training code. The MIT notices are in NOTICE.