Karaoke voice models (Vitrine)

A mirror of Darkkos/spoti-sing, kept so Vitrine's Sing can download it. The model files and NOTICE are unchanged.

The vocal separator Sing runs on the iPhone to turn a song's vocals down while it plays. It is Mel-Band RoFormer with KimberleyJensen's vocal checkpoint, its spectral core exported for Core ML with two-second windows: a float32 spectrum of shape [1, 2050, 201, 2] in, the vocals_spectrum of the same shape out, the STFT around it done by the app. The files are a compiled separator.mlmodelc. The app downloads each one and keeps it only if its size and SHA-256 match the table. The model declares iOS 18 (Core ML specification 9, the ios18 op set in model.mil). No audio leaves the phone.

File Bytes SHA-256
weights/weight.bin 488986336 970a99fb4b15724bf76d2918ceb177df592c69265d3e2fabaab6e5ba72738e62
model.mil 669061 966560ed5125174a98f19b94f5de04450a7112ade0e731f2236c202c0280a623
metadata.json 2431 52a8d5e3f09e33236d495dbed5bbce1c75bac6f2a6b6097637cf214f37c1de53
coremldata.bin 507 2090acaf7a6df72ec83857cb88d101654023a6baad222a25d0173827d3347e28
analytics/coremldata.bin 243 f7ee4ec9b5cc1c97171bd5aad93af61e183aa0db1451ef3b775bd4e77f0b7cfd

The Neural Engine model: separator-ane.mlmodelc/

The same checkpoint, exported again so the whole model runs on the iPhone's Neural Engine, which keeps the GPU free and the phone cool. Vitrine uses it for Karaoke on devices with a Neural Engine and keeps the model above for the CPU and GPU.

  • Layout: Apple's Neural Engine layout (tensors as (B, C, 1, S), 1ร—1 convolutions for the linear layers, rotary embedding folded into the projections, attention split per head), all fp16, with each RMSNorm dividing a row by its own maximum first so fp16 does not overflow (normalisation ignores scale, so the result is the same). Weights palettized to 6 bits, grouped by 16 channels.
  • Same contract: spectrum [1, 2050, 201, 2] in, vocals_spectrum out (Float16; Float32 input is accepted). Declares iOS 18.
  • Quality: 14.58 dB SDR against the true vocals on four MUSDB18 test excerpts, against 14.61 dB for the model above; 34 dB against that model's own output.
  • Speed: every operation on the Neural Engine, none on the CPU or GPU. On an iPhone 15 Pro: 226โ€“359 ms per 1.5 s of song, against 800โ€“960 ms for the model above on the GPU; the first load compiles for about 30 s, later loads take under half a second.
File Bytes SHA-256
separator-ane.mlmodelc/weights/weight.bin 209077208 bb4a0effafb5121b9aaa5ea96371fa14d0d8dc30a5ae44a93e0c15a2f4077635
separator-ane.mlmodelc/model.mil 1346642 957024c1075fe331f7ca3410b8bf5e2f83a58c61763caf154729829ce0a92d08
separator-ane.mlmodelc/metadata.json 2333 101714d1a37ad29c70bf1705ec25ec99584c3cbce27f2245858ace06dacbfb44
separator-ane.mlmodelc/coremldata.bin 388 bcfd38a8b121681dfb9fdfe6f9d88ffb29068e0dd6d24d4012f4c3a7bdcbc81b
separator-ane.mlmodelc/analytics/coremldata.bin 243 55f78aab64f5b1250ce4d6016dff23aee43fb97feb77bbceab4e36b3d9ccffb3

Converted by Vitrine from KimberleyJensen's checkpoint (below) with coremltools 9; the published checkpoint was checked to match the model above weight for weight.

Provenance

Mel-Band RoFormer by Ju-Chiang Wang, Wei-Tsung Lu and Minz Won; the vocal checkpoint by KimberleyJensen; lucidrains' BS-RoFormer implementation; ZFTurbo's training code. The MIT notices are in NOTICE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support