File size: 3,603 Bytes
6f14ba4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0c7aca9
 
6f14ba4
 
 
d22e6a0
 
 
 
 
 
 
 
6f14ba4
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
---
license: mit
library_name: coreai
pipeline_tag: audio-to-audio
base_model: KimberleyJSN/melbandroformer
tags: [core-ai, coreaikit, source-separation, stem-separation, vocals, karaoke, on-device, apple]
---

# Mel-Band RoFormer (Kim Vocal) β€” Core AI

[`KimberleyJSN/melbandroformer`](https://huggingface.co/KimberleyJSN/melbandroformer) (MIT, ~228 M)
converted to **Apple Core AI** β€” the [zoo](https://github.com/john-rocky/coreai-model-zoo)'s **first
source-separation model**. Split any song into a **vocals (acapella)** stem and an **instrumental
(karaoke)** stem, entirely on device: iPhone (AOT) and Mac.

Mel-Band RoFormer is `STFT -> band-split (mel, overlapping) -> axial rotary transformer x6 -> mask
estimator -> band-average -> complex mask multiply -> iSTFT`. The neural core lowers to Core AI
directly (no rewrite), and two moves keep the on-device host trivial:

- **Real-arithmetic core** β€” the band-average scatter becomes a constant matmul and the complex mask
  multiply becomes real ops, so the graph carries no complex tensors and no `scatter_add`.
- **STFT/iSTFT folded into the graph** as constant DFT matmuls (window baked in). The shipped graph is
  `frames[1,2,801,2048] -> recon[1,2,801,2048]`; the host only does reflect-pad, framing and
  overlap-add β€” **no FFT, no vDSP packing**.

Fixed shapes throughout (8 s chunk = 352 800 samples @ 44.1 kHz, 801 STFT frames), so there are no
dynamic-shape recompiles. The instrumental stem is `mix - vocals`.

## Contents

- `mbr_full_fp16.aimodel` β€” macOS bundle (fp16, ~470 MB).
- `mbr_full_fp16.h18p.aimodelc` β€” iOS AOT specialization (A19 / h18p, GPU).
- `metadata.json` β€” sample rate, chunk size, STFT parameters, graph I/O, host recipe.
- `golden_raw.f32` / `golden_vocals.f32` β€” an 8 s demo chunk and its expected vocals stem
  (stereo, channel-major, float32) for a host-side self-test.

## Host recipe

Reflect-pad the chunk by `n_fft // 2`, frame with `hop_length`, feed `frames[1,2,801,2048]`, then
overlap-add the output, divide by the summed squared window and trim the pad. Chunks are 8 s with
`num_overlap` crossfade; `instrumental = mix - vocals`.

## Gates

| gate | result |
|---|---|
| re-authored real-arithmetic core vs the reference model | cos **1.0000000** |
| in-graph STFT/iSTFT vs `torch.stft` reference | cos **0.9999984** |
| Core AI fp16, Mac GPU, framing + overlap-add round trip | cos **0.9999453** |
| **iPhone 17 Pro** (A19 Pro, AOT h18p, GPU) vs the Mac golden | cos **1.000000**, rms ratio **1.0000** |

On device (iPhone 17 Pro, GPU): **8 s chunk in 1.23 s ~ 6.5x real-time** warm (load 0.57 s); the
cold first run is 3.82 s ~ 2.1x real-time (load 1.24 s), before the GPU clocks ramp.

## Use it

```swift
import CoreAIKit

let separator = try await KitSeparator(catalog: "melband-roformer-vocal")
let stems = try await separator.separate(contentsOf: songURL)
// stems.vocals / stems.instrumental β€” [channel][sample] at 44.1 kHz
```

Ships in the zoo's **[coreai-audio](https://github.com/john-rocky/coreai-model-zoo/tree/main/apps/coreai-audio)**
app (**Separate** tab), which pairs it with the Music tab (Stable Audio Open Small): generate a
track, then rip its stems. Conversion recipe and the Swift host reference:
[`conversion/melband_roformer`](https://github.com/john-rocky/coreai-model-zoo/tree/main/conversion/melband_roformer).

## Attribution

Mel-Band RoFormer (Ju-Chiang Wang, Wei-Tsung Lu, Minz Won β€” ByteDance AI Labs); checkpoint by
KimberleyJensen; lucidrains BS-RoFormer implementation; ZFTurbo training code. MIT.

*Community port β€” not an Apple model.*