--- license: mit library_name: onnx pipeline_tag: audio-to-audio tags: - audio - source-separation - cinematic-audio-source-separation - onnx --- # MRX, ONNX (int8 weights) MRX (Petermann, Wichern, Wang, Le Roux, "The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks", ICASSP 2022, [arXiv:2110.09958](https://arxiv.org/abs/2110.09958)): magnitudes at three STFT resolutions into one hidden layer, a BLSTM per source, a real mask per source and resolution; 30.5 M parameters, MERL's own checkpoint trained on Divide and Remaster with the SNR loss. It splits a soundtrack into three stems: dialogue, music, effects. This is the network between its STFTs and iSTFTs, as [@audio/neural-separate](https://github.com/audiojs/neural/tree/main/packages/neural-separate) runs it (`model: 'mrx'`). | File | Size | SHA-256 | |---|---|---| | `mrx.int8.onnx` | 31.3 MB (31,287,081 bytes) | `d876c92d224f5d8cd2207d528b23369da895a6de71fb24132416fcb16f9f8424` | The float32 export it is made from is 122.3 MB; float16 weights would be 61.5 MB. ## Source - Model, code and checkpoint: [merlresearch/cocktail-fork-separation](https://github.com/merlresearch/cocktail-fork-separation) at `19b3de827ebc4bfb014570cf92dd32b4ee3b6921`: `mrx.py`, `checkpoints/default_mrx_pre_trained_weights.pth`. ## Licence and attribution MIT, Copyright (c) 2023 Mitsubishi Electric Research Laboratories (MERL) ([LICENSE](LICENSE)). The repository's [`.reuse/dep5`](https://github.com/merlresearch/cocktail-fork-separation/blob/19b3de827ebc4bfb014570cf92dd32b4ee3b6921/.reuse/dep5) names the checkpoints: > Files: checkpoints/default_mrx_pre_trained_weights.pth [and its three other checkpoints] > Copyright: 2023 Mitsubishi Electric Research Laboratories (MERL) > License: MIT MRX by Darius Petermann, Gordon Wichern, Zhong-Qiu Wang and Jonathan Le Roux (MERL); ONNX export and compaction by audiojs. Trained on Divide and Remaster v2, whose music and effects come from FMA and FSD50K clips, each under its own licence, some non-commercial. ```bibtex @inproceedings{petermann2022cocktail, title = {The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks}, author = {Petermann, Darius and Wichern, Gordon and Wang, Zhong-Qiu and Le Roux, Jonathan}, booktitle = {IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)}, year = {2022}, eprint = {2110.09958}, archivePrefix = {arXiv} } ``` ## Graph Each channel apart (B), any length (T frames): | | name | shape | | |---|---|---|---| | input | `mag_1024`, `mag_2048`, `mag_8192` | [B, 513, T], [B, 1025, T], [B, 4097, T] | magnitude STFTs: Hann windows of 1024, 2048, 8192, hop 256, scaled by 1/√n, centered with reflect padding, at 44.1 kHz | | output | `mask_1024`, `mask_2048`, `mask_8192` | [B, 3, F, T] | a real mask per source (music, speech, sfx) and resolution | A source is the sum over resolutions of the iSTFT of its mask times that resolution's spectrogram; the input is set to -27 LUFS first and the stems scaled back (upstream's `separate.py`); @audio/neural-separate runs it in 20 s chunks. ```js import separate from '@audio/neural-separate' let { stems } = await separate([mono], { sampleRate: 48000, model: 'mrx' }) // stems.dialogue, .music, .effects ``` ## Export, compaction, verification `scripts/export-mrx.py --verify` exports the network and compares the sources rebuilt from the graph with `MRX.forward` on noise and tones: max |diff| ≤ 1.2e-6 of max |y|; the package's pipeline matches upstream's `separate_soundtrack` to 101–134 dB SNR per stem. `scripts/compact.py --model mrx --calibrate ` makes this file from it: - Weights: all 39 stored in int8 (symmetric, a scale per output channel, an LSTM's per gate row and direction): rounded together they move the masked spectra 40.4 dB under their power on the calibration clip, within the 40 dB the script allows before keeping any in float16. - Computed in float32 on every backend: each weight is Cast and multiplied by its scale in the graph, which onnxruntime folds at load (the session holds float32 weights). - Node and value names shortened (its input length is free, so nothing is folded). - Against the export on the calibration clip: max |diff| 9.9e-3 of max |y|, SNR 40.4 dB (the worst resolution). ## Quality 30 clips of Divide and Remaster v3's English test set (every 40th of its 1200 clips of 60 s; Watcharasupat, Wu, Orife, 2024, CC BY-SA 4.0): each stem against its reference over the whole clip, SNR, the median over clips, dB: | | dialogue | music | effects | |---|---|---|---| | export (float32) | 10.92 | 5.17 | 5.72 | | this file | 10.92 | 5.17 | 5.70 | | change per clip: median · the clip that lost most | +0.00 · −0.01 | +0.00 · −0.04 | +0.00 · −0.02 | SI-SDR (mean over clips) 9.819 · 2.804 · 3.275 against the export's 9.815 · 2.805 · 3.273. Remixes (a stem 6 dB up or down, against the true remix): dialogue +6 dB, music −6 dB, effects −6 dB each within 0.03 dB of the export's median, no clip lower by more than 0.04 dB. Float16 weights (61.5 MB) change no median by more than 0.001 dB.