|
Download README.md from audiojs/mrx: direct link, hf CLI and curl.
- Browser
- Download file 5.21 kB
-
https://huggingface.co/audiojs/mrx/resolve/main/README.md
- Command line
-
hf download hf://audiojs/mrx/README.md
-
curl -L -o README.md https://huggingface.co/audiojs/mrx/resolve/main/README.md
5.21 kB
| license: mit | |
| library_name: onnx | |
| pipeline_tag: audio-to-audio | |
| tags: | |
| - audio | |
| - source-separation | |
| - cinematic-audio-source-separation | |
| - onnx | |
| # MRX, ONNX (int8 weights) | |
| MRX (Petermann, Wichern, Wang, Le Roux, "The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World | |
| Soundtracks", ICASSP 2022, [arXiv:2110.09958](https://arxiv.org/abs/2110.09958)): magnitudes at three STFT resolutions | |
| into one hidden layer, a BLSTM per source, a real mask per source and resolution; 30.5 M parameters, MERL's own | |
| checkpoint trained on Divide and Remaster with the SNR loss. It splits a soundtrack into three stems: dialogue, music, | |
| effects. This is the network between its STFTs and iSTFTs, as | |
| [@audio/neural-separate](https://github.com/audiojs/neural/tree/main/packages/neural-separate) runs it (`model: 'mrx'`). | |
| | File | Size | SHA-256 | | |
| |---|---|---| | |
| | `mrx.int8.onnx` | 31.3 MB (31,287,081 bytes) | `d876c92d224f5d8cd2207d528b23369da895a6de71fb24132416fcb16f9f8424` | | |
| The float32 export it is made from is 122.3 MB; float16 weights would be 61.5 MB. | |
| ## Source | |
| - Model, code and checkpoint: [merlresearch/cocktail-fork-separation](https://github.com/merlresearch/cocktail-fork-separation) | |
| at `19b3de827ebc4bfb014570cf92dd32b4ee3b6921`: `mrx.py`, `checkpoints/default_mrx_pre_trained_weights.pth`. | |
| ## Licence and attribution | |
| MIT, Copyright (c) 2023 Mitsubishi Electric Research Laboratories (MERL) ([LICENSE](LICENSE)). The repository's | |
| [`.reuse/dep5`](https://github.com/merlresearch/cocktail-fork-separation/blob/19b3de827ebc4bfb014570cf92dd32b4ee3b6921/.reuse/dep5) | |
| names the checkpoints: | |
| > Files: checkpoints/default_mrx_pre_trained_weights.pth [and its three other checkpoints] | |
| > Copyright: 2023 Mitsubishi Electric Research Laboratories (MERL) | |
| > License: MIT | |
| MRX by Darius Petermann, Gordon Wichern, Zhong-Qiu Wang and Jonathan Le Roux (MERL); ONNX export and compaction by | |
| audiojs. Trained on Divide and Remaster v2, whose music and effects come from FMA and FSD50K clips, each under its own | |
| licence, some non-commercial. | |
| ```bibtex | |
| @inproceedings{petermann2022cocktail, | |
| title = {The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks}, | |
| author = {Petermann, Darius and Wichern, Gordon and Wang, Zhong-Qiu and Le Roux, Jonathan}, | |
| booktitle = {IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)}, | |
| year = {2022}, | |
| eprint = {2110.09958}, | |
| archivePrefix = {arXiv} | |
| } | |
| ``` | |
| ## Graph | |
| Each channel apart (B), any length (T frames): | |
| | | name | shape | | | |
| |---|---|---|---| | |
| | input | `mag_1024`, `mag_2048`, `mag_8192` | [B, 513, T], [B, 1025, T], [B, 4097, T] | magnitude STFTs: Hann windows of 1024, 2048, 8192, hop 256, scaled by 1/√n, centered with reflect padding, at 44.1 kHz | | |
| | output | `mask_1024`, `mask_2048`, `mask_8192` | [B, 3, F, T] | a real mask per source (music, speech, sfx) and resolution | | |
| A source is the sum over resolutions of the iSTFT of its mask times that resolution's spectrogram; the input is set to | |
| -27 LUFS first and the stems scaled back (upstream's `separate.py`); @audio/neural-separate runs it in 20 s chunks. | |
| ```js | |
| import separate from '@audio/neural-separate' | |
| let { stems } = await separate([mono], { sampleRate: 48000, model: 'mrx' }) // stems.dialogue, .music, .effects | |
| ``` | |
| ## Export, compaction, verification | |
| `scripts/export-mrx.py --verify` exports the network and compares the sources rebuilt from the graph with | |
| `MRX.forward` on noise and tones: max |diff| ≤ 1.2e-6 of max |y|; the package's pipeline matches upstream's | |
| `separate_soundtrack` to 101–134 dB SNR per stem. `scripts/compact.py --model mrx --calibrate <a Divide and Remaster v3 | |
| tuning clip>` makes this file from it: | |
| - Weights: all 39 stored in int8 (symmetric, a scale per output channel, an LSTM's per gate row and direction): rounded | |
| together they move the masked spectra 40.4 dB under their power on the calibration clip, within the 40 dB the | |
| script allows before keeping any in float16. | |
| - Computed in float32 on every backend: each weight is Cast and multiplied by its scale in the graph, which onnxruntime | |
| folds at load (the session holds float32 weights). | |
| - Node and value names shortened (its input length is free, so nothing is folded). | |
| - Against the export on the calibration clip: max |diff| 9.9e-3 of max |y|, SNR 40.4 dB (the worst resolution). | |
| ## Quality | |
| 30 clips of Divide and Remaster v3's English test set (every 40th of its 1200 clips of 60 s; Watcharasupat, Wu, Orife, | |
| 2024, CC BY-SA 4.0): each stem against its reference over the whole clip, SNR, the median over clips, dB: | |
| | | dialogue | music | effects | | |
| |---|---|---|---| | |
| | export (float32) | 10.92 | 5.17 | 5.72 | | |
| | this file | 10.92 | 5.17 | 5.70 | | |
| | change per clip: median · the clip that lost most | +0.00 · −0.01 | +0.00 · −0.04 | +0.00 · −0.02 | | |
| SI-SDR (mean over clips) 9.819 · 2.804 · 3.275 against the export's 9.815 · 2.805 · 3.273. Remixes (a stem 6 dB up or | |
| down, against the true remix): dialogue +6 dB, music −6 dB, effects −6 dB each within 0.03 dB of the export's median, | |
| no clip lower by more than 0.04 dB. Float16 weights (61.5 MB) change no median by more than 0.001 dB. | |