mrx / README.md
dyiv's picture
ONNX export, int8 weights, model card
dedc823 verified
|
Raw History Blame Contribute Delete
5.21 kB
---
license: mit
library_name: onnx
pipeline_tag: audio-to-audio
tags:
- audio
- source-separation
- cinematic-audio-source-separation
- onnx
---
# MRX, ONNX (int8 weights)
MRX (Petermann, Wichern, Wang, Le Roux, "The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World
Soundtracks", ICASSP 2022, [arXiv:2110.09958](https://arxiv.org/abs/2110.09958)): magnitudes at three STFT resolutions
into one hidden layer, a BLSTM per source, a real mask per source and resolution; 30.5 M parameters, MERL's own
checkpoint trained on Divide and Remaster with the SNR loss. It splits a soundtrack into three stems: dialogue, music,
effects. This is the network between its STFTs and iSTFTs, as
[@audio/neural-separate](https://github.com/audiojs/neural/tree/main/packages/neural-separate) runs it (`model: 'mrx'`).
| File | Size | SHA-256 |
|---|---|---|
| `mrx.int8.onnx` | 31.3 MB (31,287,081 bytes) | `d876c92d224f5d8cd2207d528b23369da895a6de71fb24132416fcb16f9f8424` |
The float32 export it is made from is 122.3 MB; float16 weights would be 61.5 MB.
## Source
- Model, code and checkpoint: [merlresearch/cocktail-fork-separation](https://github.com/merlresearch/cocktail-fork-separation)
at `19b3de827ebc4bfb014570cf92dd32b4ee3b6921`: `mrx.py`, `checkpoints/default_mrx_pre_trained_weights.pth`.
## Licence and attribution
MIT, Copyright (c) 2023 Mitsubishi Electric Research Laboratories (MERL) ([LICENSE](LICENSE)). The repository's
[`.reuse/dep5`](https://github.com/merlresearch/cocktail-fork-separation/blob/19b3de827ebc4bfb014570cf92dd32b4ee3b6921/.reuse/dep5)
names the checkpoints:
> Files: checkpoints/default_mrx_pre_trained_weights.pth [and its three other checkpoints]
> Copyright: 2023 Mitsubishi Electric Research Laboratories (MERL)
> License: MIT
MRX by Darius Petermann, Gordon Wichern, Zhong-Qiu Wang and Jonathan Le Roux (MERL); ONNX export and compaction by
audiojs. Trained on Divide and Remaster v2, whose music and effects come from FMA and FSD50K clips, each under its own
licence, some non-commercial.
```bibtex
@inproceedings{petermann2022cocktail,
title = {The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks},
author = {Petermann, Darius and Wichern, Gordon and Wang, Zhong-Qiu and Le Roux, Jonathan},
booktitle = {IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
year = {2022},
eprint = {2110.09958},
archivePrefix = {arXiv}
}
```
## Graph
Each channel apart (B), any length (T frames):
| | name | shape | |
|---|---|---|---|
| input | `mag_1024`, `mag_2048`, `mag_8192` | [B, 513, T], [B, 1025, T], [B, 4097, T] | magnitude STFTs: Hann windows of 1024, 2048, 8192, hop 256, scaled by 1/√n, centered with reflect padding, at 44.1 kHz |
| output | `mask_1024`, `mask_2048`, `mask_8192` | [B, 3, F, T] | a real mask per source (music, speech, sfx) and resolution |
A source is the sum over resolutions of the iSTFT of its mask times that resolution's spectrogram; the input is set to
-27 LUFS first and the stems scaled back (upstream's `separate.py`); @audio/neural-separate runs it in 20 s chunks.
```js
import separate from '@audio/neural-separate'
let { stems } = await separate([mono], { sampleRate: 48000, model: 'mrx' }) // stems.dialogue, .music, .effects
```
## Export, compaction, verification
`scripts/export-mrx.py --verify` exports the network and compares the sources rebuilt from the graph with
`MRX.forward` on noise and tones: max |diff| ≤ 1.2e-6 of max |y|; the package's pipeline matches upstream's
`separate_soundtrack` to 101–134 dB SNR per stem. `scripts/compact.py --model mrx --calibrate <a Divide and Remaster v3
tuning clip>` makes this file from it:
- Weights: all 39 stored in int8 (symmetric, a scale per output channel, an LSTM's per gate row and direction): rounded
together they move the masked spectra 40.4 dB under their power on the calibration clip, within the 40 dB the
script allows before keeping any in float16.
- Computed in float32 on every backend: each weight is Cast and multiplied by its scale in the graph, which onnxruntime
folds at load (the session holds float32 weights).
- Node and value names shortened (its input length is free, so nothing is folded).
- Against the export on the calibration clip: max |diff| 9.9e-3 of max |y|, SNR 40.4 dB (the worst resolution).
## Quality
30 clips of Divide and Remaster v3's English test set (every 40th of its 1200 clips of 60 s; Watcharasupat, Wu, Orife,
2024, CC BY-SA 4.0): each stem against its reference over the whole clip, SNR, the median over clips, dB:
| | dialogue | music | effects |
|---|---|---|---|
| export (float32) | 10.92 | 5.17 | 5.72 |
| this file | 10.92 | 5.17 | 5.70 |
| change per clip: median · the clip that lost most | +0.00 · −0.01 | +0.00 · −0.04 | +0.00 · −0.02 |
SI-SDR (mean over clips) 9.819 · 2.804 · 3.275 against the export's 9.815 · 2.805 · 3.273. Remixes (a stem 6 dB up or
down, against the true remix): dialogue +6 dB, music −6 dB, effects −6 dB each within 0.03 dB of the export's median,
no clip lower by more than 0.04 dB. Float16 weights (61.5 MB) change no median by more than 0.001 dB.