mrx / README.md
dyiv's picture
ONNX export, int8 weights, model card
dedc823 verified
|
Raw History Blame Contribute Delete
5.21 kB
metadata
license: mit
library_name: onnx
pipeline_tag: audio-to-audio
tags:
  - audio
  - source-separation
  - cinematic-audio-source-separation
  - onnx

MRX, ONNX (int8 weights)

MRX (Petermann, Wichern, Wang, Le Roux, "The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks", ICASSP 2022, arXiv:2110.09958): magnitudes at three STFT resolutions into one hidden layer, a BLSTM per source, a real mask per source and resolution; 30.5 M parameters, MERL's own checkpoint trained on Divide and Remaster with the SNR loss. It splits a soundtrack into three stems: dialogue, music, effects. This is the network between its STFTs and iSTFTs, as @audio/neural-separate runs it (model: 'mrx').

File Size SHA-256
mrx.int8.onnx 31.3 MB (31,287,081 bytes) d876c92d224f5d8cd2207d528b23369da895a6de71fb24132416fcb16f9f8424

The float32 export it is made from is 122.3 MB; float16 weights would be 61.5 MB.

Source

Licence and attribution

MIT, Copyright (c) 2023 Mitsubishi Electric Research Laboratories (MERL) (LICENSE). The repository's .reuse/dep5 names the checkpoints:

Files: checkpoints/default_mrx_pre_trained_weights.pth [and its three other checkpoints] Copyright: 2023 Mitsubishi Electric Research Laboratories (MERL) License: MIT

MRX by Darius Petermann, Gordon Wichern, Zhong-Qiu Wang and Jonathan Le Roux (MERL); ONNX export and compaction by audiojs. Trained on Divide and Remaster v2, whose music and effects come from FMA and FSD50K clips, each under its own licence, some non-commercial.

@inproceedings{petermann2022cocktail,
  title     = {The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks},
  author    = {Petermann, Darius and Wichern, Gordon and Wang, Zhong-Qiu and Le Roux, Jonathan},
  booktitle = {IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  year      = {2022},
  eprint    = {2110.09958},
  archivePrefix = {arXiv}
}

Graph

Each channel apart (B), any length (T frames):

name shape
input mag_1024, mag_2048, mag_8192 [B, 513, T], [B, 1025, T], [B, 4097, T] magnitude STFTs: Hann windows of 1024, 2048, 8192, hop 256, scaled by 1/√n, centered with reflect padding, at 44.1 kHz
output mask_1024, mask_2048, mask_8192 [B, 3, F, T] a real mask per source (music, speech, sfx) and resolution

A source is the sum over resolutions of the iSTFT of its mask times that resolution's spectrogram; the input is set to -27 LUFS first and the stems scaled back (upstream's separate.py); @audio/neural-separate runs it in 20 s chunks.

import separate from '@audio/neural-separate'
let { stems } = await separate([mono], { sampleRate: 48000, model: 'mrx' })   // stems.dialogue, .music, .effects

Export, compaction, verification

scripts/export-mrx.py --verify exports the network and compares the sources rebuilt from the graph with MRX.forward on noise and tones: max |diff| ≤ 1.2e-6 of max |y|; the package's pipeline matches upstream's separate_soundtrack to 101–134 dB SNR per stem. scripts/compact.py --model mrx --calibrate <a Divide and Remaster v3 tuning clip> makes this file from it:

  • Weights: all 39 stored in int8 (symmetric, a scale per output channel, an LSTM's per gate row and direction): rounded together they move the masked spectra 40.4 dB under their power on the calibration clip, within the 40 dB the script allows before keeping any in float16.
  • Computed in float32 on every backend: each weight is Cast and multiplied by its scale in the graph, which onnxruntime folds at load (the session holds float32 weights).
  • Node and value names shortened (its input length is free, so nothing is folded).
  • Against the export on the calibration clip: max |diff| 9.9e-3 of max |y|, SNR 40.4 dB (the worst resolution).

Quality

30 clips of Divide and Remaster v3's English test set (every 40th of its 1200 clips of 60 s; Watcharasupat, Wu, Orife, 2024, CC BY-SA 4.0): each stem against its reference over the whole clip, SNR, the median over clips, dB:

dialogue music effects
export (float32) 10.92 5.17 5.72
this file 10.92 5.17 5.70
change per clip: median · the clip that lost most +0.00 · −0.01 +0.00 · −0.04 +0.00 · −0.02

SI-SDR (mean over clips) 9.819 · 2.804 · 3.275 against the export's 9.815 · 2.805 · 3.273. Remixes (a stem 6 dB up or down, against the true remix): dialogue +6 dB, music −6 dB, effects −6 dB each within 0.03 dB of the export's median, no clip lower by more than 0.04 dB. Float16 weights (61.5 MB) change no median by more than 0.001 dB.