ONNX export, int8 weights, model card
Browse files- LICENSE +21 -0
- README.md +107 -0
- scnet.int8.onnx +3 -0
LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
MIT License
|
| 2 |
+
|
| 3 |
+
Copyright (c) 2024 starrytong
|
| 4 |
+
|
| 5 |
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
| 6 |
+
of this software and associated documentation files (the "Software"), to deal
|
| 7 |
+
in the Software without restriction, including without limitation the rights
|
| 8 |
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
| 9 |
+
copies of the Software, and to permit persons to whom the Software is
|
| 10 |
+
furnished to do so, subject to the following conditions:
|
| 11 |
+
|
| 12 |
+
The above copyright notice and this permission notice shall be included in all
|
| 13 |
+
copies or substantial portions of the Software.
|
| 14 |
+
|
| 15 |
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
| 16 |
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
| 17 |
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
| 18 |
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
| 19 |
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
| 20 |
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
| 21 |
+
SOFTWARE.
|
README.md
ADDED
|
@@ -0,0 +1,107 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
library_name: onnx
|
| 4 |
+
pipeline_tag: audio-to-audio
|
| 5 |
+
tags:
|
| 6 |
+
- audio
|
| 7 |
+
- source-separation
|
| 8 |
+
- music-source-separation
|
| 9 |
+
- onnx
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
# SCNet, ONNX (int8 weights)
|
| 13 |
+
|
| 14 |
+
SCNet (Tong, Zhu, Chen, Kang, Jiang, Li, Wu, Meng, "SCNet: Sparse Compression Network for Music Source
|
| 15 |
+
Separation", ICASSP 2024, [arXiv:2401.13276](https://arxiv.org/abs/2401.13276)): band-split convolutions around dual-path
|
| 16 |
+
LSTMs on the complex spectrogram, 10.1 M parameters (SCNet-large at half the width), trained by its author on
|
| 17 |
+
MUSDB18-HQ. It splits a song into four stems: drums, bass, other, vocals. This is the network between its STFT and iSTFT, as
|
| 18 |
+
[@audio/neural-separate](https://github.com/audiojs/neural/tree/main/packages/neural-separate) runs it
|
| 19 |
+
(`model: 'scnet'`).
|
| 20 |
+
|
| 21 |
+
| File | Size | SHA-256 |
|
| 22 |
+
|---|---|---|
|
| 23 |
+
| `scnet.int8.onnx` | 12.9 MB (12,900,277 bytes) | `98228931494151762a1c4ab1ec7899a894b1f81fd4a509921a9d2f9bacc50845` |
|
| 24 |
+
|
| 25 |
+
The float32 export it is made from is 42.8 MB; float16 weights would be 23.2 MB.
|
| 26 |
+
|
| 27 |
+
## Source
|
| 28 |
+
|
| 29 |
+
- Model and code: [starrytong/SCNet](https://github.com/starrytong/SCNet) (MIT), at `5d95bf96b19c3eede63248d171efeca8e3abb948`.
|
| 30 |
+
- Checkpoint: `scnet_checkpoint_musdb18.ckpt` (SHA-256 `1bc0d1abb20bfdf966dcd07637bafd03e4bc13653d09ef18bc9b3e342eafe2aa`),
|
| 31 |
+
release v.1.0.6 of [ZFTurbo/Music-Source-Separation-Training](https://github.com/ZFTurbo/Music-Source-Separation-Training)
|
| 32 |
+
(MIT), with its `config_musdb18_scnet.yaml`; the model code that repository's `models/scnet` at
|
| 33 |
+
`84b1eac0887756b4f1a9d7a1ff49105939749ed2`.
|
| 34 |
+
|
| 35 |
+
## Licence and attribution
|
| 36 |
+
|
| 37 |
+
MIT, Copyright (c) 2024 starrytong ([LICENSE](LICENSE)). The weights' author, in
|
| 38 |
+
[starrytong/SCNet#35](https://github.com/starrytong/SCNet/issues/35#issuecomment-4999873539) (2026-07-17):
|
| 39 |
+
|
| 40 |
+
> I confirm that the released SCNet and SCNet-large pretrained weights are distributed under the MIT License,
|
| 41 |
+
> consistent with the source code. You are welcome to redistribute the original checkpoints and format-converted
|
| 42 |
+
> versions, including ONNX exports, as part of your MIT-licensed tool, with appropriate attribution.
|
| 43 |
+
|
| 44 |
+
SCNet by its authors (starrytong/SCNet); the checkpoint as Music-Source-Separation-Training (Roman Solovyev)
|
| 45 |
+
distributes it; ONNX export and compaction by audiojs. Trained on MUSDB18-HQ (Rafii et al., 2019), licensed for
|
| 46 |
+
educational use; whether that reaches the weights no project has settled.
|
| 47 |
+
|
| 48 |
+
```bibtex
|
| 49 |
+
@inproceedings{tong2024scnet,
|
| 50 |
+
title = {SCNet: Sparse Compression Network for Music Source Separation},
|
| 51 |
+
author = {Tong, Weinan and Zhu, Jiaxu and Chen, Jun and Kang, Shiyin and Jiang, Tao and Li, Yang and Wu, Zhiyong and Meng, Helen},
|
| 52 |
+
booktitle = {IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
|
| 53 |
+
year = {2024},
|
| 54 |
+
eprint = {2401.13276},
|
| 55 |
+
archivePrefix = {arXiv}
|
| 56 |
+
}
|
| 57 |
+
```
|
| 58 |
+
|
| 59 |
+
## Graph
|
| 60 |
+
|
| 61 |
+
One 11 s segment (485,100 samples at 44.1 kHz, padded to 476 frames) per run:
|
| 62 |
+
|
| 63 |
+
| | name | shape | |
|
| 64 |
+
|---|---|---|---|
|
| 65 |
+
| input | `mix_spec` | [1, 4, 2049, 476] | STFT: n 4096, hop 1024, no window, scaled by 1/β4096, centered with reflect padding; L re, L im, R re, R im |
|
| 66 |
+
| output | `stems_spec` | [1, 16, 2049, 476] | each source (drums, bass, other, vocals), channel, re and im |
|
| 67 |
+
|
| 68 |
+
The segments (every 2.75 s), their fades and the input's normalization follow Music-Source-Separation-Training's
|
| 69 |
+
`demix()`; @audio/neural-separate's README, Algorithm, has them.
|
| 70 |
+
|
| 71 |
+
```js
|
| 72 |
+
import separate from '@audio/neural-separate'
|
| 73 |
+
let { stems } = await separate([left, right], { sampleRate: 44100, model: 'scnet' })
|
| 74 |
+
```
|
| 75 |
+
|
| 76 |
+
## Export, compaction, verification
|
| 77 |
+
|
| 78 |
+
`scripts/export-scnet.py --model scnet --verify` exports the network (its rFFT over time as cosine and sine
|
| 79 |
+
products, its GroupNorm statistics reduced axis by axis) and compares the graph with `SCNet.forward` on noise and tones:
|
| 80 |
+
max |diff| β€ 2.2e-6 of max |y|; the package's pipeline matches `SCNet.forward` on its segments to 123β133 dB SNR per
|
| 81 |
+
stem. `scripts/compact.py --model scnet --calibrate <two MUSDB18 training previews>` makes this file from it:
|
| 82 |
+
|
| 83 |
+
- Weights: 77 of 82 stored in int8 (99.2 % of the values; symmetric, a scale per output channel, an LSTM's per gate row
|
| 84 |
+
and direction), the five layers ending the decoder in float16 (`decoder.2.0`'s convolution, `decoder.1.1`'s three
|
| 85 |
+
transposed convolutions, `decoder.2.1`'s first): rounded alone to int8, each moves the output 27 to 39 dB under its
|
| 86 |
+
power; all 82 together, 23.4 dB.
|
| 87 |
+
- Computed in float32 on every backend: each weight is Cast and multiplied by its scale in the graph, which onnxruntime
|
| 88 |
+
folds at load (the session holds float32 weights).
|
| 89 |
+
- Folded and named short, changing no value: with the input's shape fixed, every value computable from the weights and
|
| 90 |
+
the shapes alone is stored as the graph computes it (its DFT matrices, made in float64 from a Range, which
|
| 91 |
+
onnxruntime-web's WebGPU session cannot place); node and value names are base-36 counters.
|
| 92 |
+
- Against the export, on the calibration previews: max |diff| 1.1e-2 of max |y|, SNR 42.4 dB.
|
| 93 |
+
|
| 94 |
+
## Quality
|
| 95 |
+
|
| 96 |
+
The 50 MUSDB18 test previews, BSSEval v4 SDR (museval), the median over songs, dB:
|
| 97 |
+
|
| 98 |
+
| | vocals | drums | bass | other |
|
| 99 |
+
|---|---|---|---|---|
|
| 100 |
+
| export (float32) | 9.88 | 9.43 | 8.35 | 6.15 |
|
| 101 |
+
| this file | 9.88 | 9.44 | 8.35 | 6.14 |
|
| 102 |
+
| change per song: median Β· the song that lost most | +0.00 Β· β0.29 | β0.00 Β· β0.04 | β0.00 Β· β0.10 | β0.01 Β· β0.07 |
|
| 103 |
+
|
| 104 |
+
The β0.29 dB is a song whose vocal stem is near silence (PR - Happy Daze, β2.2 dB SDR as exported); with all 82 weights
|
| 105 |
+
in int8 (12.8 MB) it lost 2.6 dB, hence the five in float16. Remixes (the input plus (g β 1) times a stem, against the
|
| 106 |
+
true remix): vocals +6 dB 18.23 β 18.23, vocals β6 dB 21.05 β 21.05, drums β6 dB 20.91 β 20.92. Float16 weights
|
| 107 |
+
(23.2 MB) change no median by more than 0.001 dB.
|
scnet.int8.onnx
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:98228931494151762a1c4ab1ec7899a894b1f81fd4a509921a9d2f9bacc50845
|
| 3 |
+
size 12900277
|