dyiv commited on
Commit
e71bcdb
Β·
verified Β·
1 Parent(s): 3b41ee8

ONNX export, int8 weights, model card

Browse files
Files changed (3) hide show
  1. LICENSE +21 -0
  2. README.md +107 -0
  3. scnet.int8.onnx +3 -0
LICENSE ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ MIT License
2
+
3
+ Copyright (c) 2024 starrytong
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
README.md ADDED
@@ -0,0 +1,107 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ library_name: onnx
4
+ pipeline_tag: audio-to-audio
5
+ tags:
6
+ - audio
7
+ - source-separation
8
+ - music-source-separation
9
+ - onnx
10
+ ---
11
+
12
+ # SCNet, ONNX (int8 weights)
13
+
14
+ SCNet (Tong, Zhu, Chen, Kang, Jiang, Li, Wu, Meng, "SCNet: Sparse Compression Network for Music Source
15
+ Separation", ICASSP 2024, [arXiv:2401.13276](https://arxiv.org/abs/2401.13276)): band-split convolutions around dual-path
16
+ LSTMs on the complex spectrogram, 10.1 M parameters (SCNet-large at half the width), trained by its author on
17
+ MUSDB18-HQ. It splits a song into four stems: drums, bass, other, vocals. This is the network between its STFT and iSTFT, as
18
+ [@audio/neural-separate](https://github.com/audiojs/neural/tree/main/packages/neural-separate) runs it
19
+ (`model: 'scnet'`).
20
+
21
+ | File | Size | SHA-256 |
22
+ |---|---|---|
23
+ | `scnet.int8.onnx` | 12.9 MB (12,900,277 bytes) | `98228931494151762a1c4ab1ec7899a894b1f81fd4a509921a9d2f9bacc50845` |
24
+
25
+ The float32 export it is made from is 42.8 MB; float16 weights would be 23.2 MB.
26
+
27
+ ## Source
28
+
29
+ - Model and code: [starrytong/SCNet](https://github.com/starrytong/SCNet) (MIT), at `5d95bf96b19c3eede63248d171efeca8e3abb948`.
30
+ - Checkpoint: `scnet_checkpoint_musdb18.ckpt` (SHA-256 `1bc0d1abb20bfdf966dcd07637bafd03e4bc13653d09ef18bc9b3e342eafe2aa`),
31
+ release v.1.0.6 of [ZFTurbo/Music-Source-Separation-Training](https://github.com/ZFTurbo/Music-Source-Separation-Training)
32
+ (MIT), with its `config_musdb18_scnet.yaml`; the model code that repository's `models/scnet` at
33
+ `84b1eac0887756b4f1a9d7a1ff49105939749ed2`.
34
+
35
+ ## Licence and attribution
36
+
37
+ MIT, Copyright (c) 2024 starrytong ([LICENSE](LICENSE)). The weights' author, in
38
+ [starrytong/SCNet#35](https://github.com/starrytong/SCNet/issues/35#issuecomment-4999873539) (2026-07-17):
39
+
40
+ > I confirm that the released SCNet and SCNet-large pretrained weights are distributed under the MIT License,
41
+ > consistent with the source code. You are welcome to redistribute the original checkpoints and format-converted
42
+ > versions, including ONNX exports, as part of your MIT-licensed tool, with appropriate attribution.
43
+
44
+ SCNet by its authors (starrytong/SCNet); the checkpoint as Music-Source-Separation-Training (Roman Solovyev)
45
+ distributes it; ONNX export and compaction by audiojs. Trained on MUSDB18-HQ (Rafii et al., 2019), licensed for
46
+ educational use; whether that reaches the weights no project has settled.
47
+
48
+ ```bibtex
49
+ @inproceedings{tong2024scnet,
50
+ title = {SCNet: Sparse Compression Network for Music Source Separation},
51
+ author = {Tong, Weinan and Zhu, Jiaxu and Chen, Jun and Kang, Shiyin and Jiang, Tao and Li, Yang and Wu, Zhiyong and Meng, Helen},
52
+ booktitle = {IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
53
+ year = {2024},
54
+ eprint = {2401.13276},
55
+ archivePrefix = {arXiv}
56
+ }
57
+ ```
58
+
59
+ ## Graph
60
+
61
+ One 11 s segment (485,100 samples at 44.1 kHz, padded to 476 frames) per run:
62
+
63
+ | | name | shape | |
64
+ |---|---|---|---|
65
+ | input | `mix_spec` | [1, 4, 2049, 476] | STFT: n 4096, hop 1024, no window, scaled by 1/√4096, centered with reflect padding; L re, L im, R re, R im |
66
+ | output | `stems_spec` | [1, 16, 2049, 476] | each source (drums, bass, other, vocals), channel, re and im |
67
+
68
+ The segments (every 2.75 s), their fades and the input's normalization follow Music-Source-Separation-Training's
69
+ `demix()`; @audio/neural-separate's README, Algorithm, has them.
70
+
71
+ ```js
72
+ import separate from '@audio/neural-separate'
73
+ let { stems } = await separate([left, right], { sampleRate: 44100, model: 'scnet' })
74
+ ```
75
+
76
+ ## Export, compaction, verification
77
+
78
+ `scripts/export-scnet.py --model scnet --verify` exports the network (its rFFT over time as cosine and sine
79
+ products, its GroupNorm statistics reduced axis by axis) and compares the graph with `SCNet.forward` on noise and tones:
80
+ max |diff| ≀ 2.2e-6 of max |y|; the package's pipeline matches `SCNet.forward` on its segments to 123–133 dB SNR per
81
+ stem. `scripts/compact.py --model scnet --calibrate <two MUSDB18 training previews>` makes this file from it:
82
+
83
+ - Weights: 77 of 82 stored in int8 (99.2 % of the values; symmetric, a scale per output channel, an LSTM's per gate row
84
+ and direction), the five layers ending the decoder in float16 (`decoder.2.0`'s convolution, `decoder.1.1`'s three
85
+ transposed convolutions, `decoder.2.1`'s first): rounded alone to int8, each moves the output 27 to 39 dB under its
86
+ power; all 82 together, 23.4 dB.
87
+ - Computed in float32 on every backend: each weight is Cast and multiplied by its scale in the graph, which onnxruntime
88
+ folds at load (the session holds float32 weights).
89
+ - Folded and named short, changing no value: with the input's shape fixed, every value computable from the weights and
90
+ the shapes alone is stored as the graph computes it (its DFT matrices, made in float64 from a Range, which
91
+ onnxruntime-web's WebGPU session cannot place); node and value names are base-36 counters.
92
+ - Against the export, on the calibration previews: max |diff| 1.1e-2 of max |y|, SNR 42.4 dB.
93
+
94
+ ## Quality
95
+
96
+ The 50 MUSDB18 test previews, BSSEval v4 SDR (museval), the median over songs, dB:
97
+
98
+ | | vocals | drums | bass | other |
99
+ |---|---|---|---|---|
100
+ | export (float32) | 9.88 | 9.43 | 8.35 | 6.15 |
101
+ | this file | 9.88 | 9.44 | 8.35 | 6.14 |
102
+ | change per song: median Β· the song that lost most | +0.00 Β· βˆ’0.29 | βˆ’0.00 Β· βˆ’0.04 | βˆ’0.00 Β· βˆ’0.10 | βˆ’0.01 Β· βˆ’0.07 |
103
+
104
+ The βˆ’0.29 dB is a song whose vocal stem is near silence (PR - Happy Daze, βˆ’2.2 dB SDR as exported); with all 82 weights
105
+ in int8 (12.8 MB) it lost 2.6 dB, hence the five in float16. Remixes (the input plus (g βˆ’ 1) times a stem, against the
106
+ true remix): vocals +6 dB 18.23 β†’ 18.23, vocals βˆ’6 dB 21.05 β†’ 21.05, drums βˆ’6 dB 20.91 β†’ 20.92. Float16 weights
107
+ (23.2 MB) change no median by more than 0.001 dB.
scnet.int8.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:98228931494151762a1c4ab1ec7899a894b1f81fd4a509921a9d2f9bacc50845
3
+ size 12900277