lyrin commited on
Commit
32eb82a
·
verified ·
1 Parent(s): 8556b90

clean reupload

Browse files
LICENSE ADDED
@@ -0,0 +1,201 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Apache License
2
+ Version 2.0, January 2004
3
+ http://www.apache.org/licenses/
4
+
5
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
6
+
7
+ 1. Definitions.
8
+
9
+ "License" shall mean the terms and conditions for use, reproduction,
10
+ and distribution as defined by Sections 1 through 9 of this document.
11
+
12
+ "Licensor" shall mean the copyright owner or entity authorized by
13
+ the copyright owner that is granting the License.
14
+
15
+ "Legal Entity" shall mean the union of the acting entity and all
16
+ other entities that control, are controlled by, or are under common
17
+ control with that entity. For the purposes of this definition,
18
+ "control" means (i) the power, direct or indirect, to cause the
19
+ direction or management of such entity, whether by contract or
20
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
21
+ outstanding shares, or (iii) beneficial ownership of such entity.
22
+
23
+ "You" (or "Your") shall mean an individual or Legal Entity
24
+ exercising permissions granted by this License.
25
+
26
+ "Source" form shall mean the preferred form for making modifications,
27
+ including but not limited to software source code, documentation
28
+ source, and configuration files.
29
+
30
+ "Object" form shall mean any form resulting from mechanical
31
+ transformation or translation of a Source form, including but
32
+ not limited to compiled object code, generated documentation,
33
+ and conversions to other media types.
34
+
35
+ "Work" shall mean the work of authorship, whether in Source or
36
+ Object form, made available under the License, as indicated by a
37
+ copyright notice that is included in or attached to the work
38
+ (an example is provided in the Appendix below).
39
+
40
+ "Derivative Works" shall mean any work, whether in Source or Object
41
+ form, that is based on (or derived from) the Work and for which the
42
+ editorial revisions, annotations, elaborations, or other modifications
43
+ represent, as a whole, an original work of authorship. For the purposes
44
+ of this License, Derivative Works shall not include works that remain
45
+ separable from, or merely link (or bind by name) to the interfaces of,
46
+ the Work and Derivative Works thereof.
47
+
48
+ "Contribution" shall mean any work of authorship, including
49
+ the original version of the Work and any modifications or additions
50
+ to that Work or Derivative Works thereof, that is intentionally
51
+ submitted to Licensor for inclusion in the Work by the copyright owner
52
+ or by an individual or Legal Entity authorized to submit on behalf of
53
+ the copyright owner. For the purposes of this definition, "submitted"
54
+ means any form of electronic, verbal, or written communication sent
55
+ to the Licensor or its representatives, including but not limited to
56
+ communication on electronic mailing lists, source code control systems,
57
+ and issue tracking systems that are managed by, or on behalf of, the
58
+ Licensor for the purpose of discussing and improving the Work, but
59
+ excluding communication that is conspicuously marked or otherwise
60
+ designated in writing by the copyright owner as "Not a Contribution."
61
+
62
+ "Contributor" shall mean Licensor and any individual or Legal Entity
63
+ on behalf of whom a Contribution has been received by Licensor and
64
+ subsequently incorporated within the Work.
65
+
66
+ 2. Grant of Copyright License. Subject to the terms and conditions of
67
+ this License, each Contributor hereby grants to You a perpetual,
68
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
69
+ copyright license to reproduce, prepare Derivative Works of,
70
+ publicly display, publicly perform, sublicense, and distribute the
71
+ Work and such Derivative Works in Source or Object form.
72
+
73
+ 3. Grant of Patent License. Subject to the terms and conditions of
74
+ this License, each Contributor hereby grants to You a perpetual,
75
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
76
+ (except as stated in this section) patent license to make, have made,
77
+ use, offer to sell, sell, import, and otherwise transfer the Work,
78
+ where such license applies only to those patent claims licensable
79
+ by such Contributor that are necessarily infringed by their
80
+ Contribution(s) alone or by combination of their Contribution(s)
81
+ with the Work to which such Contribution(s) was submitted. If You
82
+ institute patent litigation against any entity (including a
83
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
84
+ or a Contribution incorporated within the Work constitutes direct
85
+ or contributory patent infringement, then any patent licenses
86
+ granted to You under this License for that Work shall terminate
87
+ as of the date such litigation is filed.
88
+
89
+ 4. Redistribution. You may reproduce and distribute copies of the
90
+ Work or Derivative Works thereof in any medium, with or without
91
+ modifications, and in Source or Object form, provided that You
92
+ meet the following conditions:
93
+
94
+ (a) You must give any other recipients of the Work or
95
+ Derivative Works a copy of this License; and
96
+
97
+ (b) You must cause any modified files to carry prominent notices
98
+ stating that You changed the files; and
99
+
100
+ (c) You must retain, in the Source form of any Derivative Works
101
+ that You distribute, all copyright, patent, trademark, and
102
+ attribution notices from the Source form of the Work,
103
+ excluding those notices that do not pertain to any part of
104
+ the Derivative Works; and
105
+
106
+ (d) If the Work includes a "NOTICE" text file as part of its
107
+ distribution, then any Derivative Works that You distribute must
108
+ include a readable copy of the attribution notices contained
109
+ within such NOTICE file, excluding those notices that do not
110
+ pertain to any part of the Derivative Works, in at least one
111
+ of the following places: within a NOTICE text file distributed
112
+ as part of the Derivative Works; within the Source form or
113
+ documentation, if provided along with the Derivative Works; or,
114
+ within a display generated by the Derivative Works, if and
115
+ wherever such third-party notices normally appear. The contents
116
+ of the NOTICE file are for informational purposes only and
117
+ do not modify the License. You may add Your own attribution
118
+ notices within Derivative Works that You distribute, alongside
119
+ or as an addendum to the NOTICE text from the Work, provided
120
+ that such additional attribution notices cannot be construed
121
+ as modifying the License.
122
+
123
+ You may add Your own copyright statement to Your modifications and
124
+ may provide additional or different license terms and conditions
125
+ for use, reproduction, or distribution of Your modifications, or
126
+ for any such Derivative Works as a whole, provided Your use,
127
+ reproduction, and distribution of the Work otherwise complies with
128
+ the conditions stated in this License.
129
+
130
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
131
+ any Contribution intentionally submitted for inclusion in the Work
132
+ by You to the Licensor shall be under the terms and conditions of
133
+ this License, without any additional terms or conditions.
134
+ Notwithstanding the above, nothing herein shall supersede or modify
135
+ the terms of any separate license agreement you may have executed
136
+ with Licensor regarding such Contributions.
137
+
138
+ 6. Trademarks. This License does not grant permission to use the trade
139
+ names, trademarks, service marks, or product names of the Licensor,
140
+ except as required for reasonable and customary use in describing the
141
+ origin of the Work and reproducing the content of the NOTICE file.
142
+
143
+ 7. Disclaimer of Warranty. Unless required by applicable law or
144
+ agreed to in writing, Licensor provides the Work (and each
145
+ Contributor provides its Contributions) on an "AS IS" BASIS,
146
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
147
+ implied, including, without limitation, any warranties or conditions
148
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
149
+ PARTICULAR PURPOSE. You are solely responsible for determining the
150
+ appropriateness of using or redistributing the Work and assume any
151
+ risks associated with Your exercise of permissions under this License.
152
+
153
+ 8. Limitation of Liability. In no event and under no legal theory,
154
+ whether in tort (including negligence), contract, or otherwise,
155
+ unless required by applicable law (such as deliberate and grossly
156
+ negligent acts) or agreed to in writing, shall any Contributor be
157
+ liable to You for damages, including any direct, indirect, special,
158
+ incidental, or consequential damages of any character arising as a
159
+ result of this License or out of the use or inability to use the
160
+ Work (including but not limited to damages for loss of goodwill,
161
+ work stoppage, computer failure or malfunction, or any and all
162
+ other commercial damages or losses), even if such Contributor
163
+ has been advised of the possibility of such damages.
164
+
165
+ 9. Accepting Warranty or Additional Liability. While redistributing
166
+ the Work or Derivative Works thereof, You may choose to offer,
167
+ and charge a fee for, acceptance of support, warranty, indemnity,
168
+ or other liability obligations and/or rights consistent with this
169
+ License. However, in accepting such obligations, You may act only
170
+ on Your own behalf and on Your sole responsibility, not on behalf
171
+ of any other Contributor, and only if You agree to indemnify,
172
+ defend, and hold each Contributor harmless for any liability
173
+ incurred by, or claims asserted against, such Contributor by reason
174
+ of your accepting any such warranty or additional liability.
175
+
176
+ END OF TERMS AND CONDITIONS
177
+
178
+ APPENDIX: How to apply the Apache License to your work.
179
+
180
+ To apply the Apache License to your work, attach the following
181
+ boilerplate notice, with the fields enclosed by brackets "[]"
182
+ replaced with your own identifying information. (Don't include
183
+ the brackets!) The text should be enclosed in the appropriate
184
+ comment syntax for the file format. We also recommend that a
185
+ file or class name and description of purpose be included on the
186
+ same "printed page" as the copyright notice for easier
187
+ identification within third-party archives.
188
+
189
+ Copyright [yyyy] [name of copyright owner]
190
+
191
+ Licensed under the Apache License, Version 2.0 (the "License");
192
+ you may not use this file except in compliance with the License.
193
+ You may obtain a copy of the License at
194
+
195
+ http://www.apache.org/licenses/LICENSE-2.0
196
+
197
+ Unless required by applicable law or agreed to in writing, software
198
+ distributed under the License is distributed on an "AS IS" BASIS,
199
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
200
+ See the License for the specific language governing permissions and
201
+ limitations under the License.
README.md CHANGED
@@ -1,8 +1,144 @@
1
  ---
2
  license: apache-2.0
 
 
 
 
 
 
 
 
 
 
3
  language:
4
- - zh
5
- - en
6
- base_model:
7
- - openbmb/VoxCPM2
8
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: apache-2.0
3
+ base_model: openbmb/VoxCPM2
4
+ pipeline_tag: text-to-speech
5
+ library_name: ncnn
6
+ tags:
7
+ - voxcpm2
8
+ - ncnn
9
+ - text-to-speech
10
+ - speech-synthesis
11
+ - voice-cloning
12
+ - edge-ai
13
  language:
14
+ - ar
15
+ - my
16
+ - zh
17
+ - da
18
+ - nl
19
+ - en
20
+ - fi
21
+ - fr
22
+ - de
23
+ - el
24
+ - he
25
+ - hi
26
+ - id
27
+ - it
28
+ - ja
29
+ - km
30
+ - ko
31
+ - lo
32
+ - ms
33
+ - "no"
34
+ - pl
35
+ - pt
36
+ - ru
37
+ - es
38
+ - sw
39
+ - sv
40
+ - tl
41
+ - th
42
+ - tr
43
+ - vi
44
+ ---
45
+
46
+ # VoxCPM2 NCNN
47
+
48
+ This repository contains an NCNN export of [openbmb/VoxCPM2](https://huggingface.co/openbmb/VoxCPM2) for use with the `voxcpm2-ncnn` C++ runtime.
49
+
50
+ It is a converted runtime asset package, not a newly trained or fine-tuned model. The model weights keep the same Apache-2.0 license as the upstream VoxCPM2 release.
51
+
52
+ ## Model Details
53
+
54
+ - Base model: `openbmb/VoxCPM2`
55
+ - Format: NCNN `.param` / `.bin` component graphs
56
+ - Task: multilingual text-to-speech
57
+ - Output audio: 48 kHz mono PCM, written by the runtime through FFmpeg
58
+ - Runtime target: `voxcpm2-ncnn`
59
+ - License: Apache-2.0 for the model assets
60
+
61
+ VoxCPM2 is a multilingual controllable speech generation model. The upstream release describes support for 30 languages, 9 Chinese dialects, voice design, style-controllable voice cloning, and high-fidelity continuation cloning. This NCNN package targets the modes currently exposed by the `voxcpm2-ncnn` runtime.
62
+
63
+ ## Files
64
+
65
+ The exported model directory contains the runtime assets and a local `LICENSE` copy. The repository root keeps an additional `LICENSE` copy for hosting tools that expect the license at the top level.
66
+
67
+ - `model.json`: NCNN component manifest and runtime settings
68
+ - `*.ncnn.param`, `*.ncnn.bin`: exported NCNN component graphs and weights
69
+ - `tokenizer.json`, `tokenizer_config.json`, `special_tokens_map.json`: tokenizer assets
70
+ - `config.json`, `tokenization_voxcpm2.py`: upstream configuration/tokenizer reference files
71
+ - `LICENSE`: Apache-2.0 license text for the model assets
72
+
73
+ Export-time intermediate files such as TorchScript, PNNX graphs, and generated Python wrappers are intentionally not included.
74
+
75
+ ## Usage
76
+
77
+ Download the model assets into the runtime repository:
78
+
79
+ ```bash
80
+ huggingface-cli download lyrin/voxpm2-ncnn --local-dir assets
81
+ ```
82
+
83
+ Smoke-test the NCNN components:
84
+
85
+ ```bash
86
+ xmake run voxcpm2 -m assets/voxcpm2 --smoke-components
87
+ ```
88
+
89
+ Generate speech from text:
90
+
91
+ ```bash
92
+ xmake run voxcpm2 -m assets/voxcpm2 \
93
+ -t "你好,欢迎使用 VoxCPM2 NCNN。" \
94
+ -o out.wav
95
+ ```
96
+
97
+ Use prompt continuation with prompt audio:
98
+
99
+ ```bash
100
+ xmake run voxcpm2 -m assets/voxcpm2 \
101
+ -t "这是续写测试。" \
102
+ --prompt "你好。" \
103
+ --prompt-audio prompt.wav \
104
+ -o out.wav
105
+ ```
106
+
107
+ Use reference audio:
108
+
109
+ ```bash
110
+ xmake run voxcpm2 -m assets/voxcpm2 \
111
+ -t "这是参考音频测试。" \
112
+ --reference-audio reference.wav \
113
+ -o out.flac
114
+ ```
115
+
116
+ The output format is inferred from the `-o` extension.
117
+
118
+ ## Conversion Notes
119
+
120
+ This package splits VoxCPM2 into NCNN component graphs:
121
+
122
+ - text embedding
123
+ - base and residual decoder steps
124
+ - FSQ and projection layers
125
+ - DiT estimator
126
+ - stop-token head
127
+ - AudioVAE encoder and decoder
128
+
129
+ The runtime uses a page-style KV cache internally while adapting to the current exported decoder-step NCNN graphs.
130
+
131
+ ## Limitations
132
+
133
+ - This is a conversion package; numerical behavior and performance can differ from the upstream PyTorch runtime.
134
+ - Not all upstream inference modes are necessarily exposed by the C++ runtime.
135
+ - Quality, speaker similarity, latency, and memory use depend on the NCNN build, device, Vulkan driver, and input audio quality.
136
+ - Generated speech and voice cloning should be used only with appropriate rights, consent, and safety review.
137
+
138
+ ## Attribution
139
+
140
+ The original VoxCPM2 model is by OpenBMB / ModelBest. Please refer to the upstream [VoxCPM2 model card](https://huggingface.co/openbmb/VoxCPM2), [project repository](https://github.com/OpenBMB/VoxCPM), and [technical report](https://arxiv.org/abs/2606.06928) for model architecture, training, evaluation, and intended-use details.
141
+
142
+ ## License
143
+
144
+ The model assets in this repository are released under Apache-2.0, matching the upstream VoxCPM2 release. The license text is included in `LICENSE`.
audio_vae_encoder.pnnx.bin DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:86bf64bcb7354bc480553ecd70af11da689240e32ccdc852429f1696fd235840
3
- size 192180216
 
 
 
 
audio_vae_encoder.pnnx.param DELETED
@@ -1,149 +0,0 @@
1
- 7767517
2
- 147 146
3
- pnnx.Input pnnx_input_0 0 1 0 #0=(1,5120)f32
4
- torch.unsqueeze torch.unsqueeze_57 1 1 0 1 dim=1 $input=0 #0=(1,5120)f32 #1=(1,1,5120)f32
5
- F.pad F.pad_58 1 1 1 2 mode=constant pad=(6,0) value=None $input=1 #1=(1,1,5120)f32 #2=(1,1,5126)f32
6
- nn.Conv1d conv1d_0 1 1 2 3 bias=True dilation=(1) groups=1 in_channels=1 kernel_size=(7) out_channels=128 padding=(0) padding_mode=zeros stride=(1) @bias=(128)f32 @weight=(128,1,7)f32 $input=2 #2=(1,1,5126)f32 #3=(1,128,5120)f32
7
- pnnx.Attribute audio_vae.encoder.block.1.block.0.block.0 0 1 4 @data=(1,128,1)f32 #4=(1,128,1)f32
8
- pnnx.Attribute pnnx_fold_114 0 1 5 @data=(1,128,1)f32 #5=(1,128,1)f32
9
- pnnx.Expression pnnx_expr_481 3 1 3 5 4 6 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #3=(1,128,5120)f32 #5=(1,128,1)f32 #4=(1,128,1)f32 #6=(1,128,5120)f32
10
- F.pad F.pad_59 1 1 6 7 mode=constant pad=(6,0) value=None $input=6 #6=(1,128,5120)f32 #7=(1,128,5126)f32
11
- nn.Conv1d conv1d_1 1 1 7 8 bias=True dilation=(1) groups=128 in_channels=128 kernel_size=(7) out_channels=128 padding=(0) padding_mode=zeros stride=(1) @bias=(128)f32 @weight=(128,1,7)f32 $input=7 #7=(1,128,5126)f32 #8=(1,128,5120)f32
12
- pnnx.Attribute audio_vae.encoder.block.1.block.0.block.2 0 1 9 @data=(1,128,1)f32 #9=(1,128,1)f32
13
- pnnx.Attribute pnnx_fold_151 0 1 10 @data=(1,128,1)f32 #10=(1,128,1)f32
14
- pnnx.Expression pnnx_expr_465 3 1 8 10 9 11 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #8=(1,128,5120)f32 #10=(1,128,1)f32 #9=(1,128,1)f32 #11=(1,128,5120)f32
15
- nn.Conv1d padconv1d_0 1 1 11 12 bias=True dilation=(1) groups=1 in_channels=128 kernel_size=(1) out_channels=128 padding=(0) padding_mode=zeros stride=(1) @bias=(128)f32 @weight=(128,128,1)f32 $input=11 #11=(1,128,5120)f32 #12=(1,128,5120)f32
16
- pnnx.Expression pnnx_expr_463 2 1 3 12 13 expr=add(@0,@1) #3=(1,128,5120)f32 #12=(1,128,5120)f32 #13=(1,128,5120)f32
17
- pnnx.Attribute audio_vae.encoder.block.1.block.1.block.0 0 1 14 @data=(1,128,1)f32 #14=(1,128,1)f32
18
- pnnx.Attribute pnnx_fold_202 0 1 15 @data=(1,128,1)f32 #15=(1,128,1)f32
19
- pnnx.Expression pnnx_expr_445 3 1 13 15 14 16 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #13=(1,128,5120)f32 #15=(1,128,1)f32 #14=(1,128,1)f32 #16=(1,128,5120)f32
20
- F.pad F.pad_61 1 1 16 17 mode=constant pad=(18,0) value=None $input=16 #16=(1,128,5120)f32 #17=(1,128,5138)f32
21
- nn.Conv1d conv1d_3 1 1 17 18 bias=True dilation=(3) groups=128 in_channels=128 kernel_size=(7) out_channels=128 padding=(0) padding_mode=zeros stride=(1) @bias=(128)f32 @weight=(128,1,7)f32 $input=17 #17=(1,128,5138)f32 #18=(1,128,5120)f32
22
- pnnx.Attribute audio_vae.encoder.block.1.block.1.block.2 0 1 19 @data=(1,128,1)f32 #19=(1,128,1)f32
23
- pnnx.Attribute pnnx_fold_240 0 1 20 @data=(1,128,1)f32 #20=(1,128,1)f32
24
- pnnx.Expression pnnx_expr_429 3 1 18 20 19 21 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #18=(1,128,5120)f32 #20=(1,128,1)f32 #19=(1,128,1)f32 #21=(1,128,5120)f32
25
- nn.Conv1d padconv1d_1 1 1 21 22 bias=True dilation=(1) groups=1 in_channels=128 kernel_size=(1) out_channels=128 padding=(0) padding_mode=zeros stride=(1) @bias=(128)f32 @weight=(128,128,1)f32 $input=21 #21=(1,128,5120)f32 #22=(1,128,5120)f32
26
- pnnx.Expression pnnx_expr_427 2 1 13 22 23 expr=add(@0,@1) #13=(1,128,5120)f32 #22=(1,128,5120)f32 #23=(1,128,5120)f32
27
- pnnx.Attribute audio_vae.encoder.block.1.block.2.block.0 0 1 24 @data=(1,128,1)f32 #24=(1,128,1)f32
28
- pnnx.Attribute pnnx_fold_291 0 1 25 @data=(1,128,1)f32 #25=(1,128,1)f32
29
- pnnx.Expression pnnx_expr_409 3 1 23 25 24 26 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #23=(1,128,5120)f32 #25=(1,128,1)f32 #24=(1,128,1)f32 #26=(1,128,5120)f32
30
- F.pad F.pad_63 1 1 26 27 mode=constant pad=(54,0) value=None $input=26 #26=(1,128,5120)f32 #27=(1,128,5174)f32
31
- nn.Conv1d conv1d_5 1 1 27 28 bias=True dilation=(9) groups=128 in_channels=128 kernel_size=(7) out_channels=128 padding=(0) padding_mode=zeros stride=(1) @bias=(128)f32 @weight=(128,1,7)f32 $input=27 #27=(1,128,5174)f32 #28=(1,128,5120)f32
32
- pnnx.Attribute audio_vae.encoder.block.1.block.2.block.2 0 1 29 @data=(1,128,1)f32 #29=(1,128,1)f32
33
- pnnx.Attribute pnnx_fold_329 0 1 30 @data=(1,128,1)f32 #30=(1,128,1)f32
34
- pnnx.Expression pnnx_expr_393 3 1 28 30 29 31 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #28=(1,128,5120)f32 #30=(1,128,1)f32 #29=(1,128,1)f32 #31=(1,128,5120)f32
35
- nn.Conv1d padconv1d_2 1 1 31 32 bias=True dilation=(1) groups=1 in_channels=128 kernel_size=(1) out_channels=128 padding=(0) padding_mode=zeros stride=(1) @bias=(128)f32 @weight=(128,128,1)f32 $input=31 #31=(1,128,5120)f32 #32=(1,128,5120)f32
36
- pnnx.Expression pnnx_expr_391 2 1 23 32 33 expr=add(@0,@1) #23=(1,128,5120)f32 #32=(1,128,5120)f32 #33=(1,128,5120)f32
37
- pnnx.Attribute audio_vae.encoder.block.1.block.3 0 1 34 @data=(1,128,1)f32 #34=(1,128,1)f32
38
- pnnx.Attribute pnnx_fold_365 0 1 35 @data=(1,128,1)f32 #35=(1,128,1)f32
39
- pnnx.Expression pnnx_expr_375 3 1 33 35 34 36 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #33=(1,128,5120)f32 #35=(1,128,1)f32 #34=(1,128,1)f32 #36=(1,128,5120)f32
40
- F.pad F.pad_65 1 1 36 37 mode=constant pad=(2,0) value=None $input=36 #36=(1,128,5120)f32 #37=(1,128,5122)f32
41
- nn.Conv1d conv1d_7 1 1 37 38 bias=True dilation=(1) groups=1 in_channels=128 kernel_size=(4) out_channels=256 padding=(0) padding_mode=zeros stride=(2) @bias=(256)f32 @weight=(256,128,4)f32 $input=37 #37=(1,128,5122)f32 #38=(1,256,2560)f32
42
- pnnx.Attribute audio_vae.encoder.block.2.block.0.block.0 0 1 39 @data=(1,256,1)f32 #39=(1,256,1)f32
43
- pnnx.Attribute pnnx_fold_427 0 1 40 @data=(1,256,1)f32 #40=(1,256,1)f32
44
- pnnx.Expression pnnx_expr_356 3 1 38 40 39 41 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #38=(1,256,2560)f32 #40=(1,256,1)f32 #39=(1,256,1)f32 #41=(1,256,2560)f32
45
- F.pad F.pad_66 1 1 41 42 mode=constant pad=(6,0) value=None $input=41 #41=(1,256,2560)f32 #42=(1,256,2566)f32
46
- nn.Conv1d conv1d_8 1 1 42 43 bias=True dilation=(1) groups=256 in_channels=256 kernel_size=(7) out_channels=256 padding=(0) padding_mode=zeros stride=(1) @bias=(256)f32 @weight=(256,1,7)f32 $input=42 #42=(1,256,2566)f32 #43=(1,256,2560)f32
47
- pnnx.Attribute audio_vae.encoder.block.2.block.0.block.2 0 1 44 @data=(1,256,1)f32 #44=(1,256,1)f32
48
- pnnx.Attribute pnnx_fold_464 0 1 45 @data=(1,256,1)f32 #45=(1,256,1)f32
49
- pnnx.Expression pnnx_expr_340 3 1 43 45 44 46 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #43=(1,256,2560)f32 #45=(1,256,1)f32 #44=(1,256,1)f32 #46=(1,256,2560)f32
50
- nn.Conv1d padconv1d_3 1 1 46 47 bias=True dilation=(1) groups=1 in_channels=256 kernel_size=(1) out_channels=256 padding=(0) padding_mode=zeros stride=(1) @bias=(256)f32 @weight=(256,256,1)f32 $input=46 #46=(1,256,2560)f32 #47=(1,256,2560)f32
51
- pnnx.Expression pnnx_expr_338 2 1 38 47 48 expr=add(@0,@1) #38=(1,256,2560)f32 #47=(1,256,2560)f32 #48=(1,256,2560)f32
52
- pnnx.Attribute audio_vae.encoder.block.2.block.1.block.0 0 1 49 @data=(1,256,1)f32 #49=(1,256,1)f32
53
- pnnx.Attribute pnnx_fold_515 0 1 50 @data=(1,256,1)f32 #50=(1,256,1)f32
54
- pnnx.Expression pnnx_expr_320 3 1 48 50 49 51 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #48=(1,256,2560)f32 #50=(1,256,1)f32 #49=(1,256,1)f32 #51=(1,256,2560)f32
55
- F.pad F.pad_68 1 1 51 52 mode=constant pad=(18,0) value=None $input=51 #51=(1,256,2560)f32 #52=(1,256,2578)f32
56
- nn.Conv1d conv1d_10 1 1 52 53 bias=True dilation=(3) groups=256 in_channels=256 kernel_size=(7) out_channels=256 padding=(0) padding_mode=zeros stride=(1) @bias=(256)f32 @weight=(256,1,7)f32 $input=52 #52=(1,256,2578)f32 #53=(1,256,2560)f32
57
- pnnx.Attribute audio_vae.encoder.block.2.block.1.block.2 0 1 54 @data=(1,256,1)f32 #54=(1,256,1)f32
58
- pnnx.Attribute pnnx_fold_553 0 1 55 @data=(1,256,1)f32 #55=(1,256,1)f32
59
- pnnx.Expression pnnx_expr_304 3 1 53 55 54 56 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #53=(1,256,2560)f32 #55=(1,256,1)f32 #54=(1,256,1)f32 #56=(1,256,2560)f32
60
- nn.Conv1d padconv1d_4 1 1 56 57 bias=True dilation=(1) groups=1 in_channels=256 kernel_size=(1) out_channels=256 padding=(0) padding_mode=zeros stride=(1) @bias=(256)f32 @weight=(256,256,1)f32 $input=56 #56=(1,256,2560)f32 #57=(1,256,2560)f32
61
- pnnx.Expression pnnx_expr_302 2 1 48 57 58 expr=add(@0,@1) #48=(1,256,2560)f32 #57=(1,256,2560)f32 #58=(1,256,2560)f32
62
- pnnx.Attribute audio_vae.encoder.block.2.block.2.block.0 0 1 59 @data=(1,256,1)f32 #59=(1,256,1)f32
63
- pnnx.Attribute pnnx_fold_604 0 1 60 @data=(1,256,1)f32 #60=(1,256,1)f32
64
- pnnx.Expression pnnx_expr_284 3 1 58 60 59 61 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #58=(1,256,2560)f32 #60=(1,256,1)f32 #59=(1,256,1)f32 #61=(1,256,2560)f32
65
- F.pad F.pad_70 1 1 61 62 mode=constant pad=(54,0) value=None $input=61 #61=(1,256,2560)f32 #62=(1,256,2614)f32
66
- nn.Conv1d conv1d_12 1 1 62 63 bias=True dilation=(9) groups=256 in_channels=256 kernel_size=(7) out_channels=256 padding=(0) padding_mode=zeros stride=(1) @bias=(256)f32 @weight=(256,1,7)f32 $input=62 #62=(1,256,2614)f32 #63=(1,256,2560)f32
67
- pnnx.Attribute audio_vae.encoder.block.2.block.2.block.2 0 1 64 @data=(1,256,1)f32 #64=(1,256,1)f32
68
- pnnx.Attribute pnnx_fold_642 0 1 65 @data=(1,256,1)f32 #65=(1,256,1)f32
69
- pnnx.Expression pnnx_expr_268 3 1 63 65 64 66 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #63=(1,256,2560)f32 #65=(1,256,1)f32 #64=(1,256,1)f32 #66=(1,256,2560)f32
70
- nn.Conv1d padconv1d_5 1 1 66 67 bias=True dilation=(1) groups=1 in_channels=256 kernel_size=(1) out_channels=256 padding=(0) padding_mode=zeros stride=(1) @bias=(256)f32 @weight=(256,256,1)f32 $input=66 #66=(1,256,2560)f32 #67=(1,256,2560)f32
71
- pnnx.Expression pnnx_expr_266 2 1 58 67 68 expr=add(@0,@1) #58=(1,256,2560)f32 #67=(1,256,2560)f32 #68=(1,256,2560)f32
72
- pnnx.Attribute audio_vae.encoder.block.2.block.3 0 1 69 @data=(1,256,1)f32 #69=(1,256,1)f32
73
- pnnx.Attribute pnnx_fold_678 0 1 70 @data=(1,256,1)f32 #70=(1,256,1)f32
74
- pnnx.Expression pnnx_expr_250 3 1 68 70 69 71 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #68=(1,256,2560)f32 #70=(1,256,1)f32 #69=(1,256,1)f32 #71=(1,256,2560)f32
75
- F.pad F.pad_72 1 1 71 72 mode=constant pad=(5,0) value=None $input=71 #71=(1,256,2560)f32 #72=(1,256,2565)f32
76
- nn.Conv1d conv1d_14 1 1 72 73 bias=True dilation=(1) groups=1 in_channels=256 kernel_size=(10) out_channels=512 padding=(0) padding_mode=zeros stride=(5) @bias=(512)f32 @weight=(512,256,10)f32 $input=72 #72=(1,256,2565)f32 #73=(1,512,512)f32
77
- pnnx.Attribute audio_vae.encoder.block.3.block.0.block.0 0 1 74 @data=(1,512,1)f32 #74=(1,512,1)f32
78
- pnnx.Attribute pnnx_fold_740 0 1 75 @data=(1,512,1)f32 #75=(1,512,1)f32
79
- pnnx.Expression pnnx_expr_231 3 1 73 75 74 76 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #73=(1,512,512)f32 #75=(1,512,1)f32 #74=(1,512,1)f32 #76=(1,512,512)f32
80
- F.pad F.pad_73 1 1 76 77 mode=constant pad=(6,0) value=None $input=76 #76=(1,512,512)f32 #77=(1,512,518)f32
81
- nn.Conv1d conv1d_15 1 1 77 78 bias=True dilation=(1) groups=512 in_channels=512 kernel_size=(7) out_channels=512 padding=(0) padding_mode=zeros stride=(1) @bias=(512)f32 @weight=(512,1,7)f32 $input=77 #77=(1,512,518)f32 #78=(1,512,512)f32
82
- pnnx.Attribute audio_vae.encoder.block.3.block.0.block.2 0 1 79 @data=(1,512,1)f32 #79=(1,512,1)f32
83
- pnnx.Attribute pnnx_fold_777 0 1 80 @data=(1,512,1)f32 #80=(1,512,1)f32
84
- pnnx.Expression pnnx_expr_215 3 1 78 80 79 81 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #78=(1,512,512)f32 #80=(1,512,1)f32 #79=(1,512,1)f32 #81=(1,512,512)f32
85
- nn.Conv1d padconv1d_6 1 1 81 82 bias=True dilation=(1) groups=1 in_channels=512 kernel_size=(1) out_channels=512 padding=(0) padding_mode=zeros stride=(1) @bias=(512)f32 @weight=(512,512,1)f32 $input=81 #81=(1,512,512)f32 #82=(1,512,512)f32
86
- pnnx.Expression pnnx_expr_213 2 1 73 82 83 expr=add(@0,@1) #73=(1,512,512)f32 #82=(1,512,512)f32 #83=(1,512,512)f32
87
- pnnx.Attribute audio_vae.encoder.block.3.block.1.block.0 0 1 84 @data=(1,512,1)f32 #84=(1,512,1)f32
88
- pnnx.Attribute pnnx_fold_828 0 1 85 @data=(1,512,1)f32 #85=(1,512,1)f32
89
- pnnx.Expression pnnx_expr_195 3 1 83 85 84 86 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #83=(1,512,512)f32 #85=(1,512,1)f32 #84=(1,512,1)f32 #86=(1,512,512)f32
90
- F.pad F.pad_75 1 1 86 87 mode=constant pad=(18,0) value=None $input=86 #86=(1,512,512)f32 #87=(1,512,530)f32
91
- nn.Conv1d conv1d_17 1 1 87 88 bias=True dilation=(3) groups=512 in_channels=512 kernel_size=(7) out_channels=512 padding=(0) padding_mode=zeros stride=(1) @bias=(512)f32 @weight=(512,1,7)f32 $input=87 #87=(1,512,530)f32 #88=(1,512,512)f32
92
- pnnx.Attribute audio_vae.encoder.block.3.block.1.block.2 0 1 89 @data=(1,512,1)f32 #89=(1,512,1)f32
93
- pnnx.Attribute pnnx_fold_866 0 1 90 @data=(1,512,1)f32 #90=(1,512,1)f32
94
- pnnx.Expression pnnx_expr_179 3 1 88 90 89 91 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #88=(1,512,512)f32 #90=(1,512,1)f32 #89=(1,512,1)f32 #91=(1,512,512)f32
95
- nn.Conv1d padconv1d_7 1 1 91 92 bias=True dilation=(1) groups=1 in_channels=512 kernel_size=(1) out_channels=512 padding=(0) padding_mode=zeros stride=(1) @bias=(512)f32 @weight=(512,512,1)f32 $input=91 #91=(1,512,512)f32 #92=(1,512,512)f32
96
- pnnx.Expression pnnx_expr_177 2 1 83 92 93 expr=add(@0,@1) #83=(1,512,512)f32 #92=(1,512,512)f32 #93=(1,512,512)f32
97
- pnnx.Attribute audio_vae.encoder.block.3.block.2.block.0 0 1 94 @data=(1,512,1)f32 #94=(1,512,1)f32
98
- pnnx.Attribute pnnx_fold_917 0 1 95 @data=(1,512,1)f32 #95=(1,512,1)f32
99
- pnnx.Expression pnnx_expr_159 3 1 93 95 94 96 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #93=(1,512,512)f32 #95=(1,512,1)f32 #94=(1,512,1)f32 #96=(1,512,512)f32
100
- F.pad F.pad_77 1 1 96 97 mode=constant pad=(54,0) value=None $input=96 #96=(1,512,512)f32 #97=(1,512,566)f32
101
- nn.Conv1d conv1d_19 1 1 97 98 bias=True dilation=(9) groups=512 in_channels=512 kernel_size=(7) out_channels=512 padding=(0) padding_mode=zeros stride=(1) @bias=(512)f32 @weight=(512,1,7)f32 $input=97 #97=(1,512,566)f32 #98=(1,512,512)f32
102
- pnnx.Attribute audio_vae.encoder.block.3.block.2.block.2 0 1 99 @data=(1,512,1)f32 #99=(1,512,1)f32
103
- pnnx.Attribute pnnx_fold_955 0 1 100 @data=(1,512,1)f32 #100=(1,512,1)f32
104
- pnnx.Expression pnnx_expr_143 3 1 98 100 99 101 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #98=(1,512,512)f32 #100=(1,512,1)f32 #99=(1,512,1)f32 #101=(1,512,512)f32
105
- nn.Conv1d padconv1d_8 1 1 101 102 bias=True dilation=(1) groups=1 in_channels=512 kernel_size=(1) out_channels=512 padding=(0) padding_mode=zeros stride=(1) @bias=(512)f32 @weight=(512,512,1)f32 $input=101 #101=(1,512,512)f32 #102=(1,512,512)f32
106
- pnnx.Expression pnnx_expr_141 2 1 93 102 103 expr=add(@0,@1) #93=(1,512,512)f32 #102=(1,512,512)f32 #103=(1,512,512)f32
107
- pnnx.Attribute audio_vae.encoder.block.3.block.3 0 1 104 @data=(1,512,1)f32 #104=(1,512,1)f32
108
- pnnx.Attribute pnnx_fold_991 0 1 105 @data=(1,512,1)f32 #105=(1,512,1)f32
109
- pnnx.Expression pnnx_expr_125 3 1 103 105 104 106 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #103=(1,512,512)f32 #105=(1,512,1)f32 #104=(1,512,1)f32 #106=(1,512,512)f32
110
- F.pad F.pad_79 1 1 106 107 mode=constant pad=(8,0) value=None $input=106 #106=(1,512,512)f32 #107=(1,512,520)f32
111
- nn.Conv1d conv1d_21 1 1 107 108 bias=True dilation=(1) groups=1 in_channels=512 kernel_size=(16) out_channels=1024 padding=(0) padding_mode=zeros stride=(8) @bias=(1024)f32 @weight=(1024,512,16)f32 $input=107 #107=(1,512,520)f32 #108=(1,1024,64)f32
112
- pnnx.Attribute audio_vae.encoder.block.4.block.0.block.0 0 1 109 @data=(1,1024,1)f32 #109=(1,1024,1)f32
113
- pnnx.Attribute pnnx_fold_1053 0 1 110 @data=(1,1024,1)f32 #110=(1,1024,1)f32
114
- pnnx.Expression pnnx_expr_106 3 1 108 110 109 111 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #108=(1,1024,64)f32 #110=(1,1024,1)f32 #109=(1,1024,1)f32 #111=(1,1024,64)f32
115
- F.pad F.pad_80 1 1 111 112 mode=constant pad=(6,0) value=None $input=111 #111=(1,1024,64)f32 #112=(1,1024,70)f32
116
- nn.Conv1d conv1d_22 1 1 112 113 bias=True dilation=(1) groups=1024 in_channels=1024 kernel_size=(7) out_channels=1024 padding=(0) padding_mode=zeros stride=(1) @bias=(1024)f32 @weight=(1024,1,7)f32 $input=112 #112=(1,1024,70)f32 #113=(1,1024,64)f32
117
- pnnx.Attribute audio_vae.encoder.block.4.block.0.block.2 0 1 114 @data=(1,1024,1)f32 #114=(1,1024,1)f32
118
- pnnx.Attribute pnnx_fold_1090 0 1 115 @data=(1,1024,1)f32 #115=(1,1024,1)f32
119
- pnnx.Expression pnnx_expr_90 3 1 113 115 114 116 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #113=(1,1024,64)f32 #115=(1,1024,1)f32 #114=(1,1024,1)f32 #116=(1,1024,64)f32
120
- nn.Conv1d padconv1d_9 1 1 116 117 bias=True dilation=(1) groups=1 in_channels=1024 kernel_size=(1) out_channels=1024 padding=(0) padding_mode=zeros stride=(1) @bias=(1024)f32 @weight=(1024,1024,1)f32 $input=116 #116=(1,1024,64)f32 #117=(1,1024,64)f32
121
- pnnx.Expression pnnx_expr_88 2 1 108 117 118 expr=add(@0,@1) #108=(1,1024,64)f32 #117=(1,1024,64)f32 #118=(1,1024,64)f32
122
- pnnx.Attribute audio_vae.encoder.block.4.block.1.block.0 0 1 119 @data=(1,1024,1)f32 #119=(1,1024,1)f32
123
- pnnx.Attribute pnnx_fold_1141 0 1 120 @data=(1,1024,1)f32 #120=(1,1024,1)f32
124
- pnnx.Expression pnnx_expr_70 3 1 118 120 119 121 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #118=(1,1024,64)f32 #120=(1,1024,1)f32 #119=(1,1024,1)f32 #121=(1,1024,64)f32
125
- F.pad F.pad_82 1 1 121 122 mode=constant pad=(18,0) value=None $input=121 #121=(1,1024,64)f32 #122=(1,1024,82)f32
126
- nn.Conv1d conv1d_24 1 1 122 123 bias=True dilation=(3) groups=1024 in_channels=1024 kernel_size=(7) out_channels=1024 padding=(0) padding_mode=zeros stride=(1) @bias=(1024)f32 @weight=(1024,1,7)f32 $input=122 #122=(1,1024,82)f32 #123=(1,1024,64)f32
127
- pnnx.Attribute audio_vae.encoder.block.4.block.1.block.2 0 1 124 @data=(1,1024,1)f32 #124=(1,1024,1)f32
128
- pnnx.Attribute pnnx_fold_1179 0 1 125 @data=(1,1024,1)f32 #125=(1,1024,1)f32
129
- pnnx.Expression pnnx_expr_54 3 1 123 125 124 126 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #123=(1,1024,64)f32 #125=(1,1024,1)f32 #124=(1,1024,1)f32 #126=(1,1024,64)f32
130
- nn.Conv1d padconv1d_10 1 1 126 127 bias=True dilation=(1) groups=1 in_channels=1024 kernel_size=(1) out_channels=1024 padding=(0) padding_mode=zeros stride=(1) @bias=(1024)f32 @weight=(1024,1024,1)f32 $input=126 #126=(1,1024,64)f32 #127=(1,1024,64)f32
131
- pnnx.Expression pnnx_expr_52 2 1 118 127 128 expr=add(@0,@1) #118=(1,1024,64)f32 #127=(1,1024,64)f32 #128=(1,1024,64)f32
132
- pnnx.Attribute audio_vae.encoder.block.4.block.2.block.0 0 1 129 @data=(1,1024,1)f32 #129=(1,1024,1)f32
133
- pnnx.Attribute pnnx_fold_1230 0 1 130 @data=(1,1024,1)f32 #130=(1,1024,1)f32
134
- pnnx.Expression pnnx_expr_34 3 1 128 130 129 131 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #128=(1,1024,64)f32 #130=(1,1024,1)f32 #129=(1,1024,1)f32 #131=(1,1024,64)f32
135
- F.pad F.pad_84 1 1 131 132 mode=constant pad=(54,0) value=None $input=131 #131=(1,1024,64)f32 #132=(1,1024,118)f32
136
- nn.Conv1d conv1d_26 1 1 132 133 bias=True dilation=(9) groups=1024 in_channels=1024 kernel_size=(7) out_channels=1024 padding=(0) padding_mode=zeros stride=(1) @bias=(1024)f32 @weight=(1024,1,7)f32 $input=132 #132=(1,1024,118)f32 #133=(1,1024,64)f32
137
- pnnx.Attribute audio_vae.encoder.block.4.block.2.block.2 0 1 134 @data=(1,1024,1)f32 #134=(1,1024,1)f32
138
- pnnx.Attribute pnnx_fold_1268 0 1 135 @data=(1,1024,1)f32 #135=(1,1024,1)f32
139
- pnnx.Expression pnnx_expr_18 3 1 133 135 134 136 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #133=(1,1024,64)f32 #135=(1,1024,1)f32 #134=(1,1024,1)f32 #136=(1,1024,64)f32
140
- nn.Conv1d padconv1d_11 1 1 136 137 bias=True dilation=(1) groups=1 in_channels=1024 kernel_size=(1) out_channels=1024 padding=(0) padding_mode=zeros stride=(1) @bias=(1024)f32 @weight=(1024,1024,1)f32 $input=136 #136=(1,1024,64)f32 #137=(1,1024,64)f32
141
- pnnx.Expression pnnx_expr_16 2 1 128 137 138 expr=add(@0,@1) #128=(1,1024,64)f32 #137=(1,1024,64)f32 #138=(1,1024,64)f32
142
- pnnx.Attribute audio_vae.encoder.block.4.block.3 0 1 139 @data=(1,1024,1)f32 #139=(1,1024,1)f32
143
- pnnx.Attribute pnnx_fold_1304 0 1 140 @data=(1,1024,1)f32 #140=(1,1024,1)f32
144
- pnnx.Expression pnnx_expr_0 3 1 138 140 139 141 expr=add(@0,mul(@1,pow(sin(mul(@2,@0)),2))) #138=(1,1024,64)f32 #140=(1,1024,1)f32 #139=(1,1024,1)f32 #141=(1,1024,64)f32
145
- F.pad F.pad_86 1 1 141 142 mode=constant pad=(8,0) value=None $input=141 #141=(1,1024,64)f32 #142=(1,1024,72)f32
146
- nn.Conv1d conv1d_28 1 1 142 143 bias=True dilation=(1) groups=1 in_channels=1024 kernel_size=(16) out_channels=2048 padding=(0) padding_mode=zeros stride=(8) @bias=(2048)f32 @weight=(2048,1024,16)f32 $input=142 #142=(1,1024,72)f32 #143=(1,2048,8)f32
147
- F.pad F.pad_87 1 1 143 144 mode=constant pad=(2,0) value=None $input=143 #143=(1,2048,8)f32 #144=(1,2048,10)f32
148
- nn.Conv1d conv1d_29 1 1 144 145 bias=True dilation=(1) groups=1 in_channels=2048 kernel_size=(3) out_channels=64 padding=(0) padding_mode=zeros stride=(1) @bias=(64)f32 @weight=(64,2048,3)f32 $input=144 #144=(1,2048,10)f32 #145=(1,64,8)f32
149
- pnnx.Output pnnx_output_0 1 0 145 #145=(1,64,8)f32
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
audio_vae_encoder.pt DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:27af108b21c56faa3d22fc8f140adfe47cb6e8d54735b2ce1d1c5408cf38d618
3
- size 377232113
 
 
 
 
audio_vae_encoder_ncnn.py DELETED
@@ -1,26 +0,0 @@
1
- import numpy as np
2
- import ncnn
3
- import torch
4
-
5
- def test_inference():
6
- torch.manual_seed(0)
7
- in0 = torch.rand(1, 5120, dtype=torch.float)
8
- out = []
9
-
10
- with ncnn.Net() as net:
11
- net.load_param("/home/liyulin/Tools/voxcpm-ncnn/assets/voxcpm2/audio_vae_encoder.ncnn.param")
12
- net.load_model("/home/liyulin/Tools/voxcpm-ncnn/assets/voxcpm2/audio_vae_encoder.ncnn.bin")
13
-
14
- with net.create_extractor() as ex:
15
- ex.input("in0", ncnn.Mat(in0.squeeze(0).numpy()).clone())
16
-
17
- _, out0 = ex.extract("out0")
18
- out.append(torch.from_numpy(np.array(out0)).unsqueeze(0))
19
-
20
- if len(out) == 1:
21
- return out[0]
22
- else:
23
- return tuple(out)
24
-
25
- if __name__ == "__main__":
26
- print(test_inference())
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
audio_vae_encoder_pnnx.py DELETED
@@ -1,378 +0,0 @@
1
- # pnnx model stat
2
- # model inputshape = [1,5120]f32
3
- # FLOPS = 6.569G
4
- # memory OPS = 122.184M
5
-
6
- import os
7
- import numpy as np
8
- import tempfile, zipfile
9
- import torch
10
- import torch.nn as nn
11
- import torch.nn.functional as F
12
- try:
13
- import torchvision
14
- import torchaudio
15
- except:
16
- pass
17
-
18
- class Model(nn.Module):
19
- def __init__(self):
20
- super(Model, self).__init__()
21
-
22
- self.conv1d_0 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=1, kernel_size=(7,), out_channels=128, padding=(0,), padding_mode='zeros', stride=(1,))
23
- self.conv1d_1 = nn.Conv1d(bias=True, dilation=(1,), groups=128, in_channels=128, kernel_size=(7,), out_channels=128, padding=(0,), padding_mode='zeros', stride=(1,))
24
- self.padconv1d_0 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=128, kernel_size=(1,), out_channels=128, padding=(0,), padding_mode='zeros', stride=(1,))
25
- self.conv1d_3 = nn.Conv1d(bias=True, dilation=(3,), groups=128, in_channels=128, kernel_size=(7,), out_channels=128, padding=(0,), padding_mode='zeros', stride=(1,))
26
- self.padconv1d_1 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=128, kernel_size=(1,), out_channels=128, padding=(0,), padding_mode='zeros', stride=(1,))
27
- self.conv1d_5 = nn.Conv1d(bias=True, dilation=(9,), groups=128, in_channels=128, kernel_size=(7,), out_channels=128, padding=(0,), padding_mode='zeros', stride=(1,))
28
- self.padconv1d_2 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=128, kernel_size=(1,), out_channels=128, padding=(0,), padding_mode='zeros', stride=(1,))
29
- self.conv1d_7 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=128, kernel_size=(4,), out_channels=256, padding=(0,), padding_mode='zeros', stride=(2,))
30
- self.conv1d_8 = nn.Conv1d(bias=True, dilation=(1,), groups=256, in_channels=256, kernel_size=(7,), out_channels=256, padding=(0,), padding_mode='zeros', stride=(1,))
31
- self.padconv1d_3 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=256, kernel_size=(1,), out_channels=256, padding=(0,), padding_mode='zeros', stride=(1,))
32
- self.conv1d_10 = nn.Conv1d(bias=True, dilation=(3,), groups=256, in_channels=256, kernel_size=(7,), out_channels=256, padding=(0,), padding_mode='zeros', stride=(1,))
33
- self.padconv1d_4 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=256, kernel_size=(1,), out_channels=256, padding=(0,), padding_mode='zeros', stride=(1,))
34
- self.conv1d_12 = nn.Conv1d(bias=True, dilation=(9,), groups=256, in_channels=256, kernel_size=(7,), out_channels=256, padding=(0,), padding_mode='zeros', stride=(1,))
35
- self.padconv1d_5 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=256, kernel_size=(1,), out_channels=256, padding=(0,), padding_mode='zeros', stride=(1,))
36
- self.conv1d_14 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=256, kernel_size=(10,), out_channels=512, padding=(0,), padding_mode='zeros', stride=(5,))
37
- self.conv1d_15 = nn.Conv1d(bias=True, dilation=(1,), groups=512, in_channels=512, kernel_size=(7,), out_channels=512, padding=(0,), padding_mode='zeros', stride=(1,))
38
- self.padconv1d_6 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=512, kernel_size=(1,), out_channels=512, padding=(0,), padding_mode='zeros', stride=(1,))
39
- self.conv1d_17 = nn.Conv1d(bias=True, dilation=(3,), groups=512, in_channels=512, kernel_size=(7,), out_channels=512, padding=(0,), padding_mode='zeros', stride=(1,))
40
- self.padconv1d_7 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=512, kernel_size=(1,), out_channels=512, padding=(0,), padding_mode='zeros', stride=(1,))
41
- self.conv1d_19 = nn.Conv1d(bias=True, dilation=(9,), groups=512, in_channels=512, kernel_size=(7,), out_channels=512, padding=(0,), padding_mode='zeros', stride=(1,))
42
- self.padconv1d_8 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=512, kernel_size=(1,), out_channels=512, padding=(0,), padding_mode='zeros', stride=(1,))
43
- self.conv1d_21 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=512, kernel_size=(16,), out_channels=1024, padding=(0,), padding_mode='zeros', stride=(8,))
44
- self.conv1d_22 = nn.Conv1d(bias=True, dilation=(1,), groups=1024, in_channels=1024, kernel_size=(7,), out_channels=1024, padding=(0,), padding_mode='zeros', stride=(1,))
45
- self.padconv1d_9 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=1024, kernel_size=(1,), out_channels=1024, padding=(0,), padding_mode='zeros', stride=(1,))
46
- self.conv1d_24 = nn.Conv1d(bias=True, dilation=(3,), groups=1024, in_channels=1024, kernel_size=(7,), out_channels=1024, padding=(0,), padding_mode='zeros', stride=(1,))
47
- self.padconv1d_10 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=1024, kernel_size=(1,), out_channels=1024, padding=(0,), padding_mode='zeros', stride=(1,))
48
- self.conv1d_26 = nn.Conv1d(bias=True, dilation=(9,), groups=1024, in_channels=1024, kernel_size=(7,), out_channels=1024, padding=(0,), padding_mode='zeros', stride=(1,))
49
- self.padconv1d_11 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=1024, kernel_size=(1,), out_channels=1024, padding=(0,), padding_mode='zeros', stride=(1,))
50
- self.conv1d_28 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=1024, kernel_size=(16,), out_channels=2048, padding=(0,), padding_mode='zeros', stride=(8,))
51
- self.conv1d_29 = nn.Conv1d(bias=True, dilation=(1,), groups=1, in_channels=2048, kernel_size=(3,), out_channels=64, padding=(0,), padding_mode='zeros', stride=(1,))
52
-
53
- archive = zipfile.ZipFile('/home/liyulin/Tools/voxcpm-ncnn/assets/voxcpm2/audio_vae_encoder.pnnx.bin', 'r')
54
- self.conv1d_0.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_0.bias', (128), 'float32')
55
- self.conv1d_0.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_0.weight', (128,1,7), 'float32')
56
- self.conv1d_1.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_1.bias', (128), 'float32')
57
- self.conv1d_1.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_1.weight', (128,1,7), 'float32')
58
- self.padconv1d_0.bias = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_0.bias', (128), 'float32')
59
- self.padconv1d_0.weight = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_0.weight', (128,128,1), 'float32')
60
- self.conv1d_3.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_3.bias', (128), 'float32')
61
- self.conv1d_3.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_3.weight', (128,1,7), 'float32')
62
- self.padconv1d_1.bias = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_1.bias', (128), 'float32')
63
- self.padconv1d_1.weight = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_1.weight', (128,128,1), 'float32')
64
- self.conv1d_5.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_5.bias', (128), 'float32')
65
- self.conv1d_5.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_5.weight', (128,1,7), 'float32')
66
- self.padconv1d_2.bias = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_2.bias', (128), 'float32')
67
- self.padconv1d_2.weight = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_2.weight', (128,128,1), 'float32')
68
- self.conv1d_7.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_7.bias', (256), 'float32')
69
- self.conv1d_7.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_7.weight', (256,128,4), 'float32')
70
- self.conv1d_8.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_8.bias', (256), 'float32')
71
- self.conv1d_8.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_8.weight', (256,1,7), 'float32')
72
- self.padconv1d_3.bias = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_3.bias', (256), 'float32')
73
- self.padconv1d_3.weight = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_3.weight', (256,256,1), 'float32')
74
- self.conv1d_10.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_10.bias', (256), 'float32')
75
- self.conv1d_10.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_10.weight', (256,1,7), 'float32')
76
- self.padconv1d_4.bias = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_4.bias', (256), 'float32')
77
- self.padconv1d_4.weight = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_4.weight', (256,256,1), 'float32')
78
- self.conv1d_12.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_12.bias', (256), 'float32')
79
- self.conv1d_12.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_12.weight', (256,1,7), 'float32')
80
- self.padconv1d_5.bias = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_5.bias', (256), 'float32')
81
- self.padconv1d_5.weight = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_5.weight', (256,256,1), 'float32')
82
- self.conv1d_14.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_14.bias', (512), 'float32')
83
- self.conv1d_14.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_14.weight', (512,256,10), 'float32')
84
- self.conv1d_15.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_15.bias', (512), 'float32')
85
- self.conv1d_15.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_15.weight', (512,1,7), 'float32')
86
- self.padconv1d_6.bias = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_6.bias', (512), 'float32')
87
- self.padconv1d_6.weight = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_6.weight', (512,512,1), 'float32')
88
- self.conv1d_17.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_17.bias', (512), 'float32')
89
- self.conv1d_17.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_17.weight', (512,1,7), 'float32')
90
- self.padconv1d_7.bias = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_7.bias', (512), 'float32')
91
- self.padconv1d_7.weight = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_7.weight', (512,512,1), 'float32')
92
- self.conv1d_19.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_19.bias', (512), 'float32')
93
- self.conv1d_19.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_19.weight', (512,1,7), 'float32')
94
- self.padconv1d_8.bias = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_8.bias', (512), 'float32')
95
- self.padconv1d_8.weight = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_8.weight', (512,512,1), 'float32')
96
- self.conv1d_21.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_21.bias', (1024), 'float32')
97
- self.conv1d_21.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_21.weight', (1024,512,16), 'float32')
98
- self.conv1d_22.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_22.bias', (1024), 'float32')
99
- self.conv1d_22.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_22.weight', (1024,1,7), 'float32')
100
- self.padconv1d_9.bias = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_9.bias', (1024), 'float32')
101
- self.padconv1d_9.weight = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_9.weight', (1024,1024,1), 'float32')
102
- self.conv1d_24.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_24.bias', (1024), 'float32')
103
- self.conv1d_24.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_24.weight', (1024,1,7), 'float32')
104
- self.padconv1d_10.bias = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_10.bias', (1024), 'float32')
105
- self.padconv1d_10.weight = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_10.weight', (1024,1024,1), 'float32')
106
- self.conv1d_26.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_26.bias', (1024), 'float32')
107
- self.conv1d_26.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_26.weight', (1024,1,7), 'float32')
108
- self.padconv1d_11.bias = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_11.bias', (1024), 'float32')
109
- self.padconv1d_11.weight = self.load_pnnx_bin_as_parameter(archive, 'padconv1d_11.weight', (1024,1024,1), 'float32')
110
- self.conv1d_28.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_28.bias', (2048), 'float32')
111
- self.conv1d_28.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_28.weight', (2048,1024,16), 'float32')
112
- self.conv1d_29.bias = self.load_pnnx_bin_as_parameter(archive, 'conv1d_29.bias', (64), 'float32')
113
- self.conv1d_29.weight = self.load_pnnx_bin_as_parameter(archive, 'conv1d_29.weight', (64,2048,3), 'float32')
114
- self.audio_vae_encoder_block_1_block_0_block_0_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.1.block.0.block.0.data', (1,128,1,), 'float32')
115
- self.pnnx_fold_114_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_114.data', (1,128,1,), 'float32')
116
- self.audio_vae_encoder_block_1_block_0_block_2_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.1.block.0.block.2.data', (1,128,1,), 'float32')
117
- self.pnnx_fold_151_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_151.data', (1,128,1,), 'float32')
118
- self.audio_vae_encoder_block_1_block_1_block_0_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.1.block.1.block.0.data', (1,128,1,), 'float32')
119
- self.pnnx_fold_202_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_202.data', (1,128,1,), 'float32')
120
- self.audio_vae_encoder_block_1_block_1_block_2_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.1.block.1.block.2.data', (1,128,1,), 'float32')
121
- self.pnnx_fold_240_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_240.data', (1,128,1,), 'float32')
122
- self.audio_vae_encoder_block_1_block_2_block_0_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.1.block.2.block.0.data', (1,128,1,), 'float32')
123
- self.pnnx_fold_291_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_291.data', (1,128,1,), 'float32')
124
- self.audio_vae_encoder_block_1_block_2_block_2_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.1.block.2.block.2.data', (1,128,1,), 'float32')
125
- self.pnnx_fold_329_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_329.data', (1,128,1,), 'float32')
126
- self.audio_vae_encoder_block_1_block_3_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.1.block.3.data', (1,128,1,), 'float32')
127
- self.pnnx_fold_365_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_365.data', (1,128,1,), 'float32')
128
- self.audio_vae_encoder_block_2_block_0_block_0_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.2.block.0.block.0.data', (1,256,1,), 'float32')
129
- self.pnnx_fold_427_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_427.data', (1,256,1,), 'float32')
130
- self.audio_vae_encoder_block_2_block_0_block_2_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.2.block.0.block.2.data', (1,256,1,), 'float32')
131
- self.pnnx_fold_464_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_464.data', (1,256,1,), 'float32')
132
- self.audio_vae_encoder_block_2_block_1_block_0_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.2.block.1.block.0.data', (1,256,1,), 'float32')
133
- self.pnnx_fold_515_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_515.data', (1,256,1,), 'float32')
134
- self.audio_vae_encoder_block_2_block_1_block_2_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.2.block.1.block.2.data', (1,256,1,), 'float32')
135
- self.pnnx_fold_553_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_553.data', (1,256,1,), 'float32')
136
- self.audio_vae_encoder_block_2_block_2_block_0_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.2.block.2.block.0.data', (1,256,1,), 'float32')
137
- self.pnnx_fold_604_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_604.data', (1,256,1,), 'float32')
138
- self.audio_vae_encoder_block_2_block_2_block_2_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.2.block.2.block.2.data', (1,256,1,), 'float32')
139
- self.pnnx_fold_642_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_642.data', (1,256,1,), 'float32')
140
- self.audio_vae_encoder_block_2_block_3_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.2.block.3.data', (1,256,1,), 'float32')
141
- self.pnnx_fold_678_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_678.data', (1,256,1,), 'float32')
142
- self.audio_vae_encoder_block_3_block_0_block_0_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.3.block.0.block.0.data', (1,512,1,), 'float32')
143
- self.pnnx_fold_740_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_740.data', (1,512,1,), 'float32')
144
- self.audio_vae_encoder_block_3_block_0_block_2_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.3.block.0.block.2.data', (1,512,1,), 'float32')
145
- self.pnnx_fold_777_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_777.data', (1,512,1,), 'float32')
146
- self.audio_vae_encoder_block_3_block_1_block_0_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.3.block.1.block.0.data', (1,512,1,), 'float32')
147
- self.pnnx_fold_828_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_828.data', (1,512,1,), 'float32')
148
- self.audio_vae_encoder_block_3_block_1_block_2_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.3.block.1.block.2.data', (1,512,1,), 'float32')
149
- self.pnnx_fold_866_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_866.data', (1,512,1,), 'float32')
150
- self.audio_vae_encoder_block_3_block_2_block_0_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.3.block.2.block.0.data', (1,512,1,), 'float32')
151
- self.pnnx_fold_917_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_917.data', (1,512,1,), 'float32')
152
- self.audio_vae_encoder_block_3_block_2_block_2_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.3.block.2.block.2.data', (1,512,1,), 'float32')
153
- self.pnnx_fold_955_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_955.data', (1,512,1,), 'float32')
154
- self.audio_vae_encoder_block_3_block_3_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.3.block.3.data', (1,512,1,), 'float32')
155
- self.pnnx_fold_991_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_991.data', (1,512,1,), 'float32')
156
- self.audio_vae_encoder_block_4_block_0_block_0_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.4.block.0.block.0.data', (1,1024,1,), 'float32')
157
- self.pnnx_fold_1053_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_1053.data', (1,1024,1,), 'float32')
158
- self.audio_vae_encoder_block_4_block_0_block_2_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.4.block.0.block.2.data', (1,1024,1,), 'float32')
159
- self.pnnx_fold_1090_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_1090.data', (1,1024,1,), 'float32')
160
- self.audio_vae_encoder_block_4_block_1_block_0_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.4.block.1.block.0.data', (1,1024,1,), 'float32')
161
- self.pnnx_fold_1141_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_1141.data', (1,1024,1,), 'float32')
162
- self.audio_vae_encoder_block_4_block_1_block_2_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.4.block.1.block.2.data', (1,1024,1,), 'float32')
163
- self.pnnx_fold_1179_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_1179.data', (1,1024,1,), 'float32')
164
- self.audio_vae_encoder_block_4_block_2_block_0_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.4.block.2.block.0.data', (1,1024,1,), 'float32')
165
- self.pnnx_fold_1230_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_1230.data', (1,1024,1,), 'float32')
166
- self.audio_vae_encoder_block_4_block_2_block_2_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.4.block.2.block.2.data', (1,1024,1,), 'float32')
167
- self.pnnx_fold_1268_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_1268.data', (1,1024,1,), 'float32')
168
- self.audio_vae_encoder_block_4_block_3_data = self.load_pnnx_bin_as_parameter(archive, 'audio_vae.encoder.block.4.block.3.data', (1,1024,1,), 'float32')
169
- self.pnnx_fold_1304_data = self.load_pnnx_bin_as_parameter(archive, 'pnnx_fold_1304.data', (1,1024,1,), 'float32')
170
- archive.close()
171
-
172
- def load_pnnx_bin_as_parameter(self, archive, key, shape, dtype, requires_grad=True):
173
- return nn.Parameter(self.load_pnnx_bin_as_tensor(archive, key, shape, dtype), requires_grad)
174
-
175
- def load_pnnx_bin_as_tensor(self, archive, key, shape, dtype):
176
- fd, tmppath = tempfile.mkstemp()
177
- with os.fdopen(fd, 'wb') as tmpf, archive.open(key) as keyfile:
178
- tmpf.write(keyfile.read())
179
- m = np.memmap(tmppath, dtype=dtype, mode='r', shape=shape).copy()
180
- os.remove(tmppath)
181
- return torch.from_numpy(m)
182
-
183
- def forward(self, v_0):
184
- v_1 = v_0.unsqueeze(1)
185
- v_2 = F.pad(v_1, mode='constant', pad=(6,0), value=None)
186
- v_3 = self.conv1d_0(v_2)
187
- v_4 = self.audio_vae_encoder_block_1_block_0_block_0_data
188
- v_5 = self.pnnx_fold_114_data
189
- v_6 = (v_3 + (v_5 * torch.pow(torch.sin((v_4 * v_3)), 2)))
190
- v_7 = F.pad(v_6, mode='constant', pad=(6,0), value=None)
191
- v_8 = self.conv1d_1(v_7)
192
- v_9 = self.audio_vae_encoder_block_1_block_0_block_2_data
193
- v_10 = self.pnnx_fold_151_data
194
- v_11 = (v_8 + (v_10 * torch.pow(torch.sin((v_9 * v_8)), 2)))
195
- v_12 = self.padconv1d_0(v_11)
196
- v_13 = (v_3 + v_12)
197
- v_14 = self.audio_vae_encoder_block_1_block_1_block_0_data
198
- v_15 = self.pnnx_fold_202_data
199
- v_16 = (v_13 + (v_15 * torch.pow(torch.sin((v_14 * v_13)), 2)))
200
- v_17 = F.pad(v_16, mode='constant', pad=(18,0), value=None)
201
- v_18 = self.conv1d_3(v_17)
202
- v_19 = self.audio_vae_encoder_block_1_block_1_block_2_data
203
- v_20 = self.pnnx_fold_240_data
204
- v_21 = (v_18 + (v_20 * torch.pow(torch.sin((v_19 * v_18)), 2)))
205
- v_22 = self.padconv1d_1(v_21)
206
- v_23 = (v_13 + v_22)
207
- v_24 = self.audio_vae_encoder_block_1_block_2_block_0_data
208
- v_25 = self.pnnx_fold_291_data
209
- v_26 = (v_23 + (v_25 * torch.pow(torch.sin((v_24 * v_23)), 2)))
210
- v_27 = F.pad(v_26, mode='constant', pad=(54,0), value=None)
211
- v_28 = self.conv1d_5(v_27)
212
- v_29 = self.audio_vae_encoder_block_1_block_2_block_2_data
213
- v_30 = self.pnnx_fold_329_data
214
- v_31 = (v_28 + (v_30 * torch.pow(torch.sin((v_29 * v_28)), 2)))
215
- v_32 = self.padconv1d_2(v_31)
216
- v_33 = (v_23 + v_32)
217
- v_34 = self.audio_vae_encoder_block_1_block_3_data
218
- v_35 = self.pnnx_fold_365_data
219
- v_36 = (v_33 + (v_35 * torch.pow(torch.sin((v_34 * v_33)), 2)))
220
- v_37 = F.pad(v_36, mode='constant', pad=(2,0), value=None)
221
- v_38 = self.conv1d_7(v_37)
222
- v_39 = self.audio_vae_encoder_block_2_block_0_block_0_data
223
- v_40 = self.pnnx_fold_427_data
224
- v_41 = (v_38 + (v_40 * torch.pow(torch.sin((v_39 * v_38)), 2)))
225
- v_42 = F.pad(v_41, mode='constant', pad=(6,0), value=None)
226
- v_43 = self.conv1d_8(v_42)
227
- v_44 = self.audio_vae_encoder_block_2_block_0_block_2_data
228
- v_45 = self.pnnx_fold_464_data
229
- v_46 = (v_43 + (v_45 * torch.pow(torch.sin((v_44 * v_43)), 2)))
230
- v_47 = self.padconv1d_3(v_46)
231
- v_48 = (v_38 + v_47)
232
- v_49 = self.audio_vae_encoder_block_2_block_1_block_0_data
233
- v_50 = self.pnnx_fold_515_data
234
- v_51 = (v_48 + (v_50 * torch.pow(torch.sin((v_49 * v_48)), 2)))
235
- v_52 = F.pad(v_51, mode='constant', pad=(18,0), value=None)
236
- v_53 = self.conv1d_10(v_52)
237
- v_54 = self.audio_vae_encoder_block_2_block_1_block_2_data
238
- v_55 = self.pnnx_fold_553_data
239
- v_56 = (v_53 + (v_55 * torch.pow(torch.sin((v_54 * v_53)), 2)))
240
- v_57 = self.padconv1d_4(v_56)
241
- v_58 = (v_48 + v_57)
242
- v_59 = self.audio_vae_encoder_block_2_block_2_block_0_data
243
- v_60 = self.pnnx_fold_604_data
244
- v_61 = (v_58 + (v_60 * torch.pow(torch.sin((v_59 * v_58)), 2)))
245
- v_62 = F.pad(v_61, mode='constant', pad=(54,0), value=None)
246
- v_63 = self.conv1d_12(v_62)
247
- v_64 = self.audio_vae_encoder_block_2_block_2_block_2_data
248
- v_65 = self.pnnx_fold_642_data
249
- v_66 = (v_63 + (v_65 * torch.pow(torch.sin((v_64 * v_63)), 2)))
250
- v_67 = self.padconv1d_5(v_66)
251
- v_68 = (v_58 + v_67)
252
- v_69 = self.audio_vae_encoder_block_2_block_3_data
253
- v_70 = self.pnnx_fold_678_data
254
- v_71 = (v_68 + (v_70 * torch.pow(torch.sin((v_69 * v_68)), 2)))
255
- v_72 = F.pad(v_71, mode='constant', pad=(5,0), value=None)
256
- v_73 = self.conv1d_14(v_72)
257
- v_74 = self.audio_vae_encoder_block_3_block_0_block_0_data
258
- v_75 = self.pnnx_fold_740_data
259
- v_76 = (v_73 + (v_75 * torch.pow(torch.sin((v_74 * v_73)), 2)))
260
- v_77 = F.pad(v_76, mode='constant', pad=(6,0), value=None)
261
- v_78 = self.conv1d_15(v_77)
262
- v_79 = self.audio_vae_encoder_block_3_block_0_block_2_data
263
- v_80 = self.pnnx_fold_777_data
264
- v_81 = (v_78 + (v_80 * torch.pow(torch.sin((v_79 * v_78)), 2)))
265
- v_82 = self.padconv1d_6(v_81)
266
- v_83 = (v_73 + v_82)
267
- v_84 = self.audio_vae_encoder_block_3_block_1_block_0_data
268
- v_85 = self.pnnx_fold_828_data
269
- v_86 = (v_83 + (v_85 * torch.pow(torch.sin((v_84 * v_83)), 2)))
270
- v_87 = F.pad(v_86, mode='constant', pad=(18,0), value=None)
271
- v_88 = self.conv1d_17(v_87)
272
- v_89 = self.audio_vae_encoder_block_3_block_1_block_2_data
273
- v_90 = self.pnnx_fold_866_data
274
- v_91 = (v_88 + (v_90 * torch.pow(torch.sin((v_89 * v_88)), 2)))
275
- v_92 = self.padconv1d_7(v_91)
276
- v_93 = (v_83 + v_92)
277
- v_94 = self.audio_vae_encoder_block_3_block_2_block_0_data
278
- v_95 = self.pnnx_fold_917_data
279
- v_96 = (v_93 + (v_95 * torch.pow(torch.sin((v_94 * v_93)), 2)))
280
- v_97 = F.pad(v_96, mode='constant', pad=(54,0), value=None)
281
- v_98 = self.conv1d_19(v_97)
282
- v_99 = self.audio_vae_encoder_block_3_block_2_block_2_data
283
- v_100 = self.pnnx_fold_955_data
284
- v_101 = (v_98 + (v_100 * torch.pow(torch.sin((v_99 * v_98)), 2)))
285
- v_102 = self.padconv1d_8(v_101)
286
- v_103 = (v_93 + v_102)
287
- v_104 = self.audio_vae_encoder_block_3_block_3_data
288
- v_105 = self.pnnx_fold_991_data
289
- v_106 = (v_103 + (v_105 * torch.pow(torch.sin((v_104 * v_103)), 2)))
290
- v_107 = F.pad(v_106, mode='constant', pad=(8,0), value=None)
291
- v_108 = self.conv1d_21(v_107)
292
- v_109 = self.audio_vae_encoder_block_4_block_0_block_0_data
293
- v_110 = self.pnnx_fold_1053_data
294
- v_111 = (v_108 + (v_110 * torch.pow(torch.sin((v_109 * v_108)), 2)))
295
- v_112 = F.pad(v_111, mode='constant', pad=(6,0), value=None)
296
- v_113 = self.conv1d_22(v_112)
297
- v_114 = self.audio_vae_encoder_block_4_block_0_block_2_data
298
- v_115 = self.pnnx_fold_1090_data
299
- v_116 = (v_113 + (v_115 * torch.pow(torch.sin((v_114 * v_113)), 2)))
300
- v_117 = self.padconv1d_9(v_116)
301
- v_118 = (v_108 + v_117)
302
- v_119 = self.audio_vae_encoder_block_4_block_1_block_0_data
303
- v_120 = self.pnnx_fold_1141_data
304
- v_121 = (v_118 + (v_120 * torch.pow(torch.sin((v_119 * v_118)), 2)))
305
- v_122 = F.pad(v_121, mode='constant', pad=(18,0), value=None)
306
- v_123 = self.conv1d_24(v_122)
307
- v_124 = self.audio_vae_encoder_block_4_block_1_block_2_data
308
- v_125 = self.pnnx_fold_1179_data
309
- v_126 = (v_123 + (v_125 * torch.pow(torch.sin((v_124 * v_123)), 2)))
310
- v_127 = self.padconv1d_10(v_126)
311
- v_128 = (v_118 + v_127)
312
- v_129 = self.audio_vae_encoder_block_4_block_2_block_0_data
313
- v_130 = self.pnnx_fold_1230_data
314
- v_131 = (v_128 + (v_130 * torch.pow(torch.sin((v_129 * v_128)), 2)))
315
- v_132 = F.pad(v_131, mode='constant', pad=(54,0), value=None)
316
- v_133 = self.conv1d_26(v_132)
317
- v_134 = self.audio_vae_encoder_block_4_block_2_block_2_data
318
- v_135 = self.pnnx_fold_1268_data
319
- v_136 = (v_133 + (v_135 * torch.pow(torch.sin((v_134 * v_133)), 2)))
320
- v_137 = self.padconv1d_11(v_136)
321
- v_138 = (v_128 + v_137)
322
- v_139 = self.audio_vae_encoder_block_4_block_3_data
323
- v_140 = self.pnnx_fold_1304_data
324
- v_141 = (v_138 + (v_140 * torch.pow(torch.sin((v_139 * v_138)), 2)))
325
- v_142 = F.pad(v_141, mode='constant', pad=(8,0), value=None)
326
- v_143 = self.conv1d_28(v_142)
327
- v_144 = F.pad(v_143, mode='constant', pad=(2,0), value=None)
328
- v_145 = self.conv1d_29(v_144)
329
- return v_145
330
-
331
- def export_torchscript():
332
- net = Model()
333
- net.float()
334
- net.eval()
335
-
336
- torch.manual_seed(0)
337
- v_0 = torch.rand(1, 5120, dtype=torch.float)
338
-
339
- mod = torch.jit.trace(net, v_0)
340
- mod.save("/home/liyulin/Tools/voxcpm-ncnn/assets/voxcpm2/audio_vae_encoder_pnnx.py.pt")
341
-
342
- def export_onnx():
343
- net = Model()
344
- net.float()
345
- net.eval()
346
-
347
- torch.manual_seed(0)
348
- v_0 = torch.rand(1, 5120, dtype=torch.float)
349
-
350
- torch.onnx.export(net, v_0, "/home/liyulin/Tools/voxcpm-ncnn/assets/voxcpm2/audio_vae_encoder_pnnx.py.onnx", export_params=True, operator_export_type=torch.onnx.OperatorExportTypes.ONNX_ATEN_FALLBACK, opset_version=13, input_names=['in0'], output_names=['out0'])
351
-
352
- def export_pnnx():
353
- net = Model()
354
- net.float()
355
- net.eval()
356
-
357
- torch.manual_seed(0)
358
- v_0 = torch.rand(1, 5120, dtype=torch.float)
359
-
360
- import pnnx
361
- pnnx.export(net, "/home/liyulin/Tools/voxcpm-ncnn/assets/voxcpm2/audio_vae_encoder_pnnx.py.pt", v_0)
362
-
363
- def export_ncnn():
364
- export_pnnx()
365
-
366
- @torch.no_grad()
367
- def test_inference():
368
- net = Model()
369
- net.float()
370
- net.eval()
371
-
372
- torch.manual_seed(0)
373
- v_0 = torch.rand(1, 5120, dtype=torch.float)
374
-
375
- return net(v_0)
376
-
377
- if __name__ == "__main__":
378
- print(test_inference())
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
tokenization_voxcpm2.py DELETED
@@ -1,72 +0,0 @@
1
- """Custom tokenizer for VoxCPM2 that splits multi-character Chinese tokens.
2
-
3
- VoxCPM2 was trained with ``mask_multichar_chinese_tokens`` which splits
4
- multi-character Chinese tokens (e.g. "你好" -> ["你", "好"]) into individual
5
- character IDs before embedding. The base LlamaTokenizerFast produces
6
- multi-character Chinese tokens that the model has never seen during training,
7
- yielding garbled Chinese audio output in downstream inference frameworks.
8
-
9
- This module provides ``VoxCPM2Tokenizer`` which transparently applies the
10
- character splitting inside ``encode()`` and ``__call__()``, so any downstream
11
- consumer (vLLM, vLLM-Omni, Nano-vLLM, etc.) gets correct single-character
12
- IDs without code changes.
13
- """
14
-
15
- from transformers import LlamaTokenizerFast
16
-
17
-
18
- class VoxCPM2Tokenizer(LlamaTokenizerFast):
19
-
20
- def __init__(self, *args, **kwargs):
21
- super().__init__(*args, **kwargs)
22
- self._split_map = self._build_split_map()
23
-
24
- def _build_split_map(self) -> dict[int, list[int]]:
25
- vocab = self.get_vocab()
26
- split_map: dict[int, list[int]] = {}
27
- for token, tid in vocab.items():
28
- clean = token.replace("\u2581", "")
29
- if len(clean) >= 2 and all(self._is_cjk(c) for c in clean):
30
- char_ids = self.convert_tokens_to_ids(list(clean))
31
- if all(c != self.unk_token_id for c in char_ids):
32
- split_map[tid] = char_ids
33
- return split_map
34
-
35
- @staticmethod
36
- def _is_cjk(c: str) -> bool:
37
- return (
38
- "\u4e00" <= c <= "\u9fff"
39
- or "\u3400" <= c <= "\u4dbf"
40
- or "\uf900" <= c <= "\ufaff"
41
- or "\U00020000" <= c <= "\U0002a6df"
42
- )
43
-
44
- def _expand_ids(self, ids: list[int]) -> list[int]:
45
- result: list[int] = []
46
- for tid in ids:
47
- expansion = self._split_map.get(tid)
48
- if expansion is not None:
49
- result.extend(expansion)
50
- else:
51
- result.append(tid)
52
- return result
53
-
54
- def encode(self, text, *args, **kwargs):
55
- ids = super().encode(text, *args, **kwargs)
56
- return self._expand_ids(ids)
57
-
58
- def __call__(self, text, *args, **kwargs):
59
- result = super().__call__(text, *args, **kwargs)
60
- if hasattr(result, "input_ids"):
61
- ids = result["input_ids"]
62
- if isinstance(ids, list) and ids and isinstance(ids[0], list):
63
- result["input_ids"] = [self._expand_ids(x) for x in ids]
64
- if "attention_mask" in result:
65
- result["attention_mask"] = [
66
- [1] * len(x) for x in result["input_ids"]
67
- ]
68
- elif isinstance(ids, list):
69
- result["input_ids"] = self._expand_ids(ids)
70
- if "attention_mask" in result:
71
- result["attention_mask"] = [1] * len(result["input_ids"])
72
- return result