rzgar commited on
Commit
a5fd2ee
·
verified ·
1 Parent(s): bbc00b6

Upload 11 files

Browse files
.gitattributes CHANGED
@@ -46,3 +46,7 @@ Workflow/silent_audio/silent_3sec.wav filter=lfs diff=lfs merge=lfs -text
46
  Workflow/silent_audio/silent_4sec.wav filter=lfs diff=lfs merge=lfs -text
47
  Workflow/silent_audio/silent_5sec.wav filter=lfs diff=lfs merge=lfs -text
48
  video/ComfyUI__00003-audio.mp4 filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
46
  Workflow/silent_audio/silent_4sec.wav filter=lfs diff=lfs merge=lfs -text
47
  Workflow/silent_audio/silent_5sec.wav filter=lfs diff=lfs merge=lfs -text
48
  video/ComfyUI__00003-audio.mp4 filter=lfs diff=lfs merge=lfs -text
49
+ ComfyUI-WanBerniniS2V_v2/demo_assets/ComfyUI_00207-audio.mp4 filter=lfs diff=lfs merge=lfs -text
50
+ ComfyUI-WanBerniniS2V_v2/demo_assets/dude.jpg filter=lfs diff=lfs merge=lfs -text
51
+ ComfyUI-WanBerniniS2V_v2/demo_assets/I[[:space:]]am[[:space:]]The[[:space:]]Dude[[:space:]]Playing[[:space:]]The[[:space:]]Dude.wav filter=lfs diff=lfs merge=lfs -text
52
+ ComfyUI-WanBerniniS2V_v2/demo_assets/Im[[:space:]]the[[:space:]]Dude.wav filter=lfs diff=lfs merge=lfs -text
ComfyUI-WanBerniniS2V_v2/ComfyUI-WanBerniniS2V_v2.zip ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:790d32a15fe19a572ab199ff9c9f7ae0c3099ba3169f3c9a51538a723700fa84
3
+ size 2524040
ComfyUI-WanBerniniS2V_v2/Workflow/Bernini-R-S2V-Workflow-V2.json ADDED
@@ -0,0 +1,2857 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "id": "ea7ba43e-ab56-40d4-a298-423f5bfe4462",
3
+ "revision": 0,
4
+ "last_node_id": 279,
5
+ "last_link_id": 758,
6
+ "nodes": [
7
+ {
8
+ "id": 80,
9
+ "type": "VAEDecode",
10
+ "pos": [
11
+ 1210,
12
+ 170
13
+ ],
14
+ "size": [
15
+ 210,
16
+ 50
17
+ ],
18
+ "flags": {},
19
+ "order": 41,
20
+ "mode": 0,
21
+ "inputs": [
22
+ {
23
+ "name": "samples",
24
+ "type": "LATENT",
25
+ "link": 533
26
+ },
27
+ {
28
+ "name": "vae",
29
+ "type": "VAE",
30
+ "link": 285
31
+ }
32
+ ],
33
+ "outputs": [
34
+ {
35
+ "name": "IMAGE",
36
+ "type": "IMAGE",
37
+ "links": [
38
+ 580
39
+ ]
40
+ }
41
+ ],
42
+ "properties": {
43
+ "cnr_id": "comfy-core",
44
+ "ver": "0.3.54",
45
+ "Node name for S&R": "VAEDecode"
46
+ },
47
+ "widgets_values": []
48
+ },
49
+ {
50
+ "id": 38,
51
+ "type": "CLIPLoader",
52
+ "pos": [
53
+ -530,
54
+ 470
55
+ ],
56
+ "size": [
57
+ 390,
58
+ 110
59
+ ],
60
+ "flags": {},
61
+ "order": 0,
62
+ "mode": 0,
63
+ "inputs": [],
64
+ "outputs": [
65
+ {
66
+ "name": "CLIP",
67
+ "type": "CLIP",
68
+ "slot_index": 0,
69
+ "links": [
70
+ 74,
71
+ 75
72
+ ]
73
+ }
74
+ ],
75
+ "properties": {
76
+ "cnr_id": "comfy-core",
77
+ "ver": "0.3.54",
78
+ "Node name for S&R": "CLIPLoader",
79
+ "models": [
80
+ {
81
+ "name": "umt5_xxl_fp8_e4m3fn_scaled.safetensors",
82
+ "url": "https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/resolve/main/split_files/text_encoders/umt5_xxl_fp8_e4m3fn_scaled.safetensors",
83
+ "directory": "text_encoders"
84
+ }
85
+ ]
86
+ },
87
+ "widgets_values": [
88
+ "umt5_xxl_fp8_e4m3fn_scaled.safetensors",
89
+ "wan",
90
+ "default"
91
+ ]
92
+ },
93
+ {
94
+ "id": 201,
95
+ "type": "RIFEInterpolation",
96
+ "pos": [
97
+ 1480,
98
+ 170
99
+ ],
100
+ "size": [
101
+ 270,
102
+ 180
103
+ ],
104
+ "flags": {},
105
+ "order": 42,
106
+ "mode": 0,
107
+ "inputs": [
108
+ {
109
+ "name": "images",
110
+ "type": "IMAGE",
111
+ "link": 580
112
+ }
113
+ ],
114
+ "outputs": [
115
+ {
116
+ "name": "images",
117
+ "type": "IMAGE",
118
+ "links": [
119
+ 629
120
+ ]
121
+ }
122
+ ],
123
+ "properties": {
124
+ "aux_id": "GACLove/ComfyUI-VFI",
125
+ "ver": "6176a430f12cd16003f4664c1e3c6af8e96cc3c6",
126
+ "Node name for S&R": "RIFEInterpolation"
127
+ },
128
+ "widgets_values": [
129
+ 16,
130
+ 24,
131
+ 2,
132
+ "flownet.pkl",
133
+ 4,
134
+ true
135
+ ]
136
+ },
137
+ {
138
+ "id": 210,
139
+ "type": "LoraLoaderModelOnly",
140
+ "pos": [
141
+ 1560,
142
+ 450
143
+ ],
144
+ "size": [
145
+ 360,
146
+ 90
147
+ ],
148
+ "flags": {},
149
+ "order": 18,
150
+ "mode": 0,
151
+ "inputs": [
152
+ {
153
+ "name": "model",
154
+ "type": "MODEL",
155
+ "link": 586
156
+ }
157
+ ],
158
+ "outputs": [
159
+ {
160
+ "name": "MODEL",
161
+ "type": "MODEL",
162
+ "links": [
163
+ 581
164
+ ]
165
+ }
166
+ ],
167
+ "properties": {
168
+ "cnr_id": "comfy-core",
169
+ "ver": "0.3.54",
170
+ "Node name for S&R": "LoraLoaderModelOnly",
171
+ "models": [
172
+ {
173
+ "name": "wan2.2_t2v_lightx2v_4steps_lora_v1.1_high_noise.safetensors",
174
+ "url": "https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/loras/wan2.2_t2v_lightx2v_4steps_lora_v1.1_high_noise.safetensors",
175
+ "directory": "loras"
176
+ }
177
+ ]
178
+ },
179
+ "widgets_values": [
180
+ "Wan-Speed/Bernini-R/Bernini-R_LightX2V_low_noise.safetensors",
181
+ 1
182
+ ]
183
+ },
184
+ {
185
+ "id": 54,
186
+ "type": "ModelSamplingSD3",
187
+ "pos": [
188
+ 1140,
189
+ 590
190
+ ],
191
+ "size": [
192
+ 370,
193
+ 60
194
+ ],
195
+ "flags": {},
196
+ "order": 25,
197
+ "mode": 4,
198
+ "inputs": [
199
+ {
200
+ "name": "model",
201
+ "type": "MODEL",
202
+ "link": 363
203
+ }
204
+ ],
205
+ "outputs": [
206
+ {
207
+ "name": "MODEL",
208
+ "type": "MODEL",
209
+ "slot_index": 0,
210
+ "links": [
211
+ 522
212
+ ]
213
+ }
214
+ ],
215
+ "properties": {
216
+ "cnr_id": "comfy-core",
217
+ "ver": "0.3.54",
218
+ "Node name for S&R": "ModelSamplingSD3"
219
+ },
220
+ "widgets_values": [
221
+ 8
222
+ ]
223
+ },
224
+ {
225
+ "id": 214,
226
+ "type": "ModelSamplingSD3",
227
+ "pos": [
228
+ 1560,
229
+ 590
230
+ ],
231
+ "size": [
232
+ 360,
233
+ 60
234
+ ],
235
+ "flags": {},
236
+ "order": 26,
237
+ "mode": 4,
238
+ "inputs": [
239
+ {
240
+ "name": "model",
241
+ "type": "MODEL",
242
+ "link": 581
243
+ }
244
+ ],
245
+ "outputs": [
246
+ {
247
+ "name": "MODEL",
248
+ "type": "MODEL",
249
+ "slot_index": 0,
250
+ "links": [
251
+ 582
252
+ ]
253
+ }
254
+ ],
255
+ "properties": {
256
+ "cnr_id": "comfy-core",
257
+ "ver": "0.3.54",
258
+ "Node name for S&R": "ModelSamplingSD3"
259
+ },
260
+ "widgets_values": [
261
+ 8
262
+ ]
263
+ },
264
+ {
265
+ "id": 207,
266
+ "type": "KSamplerAdvanced",
267
+ "pos": [
268
+ 1560,
269
+ 690
270
+ ],
271
+ "size": [
272
+ 360,
273
+ 340
274
+ ],
275
+ "flags": {},
276
+ "order": 40,
277
+ "mode": 0,
278
+ "inputs": [
279
+ {
280
+ "name": "model",
281
+ "type": "MODEL",
282
+ "link": 582
283
+ },
284
+ {
285
+ "name": "positive",
286
+ "type": "CONDITIONING",
287
+ "link": 741
288
+ },
289
+ {
290
+ "name": "negative",
291
+ "type": "CONDITIONING",
292
+ "link": 743
293
+ },
294
+ {
295
+ "name": "latent_image",
296
+ "type": "LATENT",
297
+ "link": 526
298
+ }
299
+ ],
300
+ "outputs": [
301
+ {
302
+ "name": "LATENT",
303
+ "type": "LATENT",
304
+ "links": [
305
+ 533
306
+ ]
307
+ }
308
+ ],
309
+ "properties": {
310
+ "cnr_id": "comfy-core",
311
+ "ver": "0.27.0",
312
+ "Node name for S&R": "KSamplerAdvanced"
313
+ },
314
+ "widgets_values": [
315
+ "disable",
316
+ 0,
317
+ "fixed",
318
+ 4,
319
+ 1,
320
+ "dpmpp_2m_sde",
321
+ "sgm_uniform",
322
+ 2,
323
+ 4,
324
+ "disable"
325
+ ]
326
+ },
327
+ {
328
+ "id": 250,
329
+ "type": "StringConcatenate",
330
+ "pos": [
331
+ -520,
332
+ 930
333
+ ],
334
+ "size": [
335
+ 400,
336
+ 250
337
+ ],
338
+ "flags": {},
339
+ "order": 24,
340
+ "mode": 0,
341
+ "inputs": [
342
+ {
343
+ "name": "string_a",
344
+ "type": "STRING",
345
+ "widget": {
346
+ "name": "string_a"
347
+ },
348
+ "link": 662
349
+ },
350
+ {
351
+ "name": "string_b",
352
+ "type": "STRING",
353
+ "widget": {
354
+ "name": "string_b"
355
+ },
356
+ "link": 664
357
+ }
358
+ ],
359
+ "outputs": [
360
+ {
361
+ "name": "STRING",
362
+ "type": "STRING",
363
+ "links": [
364
+ 663
365
+ ]
366
+ }
367
+ ],
368
+ "properties": {
369
+ "cnr_id": "comfy-core",
370
+ "ver": "0.24.0",
371
+ "Node name for S&R": "StringConcatenate"
372
+ },
373
+ "widgets_values": [
374
+ "",
375
+ "",
376
+ "."
377
+ ]
378
+ },
379
+ {
380
+ "id": 6,
381
+ "type": "CLIPTextEncode",
382
+ "pos": [
383
+ -50,
384
+ 750
385
+ ],
386
+ "size": [
387
+ 670,
388
+ 280
389
+ ],
390
+ "flags": {
391
+ "collapsed": true
392
+ },
393
+ "order": 30,
394
+ "mode": 0,
395
+ "inputs": [
396
+ {
397
+ "name": "clip",
398
+ "type": "CLIP",
399
+ "link": 74
400
+ },
401
+ {
402
+ "name": "text",
403
+ "type": "STRING",
404
+ "widget": {
405
+ "name": "text"
406
+ },
407
+ "link": 663
408
+ }
409
+ ],
410
+ "outputs": [
411
+ {
412
+ "name": "CONDITIONING",
413
+ "type": "CONDITIONING",
414
+ "slot_index": 0,
415
+ "links": [
416
+ 725
417
+ ]
418
+ }
419
+ ],
420
+ "title": "CLIP Text Encode (Positive Prompt)",
421
+ "properties": {
422
+ "cnr_id": "comfy-core",
423
+ "ver": "0.3.54",
424
+ "Node name for S&R": "CLIPTextEncode"
425
+ },
426
+ "widgets_values": [
427
+ ""
428
+ ],
429
+ "color": "#232",
430
+ "bgcolor": "#353"
431
+ },
432
+ {
433
+ "id": 251,
434
+ "type": "CustomCombo",
435
+ "pos": [
436
+ -1020,
437
+ 760
438
+ ],
439
+ "size": [
440
+ 460,
441
+ 390
442
+ ],
443
+ "flags": {},
444
+ "order": 1,
445
+ "mode": 0,
446
+ "inputs": [],
447
+ "outputs": [
448
+ {
449
+ "name": "STRING",
450
+ "type": "STRING",
451
+ "links": []
452
+ },
453
+ {
454
+ "name": "INDEX",
455
+ "type": "INT",
456
+ "links": [
457
+ 661
458
+ ]
459
+ }
460
+ ],
461
+ "properties": {
462
+ "cnr_id": "comfy-core",
463
+ "ver": "0.24.0",
464
+ "Node name for S&R": "CustomCombo"
465
+ },
466
+ "widgets_values": [
467
+ "Image to Video",
468
+ 5,
469
+ "Default",
470
+ "Text to Image",
471
+ "Text to Video",
472
+ "Image Editing",
473
+ "Subject to Image",
474
+ "Image to Video",
475
+ "Video Editing",
476
+ "Video Editing (Content Propagation)",
477
+ "Video Editing with Reference",
478
+ "Ads / Content Insertion",
479
+ "Video Editing (Action / Position)",
480
+ "Video Editing (Style / Motion)",
481
+ ""
482
+ ]
483
+ },
484
+ {
485
+ "id": 249,
486
+ "type": "6da2792c-a953-4960-a53a-ff1b5df6fe09",
487
+ "pos": [
488
+ -530,
489
+ 750
490
+ ],
491
+ "size": [
492
+ 410,
493
+ 150
494
+ ],
495
+ "flags": {},
496
+ "order": 15,
497
+ "mode": 0,
498
+ "inputs": [
499
+ {
500
+ "name": "index",
501
+ "type": "INT",
502
+ "widget": {
503
+ "name": "index"
504
+ },
505
+ "link": 661
506
+ }
507
+ ],
508
+ "outputs": [
509
+ {
510
+ "name": "selected_line",
511
+ "type": "STRING",
512
+ "links": [
513
+ 662
514
+ ]
515
+ }
516
+ ],
517
+ "properties": {
518
+ "proxyWidgets": [
519
+ [
520
+ "2",
521
+ "string"
522
+ ],
523
+ [
524
+ "248",
525
+ "value"
526
+ ]
527
+ ],
528
+ "cnr_id": "comfy-core",
529
+ "ver": "0.19.0",
530
+ "ue_properties": {
531
+ "widget_ue_connectable": {},
532
+ "input_ue_unconnectable": {}
533
+ }
534
+ },
535
+ "widgets_values": []
536
+ },
537
+ {
538
+ "id": 107,
539
+ "type": "LoraLoaderModelOnly",
540
+ "pos": [
541
+ 1140,
542
+ 450
543
+ ],
544
+ "size": [
545
+ 380,
546
+ 90
547
+ ],
548
+ "flags": {},
549
+ "order": 17,
550
+ "mode": 0,
551
+ "inputs": [
552
+ {
553
+ "name": "model",
554
+ "type": "MODEL",
555
+ "link": 585
556
+ }
557
+ ],
558
+ "outputs": [
559
+ {
560
+ "name": "MODEL",
561
+ "type": "MODEL",
562
+ "links": [
563
+ 363
564
+ ]
565
+ }
566
+ ],
567
+ "properties": {
568
+ "cnr_id": "comfy-core",
569
+ "ver": "0.3.54",
570
+ "Node name for S&R": "LoraLoaderModelOnly",
571
+ "models": [
572
+ {
573
+ "name": "wan2.2_t2v_lightx2v_4steps_lora_v1.1_high_noise.safetensors",
574
+ "url": "https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/loras/wan2.2_t2v_lightx2v_4steps_lora_v1.1_high_noise.safetensors",
575
+ "directory": "loras"
576
+ }
577
+ ]
578
+ },
579
+ "widgets_values": [
580
+ "Wan-Speed/Bernini-R/Bernini-R_LightX2V_high_noise.safetensors",
581
+ 1
582
+ ]
583
+ },
584
+ {
585
+ "id": 205,
586
+ "type": "KSamplerAdvanced",
587
+ "pos": [
588
+ 1130,
589
+ 710
590
+ ],
591
+ "size": [
592
+ 370,
593
+ 340
594
+ ],
595
+ "flags": {},
596
+ "order": 38,
597
+ "mode": 0,
598
+ "inputs": [
599
+ {
600
+ "name": "model",
601
+ "type": "MODEL",
602
+ "link": 522
603
+ },
604
+ {
605
+ "name": "positive",
606
+ "type": "CONDITIONING",
607
+ "link": 740
608
+ },
609
+ {
610
+ "name": "negative",
611
+ "type": "CONDITIONING",
612
+ "link": 742
613
+ },
614
+ {
615
+ "name": "latent_image",
616
+ "type": "LATENT",
617
+ "link": 744
618
+ }
619
+ ],
620
+ "outputs": [
621
+ {
622
+ "name": "LATENT",
623
+ "type": "LATENT",
624
+ "links": [
625
+ 526
626
+ ]
627
+ }
628
+ ],
629
+ "properties": {
630
+ "cnr_id": "comfy-core",
631
+ "ver": "0.27.0",
632
+ "Node name for S&R": "KSamplerAdvanced"
633
+ },
634
+ "widgets_values": [
635
+ "enable",
636
+ 780018084687208,
637
+ "randomize",
638
+ 4,
639
+ 1,
640
+ "dpmpp_2m_sde",
641
+ "sgm_uniform",
642
+ 0,
643
+ 2,
644
+ "enable"
645
+ ]
646
+ },
647
+ {
648
+ "id": 271,
649
+ "type": "Reroute",
650
+ "pos": [
651
+ -360,
652
+ 1610
653
+ ],
654
+ "size": [
655
+ 75,
656
+ 26
657
+ ],
658
+ "flags": {},
659
+ "order": 29,
660
+ "mode": 0,
661
+ "inputs": [
662
+ {
663
+ "name": "",
664
+ "type": "*",
665
+ "link": 737
666
+ }
667
+ ],
668
+ "outputs": [
669
+ {
670
+ "name": "",
671
+ "type": "AUDIO",
672
+ "links": [
673
+ 738,
674
+ 739
675
+ ]
676
+ }
677
+ ],
678
+ "properties": {
679
+ "showOutputText": false,
680
+ "horizontal": false
681
+ },
682
+ "color": "#323",
683
+ "bgcolor": "#535"
684
+ },
685
+ {
686
+ "id": 39,
687
+ "type": "VAELoader",
688
+ "pos": [
689
+ 530,
690
+ 730
691
+ ],
692
+ "size": [
693
+ 210,
694
+ 60
695
+ ],
696
+ "flags": {
697
+ "collapsed": true
698
+ },
699
+ "order": 2,
700
+ "mode": 0,
701
+ "inputs": [],
702
+ "outputs": [
703
+ {
704
+ "name": "VAE",
705
+ "type": "VAE",
706
+ "slot_index": 0,
707
+ "links": [
708
+ 285,
709
+ 727
710
+ ]
711
+ }
712
+ ],
713
+ "properties": {
714
+ "cnr_id": "comfy-core",
715
+ "ver": "0.3.54",
716
+ "Node name for S&R": "VAELoader",
717
+ "models": [
718
+ {
719
+ "name": "wan_2.1_vae.safetensors",
720
+ "url": "https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/vae/wan_2.1_vae.safetensors",
721
+ "directory": "vae"
722
+ }
723
+ ]
724
+ },
725
+ "widgets_values": [
726
+ "Wan/Wan2_1_VAE_fp32.safetensors"
727
+ ]
728
+ },
729
+ {
730
+ "id": 235,
731
+ "type": "Reroute",
732
+ "pos": [
733
+ -350,
734
+ 1410
735
+ ],
736
+ "size": [
737
+ 75,
738
+ 26
739
+ ],
740
+ "flags": {},
741
+ "order": 28,
742
+ "mode": 0,
743
+ "inputs": [
744
+ {
745
+ "name": "",
746
+ "type": "*",
747
+ "link": 736
748
+ }
749
+ ],
750
+ "outputs": [
751
+ {
752
+ "name": "",
753
+ "type": "AUDIO",
754
+ "links": [
755
+ 622,
756
+ 717
757
+ ]
758
+ }
759
+ ],
760
+ "properties": {
761
+ "showOutputText": false,
762
+ "horizontal": false
763
+ },
764
+ "color": "#323",
765
+ "bgcolor": "#535"
766
+ },
767
+ {
768
+ "id": 270,
769
+ "type": "Reroute",
770
+ "pos": [
771
+ -450,
772
+ 1290
773
+ ],
774
+ "size": [
775
+ 75,
776
+ 26
777
+ ],
778
+ "flags": {},
779
+ "order": 19,
780
+ "mode": 0,
781
+ "inputs": [
782
+ {
783
+ "name": "",
784
+ "type": "*",
785
+ "link": 733
786
+ }
787
+ ],
788
+ "outputs": [
789
+ {
790
+ "name": "",
791
+ "type": "AUDIO_ENCODER",
792
+ "links": [
793
+ 734,
794
+ 735
795
+ ]
796
+ }
797
+ ],
798
+ "properties": {
799
+ "showOutputText": false,
800
+ "horizontal": false
801
+ }
802
+ },
803
+ {
804
+ "id": 241,
805
+ "type": "LoadImage",
806
+ "pos": [
807
+ -40,
808
+ 840
809
+ ],
810
+ "size": [
811
+ 290,
812
+ 320
813
+ ],
814
+ "flags": {},
815
+ "order": 3,
816
+ "mode": 0,
817
+ "inputs": [],
818
+ "outputs": [
819
+ {
820
+ "name": "IMAGE",
821
+ "type": "IMAGE",
822
+ "links": [
823
+ 731
824
+ ]
825
+ },
826
+ {
827
+ "name": "MASK",
828
+ "type": "MASK",
829
+ "links": [
830
+ 702,
831
+ 729
832
+ ]
833
+ }
834
+ ],
835
+ "title": "image0 & mask_1",
836
+ "properties": {
837
+ "cnr_id": "comfy-core",
838
+ "ver": "0.27.0",
839
+ "Node name for S&R": "LoadImage",
840
+ "image": "clipspace/clipspace-painted-masked-1783665396446.png [input]"
841
+ },
842
+ "widgets_values": [
843
+ "clipspace/clipspace-painted-masked-1783665396446.png [input]",
844
+ "image"
845
+ ]
846
+ },
847
+ {
848
+ "id": 265,
849
+ "type": "InvertMask",
850
+ "pos": [
851
+ 240,
852
+ 790
853
+ ],
854
+ "size": [
855
+ 140,
856
+ 30
857
+ ],
858
+ "flags": {
859
+ "collapsed": true
860
+ },
861
+ "order": 16,
862
+ "mode": 4,
863
+ "inputs": [
864
+ {
865
+ "name": "mask",
866
+ "type": "MASK",
867
+ "link": 702
868
+ }
869
+ ],
870
+ "outputs": [
871
+ {
872
+ "name": "MASK",
873
+ "type": "MASK",
874
+ "links": []
875
+ }
876
+ ],
877
+ "properties": {
878
+ "cnr_id": "comfy-core",
879
+ "ver": "0.27.0",
880
+ "Node name for S&R": "InvertMask"
881
+ },
882
+ "widgets_values": []
883
+ },
884
+ {
885
+ "id": 7,
886
+ "type": "CLIPTextEncode",
887
+ "pos": [
888
+ 10,
889
+ 500
890
+ ],
891
+ "size": [
892
+ 650,
893
+ 180
894
+ ],
895
+ "flags": {
896
+ "collapsed": true
897
+ },
898
+ "order": 14,
899
+ "mode": 0,
900
+ "inputs": [
901
+ {
902
+ "name": "clip",
903
+ "type": "CLIP",
904
+ "link": 75
905
+ }
906
+ ],
907
+ "outputs": [
908
+ {
909
+ "name": "CONDITIONING",
910
+ "type": "CONDITIONING",
911
+ "slot_index": 0,
912
+ "links": [
913
+ 726
914
+ ]
915
+ }
916
+ ],
917
+ "title": "CLIP Text Encode (Negative Prompt)",
918
+ "properties": {
919
+ "cnr_id": "comfy-core",
920
+ "ver": "0.3.54",
921
+ "Node name for S&R": "CLIPTextEncode"
922
+ },
923
+ "widgets_values": [
924
+ "Vivid color tone, overexposed, static, unclear details, subtitles, style, artwork, painting, image, still, overall grayish, worst quality, low quality, leftover JPEG compression artifacts, ugly, incomplete, missing parts, extra fingers, poorly drawn hands, poorly drawn face, disfigured, malformed body parts, fused fingers, a completely motionless image, messy background, three legs, many people in the background, walking backward."
925
+ ],
926
+ "color": "#223",
927
+ "bgcolor": "#335"
928
+ },
929
+ {
930
+ "id": 37,
931
+ "type": "UNETLoader",
932
+ "pos": [
933
+ -530,
934
+ 180
935
+ ],
936
+ "size": [
937
+ 390,
938
+ 90
939
+ ],
940
+ "flags": {},
941
+ "order": 4,
942
+ "mode": 0,
943
+ "inputs": [],
944
+ "outputs": [
945
+ {
946
+ "name": "MODEL",
947
+ "type": "MODEL",
948
+ "slot_index": 0,
949
+ "links": [
950
+ 585
951
+ ]
952
+ }
953
+ ],
954
+ "properties": {
955
+ "cnr_id": "comfy-core",
956
+ "ver": "0.3.54",
957
+ "Node name for S&R": "UNETLoader",
958
+ "models": [
959
+ {
960
+ "name": "wan2.2_s2v_14B_fp8_scaled.safetensors",
961
+ "url": "https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/diffusion_models/wan2.2_s2v_14B_fp8_scaled.safetensors",
962
+ "directory": "diffusion_models"
963
+ }
964
+ ]
965
+ },
966
+ "widgets_values": [
967
+ "Bernini-R_fp8/Bernini-R-S2V-FP8/wan2.2_bernini_r_high_noise_fp8_scaled_s2v.safetensors",
968
+ "default"
969
+ ]
970
+ },
971
+ {
972
+ "id": 204,
973
+ "type": "UNETLoader",
974
+ "pos": [
975
+ -530,
976
+ 320
977
+ ],
978
+ "size": [
979
+ 390,
980
+ 90
981
+ ],
982
+ "flags": {},
983
+ "order": 5,
984
+ "mode": 0,
985
+ "inputs": [],
986
+ "outputs": [
987
+ {
988
+ "name": "MODEL",
989
+ "type": "MODEL",
990
+ "slot_index": 0,
991
+ "links": [
992
+ 586
993
+ ]
994
+ }
995
+ ],
996
+ "properties": {
997
+ "cnr_id": "comfy-core",
998
+ "ver": "0.3.54",
999
+ "Node name for S&R": "UNETLoader",
1000
+ "models": [
1001
+ {
1002
+ "name": "wan2.2_s2v_14B_fp8_scaled.safetensors",
1003
+ "url": "https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/diffusion_models/wan2.2_s2v_14B_fp8_scaled.safetensors",
1004
+ "directory": "diffusion_models"
1005
+ }
1006
+ ]
1007
+ },
1008
+ "widgets_values": [
1009
+ "Bernini-R_fp8/Bernini-R-S2V-FP8/wan2.2_bernini_r_low_noise_fp8_scaled_s2v.safetensors",
1010
+ "default"
1011
+ ]
1012
+ },
1013
+ {
1014
+ "id": 267,
1015
+ "type": "AudioConcat",
1016
+ "pos": [
1017
+ -210,
1018
+ 1290
1019
+ ],
1020
+ "size": [
1021
+ 290,
1022
+ 80
1023
+ ],
1024
+ "flags": {},
1025
+ "order": 34,
1026
+ "mode": 0,
1027
+ "inputs": [
1028
+ {
1029
+ "name": "audio1",
1030
+ "type": "AUDIO",
1031
+ "link": 717
1032
+ },
1033
+ {
1034
+ "name": "audio2",
1035
+ "type": "AUDIO",
1036
+ "link": 739
1037
+ }
1038
+ ],
1039
+ "outputs": [
1040
+ {
1041
+ "name": "AUDIO",
1042
+ "type": "AUDIO",
1043
+ "links": [
1044
+ 745,
1045
+ 754
1046
+ ]
1047
+ }
1048
+ ],
1049
+ "title": "audio_1 + audio_2",
1050
+ "properties": {
1051
+ "cnr_id": "comfy-core",
1052
+ "ver": "0.27.0",
1053
+ "Node name for S&R": "AudioConcat"
1054
+ },
1055
+ "widgets_values": [
1056
+ "after"
1057
+ ]
1058
+ },
1059
+ {
1060
+ "id": 238,
1061
+ "type": "ComfyMathExpression",
1062
+ "pos": [
1063
+ 1110,
1064
+ 1110
1065
+ ],
1066
+ "size": [
1067
+ 400,
1068
+ 200
1069
+ ],
1070
+ "flags": {
1071
+ "collapsed": true
1072
+ },
1073
+ "order": 21,
1074
+ "mode": 0,
1075
+ "inputs": [
1076
+ {
1077
+ "label": "a",
1078
+ "name": "values.a",
1079
+ "type": "FLOAT,INT,BOOLEAN",
1080
+ "link": 627
1081
+ },
1082
+ {
1083
+ "label": "b",
1084
+ "name": "values.b",
1085
+ "shape": 7,
1086
+ "type": "FLOAT,INT,BOOLEAN",
1087
+ "link": null
1088
+ }
1089
+ ],
1090
+ "outputs": [
1091
+ {
1092
+ "name": "FLOAT",
1093
+ "type": "FLOAT",
1094
+ "links": null
1095
+ },
1096
+ {
1097
+ "name": "INT",
1098
+ "type": "INT",
1099
+ "links": [
1100
+ 641,
1101
+ 724
1102
+ ]
1103
+ },
1104
+ {
1105
+ "name": "BOOL",
1106
+ "type": "BOOLEAN",
1107
+ "links": null
1108
+ }
1109
+ ],
1110
+ "title": "16 F/S",
1111
+ "properties": {
1112
+ "cnr_id": "comfy-core",
1113
+ "ver": "0.27.0",
1114
+ "Node name for S&R": "ComfyMathExpression"
1115
+ },
1116
+ "widgets_values": [
1117
+ "a * 16 + 1"
1118
+ ]
1119
+ },
1120
+ {
1121
+ "id": 57,
1122
+ "type": "AudioEncoderLoader",
1123
+ "pos": [
1124
+ -1030,
1125
+ 1290
1126
+ ],
1127
+ "size": [
1128
+ 290,
1129
+ 60
1130
+ ],
1131
+ "flags": {},
1132
+ "order": 6,
1133
+ "mode": 0,
1134
+ "inputs": [],
1135
+ "outputs": [
1136
+ {
1137
+ "name": "AUDIO_ENCODER",
1138
+ "type": "AUDIO_ENCODER",
1139
+ "links": [
1140
+ 733
1141
+ ]
1142
+ }
1143
+ ],
1144
+ "properties": {
1145
+ "cnr_id": "comfy-core",
1146
+ "ver": "0.3.54",
1147
+ "Node name for S&R": "AudioEncoderLoader",
1148
+ "models": [
1149
+ {
1150
+ "name": "wav2vec2_large_english_fp16.safetensors",
1151
+ "url": "https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/audio_encoders/wav2vec2_large_english_fp16.safetensors",
1152
+ "directory": "audio_encoders"
1153
+ }
1154
+ ]
1155
+ },
1156
+ "widgets_values": [
1157
+ "wav2vec2_large_english_fp16.safetensors"
1158
+ ]
1159
+ },
1160
+ {
1161
+ "id": 266,
1162
+ "type": "AudioEncoderEncode",
1163
+ "pos": [
1164
+ -220,
1165
+ 1610
1166
+ ],
1167
+ "size": [
1168
+ 290,
1169
+ 50
1170
+ ],
1171
+ "flags": {},
1172
+ "order": 33,
1173
+ "mode": 0,
1174
+ "inputs": [
1175
+ {
1176
+ "name": "audio_encoder",
1177
+ "type": "AUDIO_ENCODER",
1178
+ "link": 735
1179
+ },
1180
+ {
1181
+ "name": "audio",
1182
+ "type": "AUDIO",
1183
+ "link": 738
1184
+ }
1185
+ ],
1186
+ "outputs": [
1187
+ {
1188
+ "name": "AUDIO_ENCODER_OUTPUT",
1189
+ "type": "AUDIO_ENCODER_OUTPUT",
1190
+ "links": [
1191
+ 730
1192
+ ]
1193
+ }
1194
+ ],
1195
+ "properties": {
1196
+ "cnr_id": "comfy-core",
1197
+ "ver": "0.27.0",
1198
+ "Node name for S&R": "AudioEncoderEncode"
1199
+ },
1200
+ "widgets_values": []
1201
+ },
1202
+ {
1203
+ "id": 275,
1204
+ "type": "AudioEncoderLoader",
1205
+ "pos": [
1206
+ 120,
1207
+ 1310
1208
+ ],
1209
+ "size": [
1210
+ 290,
1211
+ 60
1212
+ ],
1213
+ "flags": {},
1214
+ "order": 7,
1215
+ "mode": 2,
1216
+ "inputs": [],
1217
+ "outputs": [
1218
+ {
1219
+ "name": "AUDIO_ENCODER",
1220
+ "type": "AUDIO_ENCODER",
1221
+ "links": [
1222
+ 746
1223
+ ]
1224
+ }
1225
+ ],
1226
+ "properties": {
1227
+ "cnr_id": "comfy-core",
1228
+ "ver": "0.3.54",
1229
+ "Node name for S&R": "AudioEncoderLoader",
1230
+ "models": [
1231
+ {
1232
+ "name": "wav2vec2_large_english_fp16.safetensors",
1233
+ "url": "https://huggingface.co/Comfy-Org/Wan_2.2_ComfyUI_Repackaged/resolve/main/split_files/audio_encoders/wav2vec2_large_english_fp16.safetensors",
1234
+ "directory": "audio_encoders"
1235
+ }
1236
+ ]
1237
+ },
1238
+ "widgets_values": [
1239
+ "wav2vec2_large_english_fp16.safetensors"
1240
+ ]
1241
+ },
1242
+ {
1243
+ "id": 273,
1244
+ "type": "LoadAudio",
1245
+ "pos": [
1246
+ 120,
1247
+ 1420
1248
+ ],
1249
+ "size": [
1250
+ 290,
1251
+ 140
1252
+ ],
1253
+ "flags": {},
1254
+ "order": 8,
1255
+ "mode": 2,
1256
+ "inputs": [],
1257
+ "outputs": [
1258
+ {
1259
+ "name": "AUDIO",
1260
+ "type": "AUDIO",
1261
+ "links": [
1262
+ 747
1263
+ ]
1264
+ }
1265
+ ],
1266
+ "title": "audio_1",
1267
+ "properties": {
1268
+ "cnr_id": "comfy-core",
1269
+ "ver": "0.3.54",
1270
+ "Node name for S&R": "LoadAudio"
1271
+ },
1272
+ "widgets_values": [
1273
+ "anthony.wav",
1274
+ null,
1275
+ null
1276
+ ]
1277
+ },
1278
+ {
1279
+ "id": 274,
1280
+ "type": "TrimAudioDuration",
1281
+ "pos": [
1282
+ 440,
1283
+ 1420
1284
+ ],
1285
+ "size": [
1286
+ 270,
1287
+ 90
1288
+ ],
1289
+ "flags": {},
1290
+ "order": 20,
1291
+ "mode": 2,
1292
+ "inputs": [
1293
+ {
1294
+ "name": "audio",
1295
+ "type": "AUDIO",
1296
+ "link": 747
1297
+ }
1298
+ ],
1299
+ "outputs": [
1300
+ {
1301
+ "name": "AUDIO",
1302
+ "type": "AUDIO",
1303
+ "links": [
1304
+ 751
1305
+ ]
1306
+ }
1307
+ ],
1308
+ "properties": {
1309
+ "cnr_id": "comfy-core",
1310
+ "ver": "0.27.0",
1311
+ "Node name for S&R": "TrimAudioDuration"
1312
+ },
1313
+ "widgets_values": [
1314
+ 0,
1315
+ 9
1316
+ ]
1317
+ },
1318
+ {
1319
+ "id": 276,
1320
+ "type": "AudioEncoderEncode",
1321
+ "pos": [
1322
+ 740,
1323
+ 1300
1324
+ ],
1325
+ "size": [
1326
+ 290,
1327
+ 50
1328
+ ],
1329
+ "flags": {},
1330
+ "order": 31,
1331
+ "mode": 2,
1332
+ "inputs": [
1333
+ {
1334
+ "name": "audio_encoder",
1335
+ "type": "AUDIO_ENCODER",
1336
+ "link": 746
1337
+ },
1338
+ {
1339
+ "name": "audio",
1340
+ "type": "AUDIO",
1341
+ "link": 752
1342
+ }
1343
+ ],
1344
+ "outputs": [
1345
+ {
1346
+ "name": "AUDIO_ENCODER_OUTPUT",
1347
+ "type": "AUDIO_ENCODER_OUTPUT",
1348
+ "links": []
1349
+ }
1350
+ ],
1351
+ "properties": {
1352
+ "cnr_id": "comfy-core",
1353
+ "ver": "0.27.0",
1354
+ "Node name for S&R": "AudioEncoderEncode"
1355
+ },
1356
+ "widgets_values": []
1357
+ },
1358
+ {
1359
+ "id": 279,
1360
+ "type": "Reroute",
1361
+ "pos": [
1362
+ 1920,
1363
+ 250
1364
+ ],
1365
+ "size": [
1366
+ 75,
1367
+ 26
1368
+ ],
1369
+ "flags": {},
1370
+ "order": 39,
1371
+ "mode": 0,
1372
+ "inputs": [
1373
+ {
1374
+ "name": "",
1375
+ "type": "*",
1376
+ "link": 757
1377
+ }
1378
+ ],
1379
+ "outputs": [
1380
+ {
1381
+ "name": "",
1382
+ "type": "AUDIO",
1383
+ "links": [
1384
+ 755
1385
+ ]
1386
+ }
1387
+ ],
1388
+ "properties": {
1389
+ "showOutputText": false,
1390
+ "horizontal": false
1391
+ },
1392
+ "color": "#323",
1393
+ "bgcolor": "#535"
1394
+ },
1395
+ {
1396
+ "id": 277,
1397
+ "type": "Reroute",
1398
+ "pos": [
1399
+ 770,
1400
+ 1430
1401
+ ],
1402
+ "size": [
1403
+ 75,
1404
+ 26
1405
+ ],
1406
+ "flags": {},
1407
+ "order": 27,
1408
+ "mode": 2,
1409
+ "inputs": [
1410
+ {
1411
+ "name": "",
1412
+ "type": "*",
1413
+ "link": 751
1414
+ }
1415
+ ],
1416
+ "outputs": [
1417
+ {
1418
+ "name": "",
1419
+ "type": "AUDIO",
1420
+ "links": [
1421
+ 752
1422
+ ]
1423
+ }
1424
+ ],
1425
+ "properties": {
1426
+ "showOutputText": false,
1427
+ "horizontal": false
1428
+ },
1429
+ "color": "#323",
1430
+ "bgcolor": "#535"
1431
+ },
1432
+ {
1433
+ "id": 239,
1434
+ "type": "PrimitiveInt",
1435
+ "pos": [
1436
+ 710,
1437
+ 1080
1438
+ ],
1439
+ "size": [
1440
+ 350,
1441
+ 90
1442
+ ],
1443
+ "flags": {},
1444
+ "order": 9,
1445
+ "mode": 0,
1446
+ "inputs": [],
1447
+ "outputs": [
1448
+ {
1449
+ "name": "INT",
1450
+ "type": "INT",
1451
+ "links": [
1452
+ 627
1453
+ ]
1454
+ }
1455
+ ],
1456
+ "title": "Video Length in Seconds",
1457
+ "properties": {
1458
+ "cnr_id": "comfy-core",
1459
+ "ver": "0.27.0",
1460
+ "Node name for S&R": "PrimitiveInt"
1461
+ },
1462
+ "widgets_values": [
1463
+ 9,
1464
+ "fixed"
1465
+ ]
1466
+ },
1467
+ {
1468
+ "id": 269,
1469
+ "type": "BerniniS2VConditioningV2",
1470
+ "pos": [
1471
+ 710,
1472
+ 600
1473
+ ],
1474
+ "size": [
1475
+ 360,
1476
+ 450
1477
+ ],
1478
+ "flags": {},
1479
+ "order": 35,
1480
+ "mode": 0,
1481
+ "inputs": [
1482
+ {
1483
+ "name": "positive",
1484
+ "type": "CONDITIONING",
1485
+ "link": 725
1486
+ },
1487
+ {
1488
+ "name": "negative",
1489
+ "type": "CONDITIONING",
1490
+ "link": 726
1491
+ },
1492
+ {
1493
+ "name": "vae",
1494
+ "type": "VAE",
1495
+ "link": 727
1496
+ },
1497
+ {
1498
+ "name": "audio_1",
1499
+ "type": "AUDIO_ENCODER_OUTPUT",
1500
+ "link": 758
1501
+ },
1502
+ {
1503
+ "name": "mask_1",
1504
+ "type": "MASK",
1505
+ "link": 729
1506
+ },
1507
+ {
1508
+ "name": "audio_2",
1509
+ "shape": 7,
1510
+ "type": "AUDIO_ENCODER_OUTPUT",
1511
+ "link": 730
1512
+ },
1513
+ {
1514
+ "name": "mask_2",
1515
+ "shape": 7,
1516
+ "type": "MASK",
1517
+ "link": 732
1518
+ },
1519
+ {
1520
+ "name": "source_video",
1521
+ "shape": 7,
1522
+ "type": "IMAGE",
1523
+ "link": null
1524
+ },
1525
+ {
1526
+ "name": "reference_video",
1527
+ "shape": 7,
1528
+ "type": "IMAGE",
1529
+ "link": null
1530
+ },
1531
+ {
1532
+ "label": "reference_image_0",
1533
+ "name": "reference_images.reference_image_0",
1534
+ "shape": 7,
1535
+ "type": "IMAGE",
1536
+ "link": 731
1537
+ },
1538
+ {
1539
+ "label": "reference_image_1",
1540
+ "name": "reference_images.reference_image_1",
1541
+ "shape": 7,
1542
+ "type": "IMAGE",
1543
+ "link": null
1544
+ },
1545
+ {
1546
+ "name": "length",
1547
+ "type": "INT",
1548
+ "widget": {
1549
+ "name": "length"
1550
+ },
1551
+ "link": 724
1552
+ }
1553
+ ],
1554
+ "outputs": [
1555
+ {
1556
+ "name": "positive",
1557
+ "type": "CONDITIONING",
1558
+ "links": [
1559
+ 740,
1560
+ 741
1561
+ ]
1562
+ },
1563
+ {
1564
+ "name": "negative",
1565
+ "type": "CONDITIONING",
1566
+ "links": [
1567
+ 742,
1568
+ 743
1569
+ ]
1570
+ },
1571
+ {
1572
+ "name": "latent",
1573
+ "type": "LATENT",
1574
+ "links": [
1575
+ 744
1576
+ ]
1577
+ }
1578
+ ],
1579
+ "properties": {
1580
+ "Node name for S&R": "BerniniS2VConditioningV2"
1581
+ },
1582
+ "widgets_values": [
1583
+ 720,
1584
+ 480,
1585
+ 81,
1586
+ 1,
1587
+ -1,
1588
+ 0,
1589
+ 1,
1590
+ 848
1591
+ ]
1592
+ },
1593
+ {
1594
+ "id": 240,
1595
+ "type": "VHS_VideoCombine",
1596
+ "pos": [
1597
+ 2040,
1598
+ 230
1599
+ ],
1600
+ "size": [
1601
+ 470,
1602
+ 648
1603
+ ],
1604
+ "flags": {},
1605
+ "order": 43,
1606
+ "mode": 0,
1607
+ "inputs": [
1608
+ {
1609
+ "name": "images",
1610
+ "type": "IMAGE",
1611
+ "link": 629
1612
+ },
1613
+ {
1614
+ "name": "audio",
1615
+ "shape": 7,
1616
+ "type": "AUDIO",
1617
+ "link": 755
1618
+ },
1619
+ {
1620
+ "name": "meta_batch",
1621
+ "shape": 7,
1622
+ "type": "VHS_BatchManager",
1623
+ "link": null
1624
+ },
1625
+ {
1626
+ "name": "vae",
1627
+ "shape": 7,
1628
+ "type": "VAE",
1629
+ "link": null
1630
+ }
1631
+ ],
1632
+ "outputs": [
1633
+ {
1634
+ "name": "Filenames",
1635
+ "type": "VHS_FILENAMES",
1636
+ "links": null
1637
+ }
1638
+ ],
1639
+ "properties": {
1640
+ "cnr_id": "comfyui-videohelpersuite",
1641
+ "ver": "1.7.9",
1642
+ "Node name for S&R": "VHS_VideoCombine"
1643
+ },
1644
+ "widgets_values": {
1645
+ "frame_rate": 24,
1646
+ "loop_count": 0,
1647
+ "filename_prefix": "video/ComfyUI",
1648
+ "format": "video/h264-mp4",
1649
+ "pix_fmt": "yuv420p",
1650
+ "crf": 19,
1651
+ "save_metadata": true,
1652
+ "trim_to_audio": true,
1653
+ "pingpong": false,
1654
+ "save_output": true,
1655
+ "videopreview": {
1656
+ "hidden": false,
1657
+ "paused": false,
1658
+ "params": {
1659
+ "filename": "ComfyUI_00207-audio.mp4",
1660
+ "subfolder": "video",
1661
+ "type": "output",
1662
+ "format": "video/h264-mp4",
1663
+ "frame_rate": 24,
1664
+ "workflow": "ComfyUI_00207.png",
1665
+ "fullpath": "/media/dgt/N33B/AI/ComfyUI/output/video/ComfyUI_00207-audio.mp4"
1666
+ }
1667
+ }
1668
+ }
1669
+ },
1670
+ {
1671
+ "id": 268,
1672
+ "type": "LoadImage",
1673
+ "pos": [
1674
+ 320,
1675
+ 840
1676
+ ],
1677
+ "size": [
1678
+ 290,
1679
+ 320
1680
+ ],
1681
+ "flags": {},
1682
+ "order": 10,
1683
+ "mode": 0,
1684
+ "inputs": [],
1685
+ "outputs": [
1686
+ {
1687
+ "name": "IMAGE",
1688
+ "type": "IMAGE",
1689
+ "links": []
1690
+ },
1691
+ {
1692
+ "name": "MASK",
1693
+ "type": "MASK",
1694
+ "links": [
1695
+ 732
1696
+ ]
1697
+ }
1698
+ ],
1699
+ "title": "Mask_2",
1700
+ "properties": {
1701
+ "cnr_id": "comfy-core",
1702
+ "ver": "0.27.0",
1703
+ "Node name for S&R": "LoadImage",
1704
+ "image": "clipspace/clipspace-painted-masked-1783665412223.png [input]"
1705
+ },
1706
+ "widgets_values": [
1707
+ "clipspace/clipspace-painted-masked-1783665412223.png [input]",
1708
+ "image"
1709
+ ]
1710
+ },
1711
+ {
1712
+ "id": 278,
1713
+ "type": "Reroute",
1714
+ "pos": [
1715
+ 0,
1716
+ 1220
1717
+ ],
1718
+ "size": [
1719
+ 75,
1720
+ 26
1721
+ ],
1722
+ "flags": {},
1723
+ "order": 37,
1724
+ "mode": 0,
1725
+ "inputs": [
1726
+ {
1727
+ "name": "",
1728
+ "type": "*",
1729
+ "link": 754
1730
+ }
1731
+ ],
1732
+ "outputs": [
1733
+ {
1734
+ "name": "",
1735
+ "type": "AUDIO",
1736
+ "links": [
1737
+ 757
1738
+ ]
1739
+ }
1740
+ ],
1741
+ "properties": {
1742
+ "showOutputText": false,
1743
+ "horizontal": false
1744
+ },
1745
+ "color": "#323",
1746
+ "bgcolor": "#535"
1747
+ },
1748
+ {
1749
+ "id": 58,
1750
+ "type": "LoadAudio",
1751
+ "pos": [
1752
+ -1030,
1753
+ 1410
1754
+ ],
1755
+ "size": [
1756
+ 290,
1757
+ 140
1758
+ ],
1759
+ "flags": {},
1760
+ "order": 11,
1761
+ "mode": 0,
1762
+ "inputs": [],
1763
+ "outputs": [
1764
+ {
1765
+ "name": "AUDIO",
1766
+ "type": "AUDIO",
1767
+ "links": [
1768
+ 671
1769
+ ]
1770
+ }
1771
+ ],
1772
+ "title": "audio_1",
1773
+ "properties": {
1774
+ "cnr_id": "comfy-core",
1775
+ "ver": "0.3.54",
1776
+ "Node name for S&R": "LoadAudio"
1777
+ },
1778
+ "widgets_values": [
1779
+ "I am The Dude Playing The Dude.wav",
1780
+ null,
1781
+ null
1782
+ ]
1783
+ },
1784
+ {
1785
+ "id": 253,
1786
+ "type": "TrimAudioDuration",
1787
+ "pos": [
1788
+ -660,
1789
+ 1610
1790
+ ],
1791
+ "size": [
1792
+ 270,
1793
+ 90
1794
+ ],
1795
+ "flags": {},
1796
+ "order": 23,
1797
+ "mode": 0,
1798
+ "inputs": [
1799
+ {
1800
+ "name": "audio",
1801
+ "type": "AUDIO",
1802
+ "link": 667
1803
+ }
1804
+ ],
1805
+ "outputs": [
1806
+ {
1807
+ "name": "AUDIO",
1808
+ "type": "AUDIO",
1809
+ "links": [
1810
+ 737
1811
+ ]
1812
+ }
1813
+ ],
1814
+ "properties": {
1815
+ "cnr_id": "comfy-core",
1816
+ "ver": "0.27.0",
1817
+ "Node name for S&R": "TrimAudioDuration"
1818
+ },
1819
+ "widgets_values": [
1820
+ 0,
1821
+ 3
1822
+ ]
1823
+ },
1824
+ {
1825
+ "id": 223,
1826
+ "type": "TrimAudioDuration",
1827
+ "pos": [
1828
+ -660,
1829
+ 1410
1830
+ ],
1831
+ "size": [
1832
+ 270,
1833
+ 90
1834
+ ],
1835
+ "flags": {},
1836
+ "order": 22,
1837
+ "mode": 0,
1838
+ "inputs": [
1839
+ {
1840
+ "name": "audio",
1841
+ "type": "AUDIO",
1842
+ "link": 671
1843
+ }
1844
+ ],
1845
+ "outputs": [
1846
+ {
1847
+ "name": "AUDIO",
1848
+ "type": "AUDIO",
1849
+ "links": [
1850
+ 736
1851
+ ]
1852
+ }
1853
+ ],
1854
+ "properties": {
1855
+ "cnr_id": "comfy-core",
1856
+ "ver": "0.27.0",
1857
+ "Node name for S&R": "TrimAudioDuration"
1858
+ },
1859
+ "widgets_values": [
1860
+ 0,
1861
+ 6
1862
+ ]
1863
+ },
1864
+ {
1865
+ "id": 247,
1866
+ "type": "LoadAudio",
1867
+ "pos": [
1868
+ -1030,
1869
+ 1610
1870
+ ],
1871
+ "size": [
1872
+ 290,
1873
+ 140
1874
+ ],
1875
+ "flags": {},
1876
+ "order": 12,
1877
+ "mode": 0,
1878
+ "inputs": [],
1879
+ "outputs": [
1880
+ {
1881
+ "name": "AUDIO",
1882
+ "type": "AUDIO",
1883
+ "links": [
1884
+ 667
1885
+ ]
1886
+ }
1887
+ ],
1888
+ "title": "audio_2",
1889
+ "properties": {
1890
+ "cnr_id": "comfy-core",
1891
+ "ver": "0.27.0",
1892
+ "Node name for S&R": "LoadAudio"
1893
+ },
1894
+ "widgets_values": [
1895
+ "Im the Dude.wav",
1896
+ null,
1897
+ null
1898
+ ]
1899
+ },
1900
+ {
1901
+ "id": 252,
1902
+ "type": "PrimitiveStringMultiline",
1903
+ "pos": [
1904
+ -10,
1905
+ 170
1906
+ ],
1907
+ "size": [
1908
+ 710,
1909
+ 290
1910
+ ],
1911
+ "flags": {},
1912
+ "order": 13,
1913
+ "mode": 0,
1914
+ "inputs": [],
1915
+ "outputs": [
1916
+ {
1917
+ "name": "STRING",
1918
+ "type": "STRING",
1919
+ "links": [
1920
+ 664
1921
+ ]
1922
+ }
1923
+ ],
1924
+ "properties": {
1925
+ "cnr_id": "comfy-core",
1926
+ "ver": "0.27.0",
1927
+ "Node name for S&R": "PrimitiveStringMultiline"
1928
+ },
1929
+ "widgets_values": [
1930
+ "two men in image0 sitting in a restaurant."
1931
+ ],
1932
+ "color": "#232",
1933
+ "bgcolor": "#353"
1934
+ },
1935
+ {
1936
+ "id": 56,
1937
+ "type": "AudioEncoderEncode",
1938
+ "pos": [
1939
+ -210,
1940
+ 1410
1941
+ ],
1942
+ "size": [
1943
+ 290,
1944
+ 50
1945
+ ],
1946
+ "flags": {},
1947
+ "order": 32,
1948
+ "mode": 0,
1949
+ "inputs": [
1950
+ {
1951
+ "name": "audio_encoder",
1952
+ "type": "AUDIO_ENCODER",
1953
+ "link": 734
1954
+ },
1955
+ {
1956
+ "name": "audio",
1957
+ "type": "AUDIO",
1958
+ "link": 622
1959
+ }
1960
+ ],
1961
+ "outputs": [
1962
+ {
1963
+ "name": "AUDIO_ENCODER_OUTPUT",
1964
+ "type": "AUDIO_ENCODER_OUTPUT",
1965
+ "links": [
1966
+ 758
1967
+ ]
1968
+ }
1969
+ ],
1970
+ "properties": {
1971
+ "cnr_id": "comfy-core",
1972
+ "ver": "0.3.54",
1973
+ "Node name for S&R": "AudioEncoderEncode"
1974
+ },
1975
+ "widgets_values": []
1976
+ },
1977
+ {
1978
+ "id": 272,
1979
+ "type": "PreviewAudio",
1980
+ "pos": [
1981
+ -180,
1982
+ 1720
1983
+ ],
1984
+ "size": [
1985
+ 270,
1986
+ 90
1987
+ ],
1988
+ "flags": {},
1989
+ "order": 36,
1990
+ "mode": 0,
1991
+ "inputs": [
1992
+ {
1993
+ "name": "audio",
1994
+ "type": "AUDIO",
1995
+ "link": 745
1996
+ }
1997
+ ],
1998
+ "outputs": [
1999
+ {
2000
+ "name": "audio",
2001
+ "type": "AUDIO",
2002
+ "links": null
2003
+ }
2004
+ ],
2005
+ "properties": {
2006
+ "cnr_id": "comfy-core",
2007
+ "ver": "0.27.0",
2008
+ "Node name for S&R": "PreviewAudio"
2009
+ },
2010
+ "widgets_values": []
2011
+ }
2012
+ ],
2013
+ "links": [
2014
+ [
2015
+ 74,
2016
+ 38,
2017
+ 0,
2018
+ 6,
2019
+ 0,
2020
+ "CLIP"
2021
+ ],
2022
+ [
2023
+ 75,
2024
+ 38,
2025
+ 0,
2026
+ 7,
2027
+ 0,
2028
+ "CLIP"
2029
+ ],
2030
+ [
2031
+ 285,
2032
+ 39,
2033
+ 0,
2034
+ 80,
2035
+ 1,
2036
+ "VAE"
2037
+ ],
2038
+ [
2039
+ 363,
2040
+ 107,
2041
+ 0,
2042
+ 54,
2043
+ 0,
2044
+ "MODEL"
2045
+ ],
2046
+ [
2047
+ 522,
2048
+ 54,
2049
+ 0,
2050
+ 205,
2051
+ 0,
2052
+ "MODEL"
2053
+ ],
2054
+ [
2055
+ 526,
2056
+ 205,
2057
+ 0,
2058
+ 207,
2059
+ 3,
2060
+ "LATENT"
2061
+ ],
2062
+ [
2063
+ 533,
2064
+ 207,
2065
+ 0,
2066
+ 80,
2067
+ 0,
2068
+ "LATENT"
2069
+ ],
2070
+ [
2071
+ 580,
2072
+ 80,
2073
+ 0,
2074
+ 201,
2075
+ 0,
2076
+ "IMAGE"
2077
+ ],
2078
+ [
2079
+ 581,
2080
+ 210,
2081
+ 0,
2082
+ 214,
2083
+ 0,
2084
+ "MODEL"
2085
+ ],
2086
+ [
2087
+ 582,
2088
+ 214,
2089
+ 0,
2090
+ 207,
2091
+ 0,
2092
+ "MODEL"
2093
+ ],
2094
+ [
2095
+ 585,
2096
+ 37,
2097
+ 0,
2098
+ 107,
2099
+ 0,
2100
+ "MODEL"
2101
+ ],
2102
+ [
2103
+ 586,
2104
+ 204,
2105
+ 0,
2106
+ 210,
2107
+ 0,
2108
+ "MODEL"
2109
+ ],
2110
+ [
2111
+ 622,
2112
+ 235,
2113
+ 0,
2114
+ 56,
2115
+ 1,
2116
+ "AUDIO"
2117
+ ],
2118
+ [
2119
+ 627,
2120
+ 239,
2121
+ 0,
2122
+ 238,
2123
+ 0,
2124
+ "INT"
2125
+ ],
2126
+ [
2127
+ 629,
2128
+ 201,
2129
+ 0,
2130
+ 240,
2131
+ 0,
2132
+ "IMAGE"
2133
+ ],
2134
+ [
2135
+ 661,
2136
+ 251,
2137
+ 1,
2138
+ 249,
2139
+ 0,
2140
+ "INT"
2141
+ ],
2142
+ [
2143
+ 662,
2144
+ 249,
2145
+ 0,
2146
+ 250,
2147
+ 0,
2148
+ "STRING"
2149
+ ],
2150
+ [
2151
+ 663,
2152
+ 250,
2153
+ 0,
2154
+ 6,
2155
+ 1,
2156
+ "STRING"
2157
+ ],
2158
+ [
2159
+ 664,
2160
+ 252,
2161
+ 0,
2162
+ 250,
2163
+ 1,
2164
+ "STRING"
2165
+ ],
2166
+ [
2167
+ 667,
2168
+ 247,
2169
+ 0,
2170
+ 253,
2171
+ 0,
2172
+ "AUDIO"
2173
+ ],
2174
+ [
2175
+ 671,
2176
+ 58,
2177
+ 0,
2178
+ 223,
2179
+ 0,
2180
+ "AUDIO"
2181
+ ],
2182
+ [
2183
+ 702,
2184
+ 241,
2185
+ 1,
2186
+ 265,
2187
+ 0,
2188
+ "MASK"
2189
+ ],
2190
+ [
2191
+ 717,
2192
+ 235,
2193
+ 0,
2194
+ 267,
2195
+ 0,
2196
+ "AUDIO"
2197
+ ],
2198
+ [
2199
+ 724,
2200
+ 238,
2201
+ 1,
2202
+ 269,
2203
+ 11,
2204
+ "INT"
2205
+ ],
2206
+ [
2207
+ 725,
2208
+ 6,
2209
+ 0,
2210
+ 269,
2211
+ 0,
2212
+ "CONDITIONING"
2213
+ ],
2214
+ [
2215
+ 726,
2216
+ 7,
2217
+ 0,
2218
+ 269,
2219
+ 1,
2220
+ "CONDITIONING"
2221
+ ],
2222
+ [
2223
+ 727,
2224
+ 39,
2225
+ 0,
2226
+ 269,
2227
+ 2,
2228
+ "VAE"
2229
+ ],
2230
+ [
2231
+ 729,
2232
+ 241,
2233
+ 1,
2234
+ 269,
2235
+ 4,
2236
+ "MASK"
2237
+ ],
2238
+ [
2239
+ 730,
2240
+ 266,
2241
+ 0,
2242
+ 269,
2243
+ 5,
2244
+ "AUDIO_ENCODER_OUTPUT"
2245
+ ],
2246
+ [
2247
+ 731,
2248
+ 241,
2249
+ 0,
2250
+ 269,
2251
+ 9,
2252
+ "IMAGE"
2253
+ ],
2254
+ [
2255
+ 732,
2256
+ 268,
2257
+ 1,
2258
+ 269,
2259
+ 6,
2260
+ "MASK"
2261
+ ],
2262
+ [
2263
+ 733,
2264
+ 57,
2265
+ 0,
2266
+ 270,
2267
+ 0,
2268
+ "AUDIO_ENCODER"
2269
+ ],
2270
+ [
2271
+ 734,
2272
+ 270,
2273
+ 0,
2274
+ 56,
2275
+ 0,
2276
+ "AUDIO_ENCODER"
2277
+ ],
2278
+ [
2279
+ 735,
2280
+ 270,
2281
+ 0,
2282
+ 266,
2283
+ 0,
2284
+ "AUDIO_ENCODER"
2285
+ ],
2286
+ [
2287
+ 736,
2288
+ 223,
2289
+ 0,
2290
+ 235,
2291
+ 0,
2292
+ "AUDIO"
2293
+ ],
2294
+ [
2295
+ 737,
2296
+ 253,
2297
+ 0,
2298
+ 271,
2299
+ 0,
2300
+ "AUDIO"
2301
+ ],
2302
+ [
2303
+ 738,
2304
+ 271,
2305
+ 0,
2306
+ 266,
2307
+ 1,
2308
+ "AUDIO"
2309
+ ],
2310
+ [
2311
+ 739,
2312
+ 271,
2313
+ 0,
2314
+ 267,
2315
+ 1,
2316
+ "AUDIO"
2317
+ ],
2318
+ [
2319
+ 740,
2320
+ 269,
2321
+ 0,
2322
+ 205,
2323
+ 1,
2324
+ "CONDITIONING"
2325
+ ],
2326
+ [
2327
+ 741,
2328
+ 269,
2329
+ 0,
2330
+ 207,
2331
+ 1,
2332
+ "CONDITIONING"
2333
+ ],
2334
+ [
2335
+ 742,
2336
+ 269,
2337
+ 1,
2338
+ 205,
2339
+ 2,
2340
+ "CONDITIONING"
2341
+ ],
2342
+ [
2343
+ 743,
2344
+ 269,
2345
+ 1,
2346
+ 207,
2347
+ 2,
2348
+ "CONDITIONING"
2349
+ ],
2350
+ [
2351
+ 744,
2352
+ 269,
2353
+ 2,
2354
+ 205,
2355
+ 3,
2356
+ "LATENT"
2357
+ ],
2358
+ [
2359
+ 745,
2360
+ 267,
2361
+ 0,
2362
+ 272,
2363
+ 0,
2364
+ "AUDIO"
2365
+ ],
2366
+ [
2367
+ 746,
2368
+ 275,
2369
+ 0,
2370
+ 276,
2371
+ 0,
2372
+ "AUDIO_ENCODER"
2373
+ ],
2374
+ [
2375
+ 747,
2376
+ 273,
2377
+ 0,
2378
+ 274,
2379
+ 0,
2380
+ "AUDIO"
2381
+ ],
2382
+ [
2383
+ 751,
2384
+ 274,
2385
+ 0,
2386
+ 277,
2387
+ 0,
2388
+ "AUDIO"
2389
+ ],
2390
+ [
2391
+ 752,
2392
+ 277,
2393
+ 0,
2394
+ 276,
2395
+ 1,
2396
+ "AUDIO"
2397
+ ],
2398
+ [
2399
+ 754,
2400
+ 267,
2401
+ 0,
2402
+ 278,
2403
+ 0,
2404
+ "AUDIO"
2405
+ ],
2406
+ [
2407
+ 755,
2408
+ 279,
2409
+ 0,
2410
+ 240,
2411
+ 1,
2412
+ "AUDIO"
2413
+ ],
2414
+ [
2415
+ 757,
2416
+ 278,
2417
+ 0,
2418
+ 279,
2419
+ 0,
2420
+ "AUDIO"
2421
+ ],
2422
+ [
2423
+ 758,
2424
+ 56,
2425
+ 0,
2426
+ 269,
2427
+ 3,
2428
+ "AUDIO_ENCODER_OUTPUT"
2429
+ ]
2430
+ ],
2431
+ "groups": [
2432
+ {
2433
+ "id": 2,
2434
+ "title": "-",
2435
+ "bounding": [
2436
+ -110,
2437
+ 100,
2438
+ 2650,
2439
+ 1090
2440
+ ],
2441
+ "color": "#3f789e",
2442
+ "flags": {}
2443
+ },
2444
+ {
2445
+ "id": 3,
2446
+ "title": "-",
2447
+ "bounding": [
2448
+ -540,
2449
+ 100,
2450
+ 430,
2451
+ 490
2452
+ ],
2453
+ "color": "#3f789e",
2454
+ "flags": {}
2455
+ },
2456
+ {
2457
+ "id": 10,
2458
+ "title": "Tasks",
2459
+ "bounding": [
2460
+ -1030,
2461
+ 680,
2462
+ 920,
2463
+ 510
2464
+ ],
2465
+ "color": "#3f789e",
2466
+ "flags": {}
2467
+ },
2468
+ {
2469
+ "id": 11,
2470
+ "title": "2 Speakers",
2471
+ "bounding": [
2472
+ -1040,
2473
+ 1220,
2474
+ 1130,
2475
+ 540
2476
+ ],
2477
+ "color": "#3f789e",
2478
+ "flags": {}
2479
+ },
2480
+ {
2481
+ "id": 12,
2482
+ "title": "1 Speaker",
2483
+ "bounding": [
2484
+ 90,
2485
+ 1220,
2486
+ 1170,
2487
+ 530
2488
+ ],
2489
+ "color": "#3f789e",
2490
+ "flags": {}
2491
+ }
2492
+ ],
2493
+ "definitions": {
2494
+ "subgraphs": [
2495
+ {
2496
+ "id": "6da2792c-a953-4960-a53a-ff1b5df6fe09",
2497
+ "version": 1,
2498
+ "state": {
2499
+ "lastGroupId": 12,
2500
+ "lastNodeId": 279,
2501
+ "lastLinkId": 758,
2502
+ "lastRerouteId": 0
2503
+ },
2504
+ "revision": 0,
2505
+ "config": {},
2506
+ "name": "Select Per-Line Text by Index",
2507
+ "description": "Selects one line from multiline text by zero-based index for batch or list-driven prompt workflows.",
2508
+ "inputNode": {
2509
+ "id": -10,
2510
+ "bounding": [
2511
+ -990,
2512
+ 8595,
2513
+ 128,
2514
+ 88
2515
+ ]
2516
+ },
2517
+ "outputNode": {
2518
+ "id": -20,
2519
+ "bounding": [
2520
+ 710,
2521
+ 8585,
2522
+ 128,
2523
+ 68
2524
+ ]
2525
+ },
2526
+ "inputs": [
2527
+ {
2528
+ "id": "75417d82-a934-4ac9-b667-d8dcd5a3bfb3",
2529
+ "name": "text_per_line",
2530
+ "type": "STRING",
2531
+ "linkIds": [
2532
+ 13
2533
+ ],
2534
+ "localized_name": "text_per_line",
2535
+ "pos": [
2536
+ -886,
2537
+ 8619
2538
+ ]
2539
+ },
2540
+ {
2541
+ "id": "46e69a73-1804-4ca6-9175-31445bf0be96",
2542
+ "name": "index",
2543
+ "type": "INT",
2544
+ "linkIds": [
2545
+ 14
2546
+ ],
2547
+ "localized_name": "index",
2548
+ "pos": [
2549
+ -886,
2550
+ 8639
2551
+ ]
2552
+ }
2553
+ ],
2554
+ "outputs": [
2555
+ {
2556
+ "id": "e34e8ad1-84d2-4bd2-a460-eb7de6067c10",
2557
+ "name": "selected_line",
2558
+ "type": "STRING",
2559
+ "linkIds": [
2560
+ 10
2561
+ ],
2562
+ "localized_name": "selected_line",
2563
+ "pos": [
2564
+ 734,
2565
+ 8609
2566
+ ]
2567
+ }
2568
+ ],
2569
+ "widgets": [],
2570
+ "nodes": [
2571
+ {
2572
+ "id": 1,
2573
+ "type": "PreviewAny",
2574
+ "pos": [
2575
+ -500,
2576
+ 8400
2577
+ ],
2578
+ "size": [
2579
+ 230,
2580
+ 180
2581
+ ],
2582
+ "flags": {},
2583
+ "order": 0,
2584
+ "mode": 0,
2585
+ "inputs": [
2586
+ {
2587
+ "localized_name": "source",
2588
+ "name": "source",
2589
+ "type": "*",
2590
+ "link": 1
2591
+ }
2592
+ ],
2593
+ "outputs": [
2594
+ {
2595
+ "localized_name": "STRING",
2596
+ "name": "STRING",
2597
+ "type": "STRING",
2598
+ "links": [
2599
+ 6
2600
+ ]
2601
+ }
2602
+ ],
2603
+ "properties": {
2604
+ "cnr_id": "comfy-core",
2605
+ "ver": "0.19.0",
2606
+ "Node name for S&R": "PreviewAny",
2607
+ "ue_properties": {
2608
+ "widget_ue_connectable": {},
2609
+ "input_ue_unconnectable": {}
2610
+ }
2611
+ },
2612
+ "widgets_values": [
2613
+ null,
2614
+ null,
2615
+ null
2616
+ ]
2617
+ },
2618
+ {
2619
+ "id": 2,
2620
+ "type": "RegexExtract",
2621
+ "pos": [
2622
+ -240,
2623
+ 8740
2624
+ ],
2625
+ "size": [
2626
+ 470,
2627
+ 460
2628
+ ],
2629
+ "flags": {},
2630
+ "order": 1,
2631
+ "mode": 0,
2632
+ "showAdvanced": false,
2633
+ "inputs": [
2634
+ {
2635
+ "localized_name": "string",
2636
+ "name": "string",
2637
+ "type": "STRING",
2638
+ "widget": {
2639
+ "name": "string"
2640
+ },
2641
+ "link": 13
2642
+ },
2643
+ {
2644
+ "localized_name": "regex_pattern",
2645
+ "name": "regex_pattern",
2646
+ "type": "STRING",
2647
+ "widget": {
2648
+ "name": "regex_pattern"
2649
+ },
2650
+ "link": 9
2651
+ }
2652
+ ],
2653
+ "outputs": [
2654
+ {
2655
+ "localized_name": "STRING",
2656
+ "name": "STRING",
2657
+ "type": "STRING",
2658
+ "links": [
2659
+ 10
2660
+ ]
2661
+ }
2662
+ ],
2663
+ "properties": {
2664
+ "cnr_id": "comfy-core",
2665
+ "ver": "0.19.0",
2666
+ "Node name for S&R": "RegexExtract",
2667
+ "ue_properties": {
2668
+ "widget_ue_connectable": {},
2669
+ "input_ue_unconnectable": {}
2670
+ }
2671
+ },
2672
+ "widgets_values": [
2673
+ "You are a helpful assistant.\nYou are a helpful assistant specialized in text-to-image generation.\nYou are a helpful assistant specialized in text-to-video generation.\nYou are a helpful assistant specialized in image editing.\nYou are a helpful assistant specialized in subject-to-image generation.\nYou are a helpful assistant specialized in image-to-video generation.\nYou are a helpful assistant specialized in video editing.\nYou are a helpful assistant specialized in video editing on content propagation.\nYou are a helpful assistant specialized in video editing with reference.\nYou are a helpful assistant specialized in ads insertion.\nYou are a helpful assistant for editing. You may need to adjust the subject's action or position.\nYou are a helpful assistant for editing. You might need to adjust the video's style, lighting, colors, textures, and the subject's pose or action.",
2674
+ "",
2675
+ "First Group",
2676
+ false,
2677
+ false,
2678
+ false,
2679
+ 1
2680
+ ]
2681
+ },
2682
+ {
2683
+ "id": 248,
2684
+ "type": "PrimitiveInt",
2685
+ "pos": [
2686
+ -810,
2687
+ 8400
2688
+ ],
2689
+ "size": [
2690
+ 270,
2691
+ 110
2692
+ ],
2693
+ "flags": {},
2694
+ "order": 3,
2695
+ "mode": 0,
2696
+ "inputs": [
2697
+ {
2698
+ "localized_name": "value",
2699
+ "name": "value",
2700
+ "type": "INT",
2701
+ "widget": {
2702
+ "name": "value"
2703
+ },
2704
+ "link": 14
2705
+ }
2706
+ ],
2707
+ "outputs": [
2708
+ {
2709
+ "localized_name": "INT",
2710
+ "name": "INT",
2711
+ "type": "INT",
2712
+ "links": [
2713
+ 1
2714
+ ]
2715
+ }
2716
+ ],
2717
+ "title": "Int (line index)",
2718
+ "properties": {
2719
+ "cnr_id": "comfy-core",
2720
+ "ver": "0.19.0",
2721
+ "Node name for S&R": "Int (line index)",
2722
+ "ue_properties": {
2723
+ "widget_ue_connectable": {},
2724
+ "input_ue_unconnectable": {}
2725
+ }
2726
+ },
2727
+ "widgets_values": [
2728
+ 0,
2729
+ "fixed"
2730
+ ]
2731
+ },
2732
+ {
2733
+ "id": 8,
2734
+ "type": "StringReplace",
2735
+ "pos": [
2736
+ -240,
2737
+ 8400
2738
+ ],
2739
+ "size": [
2740
+ 400,
2741
+ 280
2742
+ ],
2743
+ "flags": {},
2744
+ "order": 2,
2745
+ "mode": 0,
2746
+ "inputs": [
2747
+ {
2748
+ "localized_name": "replace",
2749
+ "name": "replace",
2750
+ "type": "STRING",
2751
+ "widget": {
2752
+ "name": "replace"
2753
+ },
2754
+ "link": 6
2755
+ }
2756
+ ],
2757
+ "outputs": [
2758
+ {
2759
+ "localized_name": "STRING",
2760
+ "name": "STRING",
2761
+ "type": "STRING",
2762
+ "links": [
2763
+ 9
2764
+ ]
2765
+ }
2766
+ ],
2767
+ "properties": {
2768
+ "cnr_id": "comfy-core",
2769
+ "ver": "0.19.0",
2770
+ "Node name for S&R": "StringReplace",
2771
+ "ue_properties": {
2772
+ "widget_ue_connectable": {},
2773
+ "input_ue_unconnectable": {}
2774
+ }
2775
+ },
2776
+ "widgets_values": [
2777
+ "^(?:[^\\n]*\\n){index}([^\\n]*)(?:\\n|$)",
2778
+ "index",
2779
+ ""
2780
+ ]
2781
+ }
2782
+ ],
2783
+ "groups": [],
2784
+ "links": [
2785
+ {
2786
+ "id": 1,
2787
+ "origin_id": 248,
2788
+ "origin_slot": 0,
2789
+ "target_id": 1,
2790
+ "target_slot": 0,
2791
+ "type": "INT"
2792
+ },
2793
+ {
2794
+ "id": 9,
2795
+ "origin_id": 8,
2796
+ "origin_slot": 0,
2797
+ "target_id": 2,
2798
+ "target_slot": 1,
2799
+ "type": "STRING"
2800
+ },
2801
+ {
2802
+ "id": 6,
2803
+ "origin_id": 1,
2804
+ "origin_slot": 0,
2805
+ "target_id": 8,
2806
+ "target_slot": 0,
2807
+ "type": "STRING"
2808
+ },
2809
+ {
2810
+ "id": 10,
2811
+ "origin_id": 2,
2812
+ "origin_slot": 0,
2813
+ "target_id": -20,
2814
+ "target_slot": 0,
2815
+ "type": "STRING"
2816
+ },
2817
+ {
2818
+ "id": 13,
2819
+ "origin_id": -10,
2820
+ "origin_slot": 0,
2821
+ "target_id": 2,
2822
+ "target_slot": 0,
2823
+ "type": "STRING"
2824
+ },
2825
+ {
2826
+ "id": 14,
2827
+ "origin_id": -10,
2828
+ "origin_slot": 1,
2829
+ "target_id": 248,
2830
+ "target_slot": 0,
2831
+ "type": "INT"
2832
+ }
2833
+ ],
2834
+ "extra": {
2835
+ "ue_links": [],
2836
+ "links_added_by_ue": []
2837
+ }
2838
+ }
2839
+ ]
2840
+ },
2841
+ "config": {},
2842
+ "extra": {
2843
+ "ds": {
2844
+ "scale": 0.6618166065230274,
2845
+ "offset": [
2846
+ 567.2760482151011,
2847
+ 4.735915514621264
2848
+ ]
2849
+ },
2850
+ "frontendVersion": "1.45.20",
2851
+ "VHS_latentpreview": false,
2852
+ "VHS_latentpreviewrate": 0,
2853
+ "VHS_MetadataImage": true,
2854
+ "VHS_KeepIntermediate": true
2855
+ },
2856
+ "version": 0.4
2857
+ }
ComfyUI-WanBerniniS2V_v2/__init__.py ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ from .model_patch import apply_model_patches
2
+
3
+ apply_model_patches()
4
+
5
+ from .nodes import comfy_entrypoint
6
+
7
+ __all__ = ["comfy_entrypoint"]
ComfyUI-WanBerniniS2V_v2/audio_mask.py ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import torch
2
+
3
+ import comfy.utils
4
+
5
+ WAN_VAE_SCALE = 8
6
+ WAN_PATCH_SPATIAL = 2
7
+
8
+
9
+ def _padded_latent_dim(pixels: int) -> int:
10
+ latent = pixels // WAN_VAE_SCALE
11
+ return latent + (WAN_PATCH_SPATIAL - latent % WAN_PATCH_SPATIAL) % WAN_PATCH_SPATIAL
12
+
13
+
14
+ def token_grid_size(width: int, height: int) -> tuple[int, int]:
15
+ return _padded_latent_dim(height) // WAN_PATCH_SPATIAL, _padded_latent_dim(width) // WAN_PATCH_SPATIAL
16
+
17
+
18
+ def mask_to_token_grid(mask_image: torch.Tensor, width: int, height: int) -> torch.Tensor:
19
+ token_h, token_w = token_grid_size(width, height)
20
+ mask = mask_image[0] if mask_image.ndim == 3 else mask_image
21
+ mask = mask.unsqueeze(0).unsqueeze(0)
22
+ mask = comfy.utils.common_upscale(mask, width, height, "area", "center")
23
+ mask = comfy.utils.common_upscale(mask, token_w, token_h, "area", "center")
24
+ return (mask > 0.5).to(dtype=torch.float32).flatten(2).squeeze(1)
25
+
26
+
27
+ def _latent_frame_weight(video_frame: int, start_frame: int, end_frame: int, crossfade_frames: int) -> float:
28
+ if video_frame < start_frame or video_frame >= end_frame:
29
+ return 0.0
30
+ if crossfade_frames <= 0:
31
+ return 1.0
32
+ weight = 1.0
33
+ if video_frame < start_frame + crossfade_frames:
34
+ weight = min(weight, (video_frame - start_frame + 1) / crossfade_frames)
35
+ if video_frame >= end_frame - crossfade_frames:
36
+ weight = min(weight, (end_frame - video_frame) / crossfade_frames)
37
+ return max(0.0, weight)
38
+
39
+
40
+ def build_timeline_audio_inject_mask(
41
+ width: int,
42
+ height: int,
43
+ length: int,
44
+ segments,
45
+ crossfade_frames: int = 0,
46
+ device=None,
47
+ ) -> torch.Tensor:
48
+ latent_t = ((length - 1) // 4) + 1
49
+ token_h, token_w = token_grid_size(width, height)
50
+ n_tokens = token_h * token_w
51
+ mask = torch.zeros(1, latent_t, n_tokens, 1)
52
+
53
+ for segment in segments:
54
+ tokens = mask_to_token_grid(segment["mask_image"], width, height)
55
+ start_frame = int(segment["start_frame"])
56
+ end_frame = int(segment["end_frame"])
57
+ for latent_idx in range(latent_t):
58
+ vf0 = latent_idx * 4
59
+ vf1 = vf0 + 4
60
+ weight = 0.0
61
+ for video_frame in range(vf0, vf1):
62
+ weight = max(weight, _latent_frame_weight(video_frame, start_frame, end_frame, crossfade_frames))
63
+ if weight > 0.0:
64
+ mask[:, latent_idx, :, 0] = torch.maximum(mask[:, latent_idx, :, 0], tokens * weight)
65
+
66
+ if device is not None:
67
+ mask = mask.to(device)
68
+ return mask
ComfyUI-WanBerniniS2V_v2/demo_assets/ComfyUI_00207-audio.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a01ab587948bca585beb7422d10cdb7a0cd05dc06740bb97400e8b894e1417ff
3
+ size 1877791
ComfyUI-WanBerniniS2V_v2/demo_assets/I am The Dude Playing The Dude.wav ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1b1dcc8ec7a4ca3b26196713e65ab214f7a2516778f45d9273829255d098bab2
3
+ size 191604
ComfyUI-WanBerniniS2V_v2/demo_assets/Im the Dude.wav ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4590983cee3e79dda5b59979b84ffa8b811e86c7f19f97483a1fb7ce0c824d97
3
+ size 404056
ComfyUI-WanBerniniS2V_v2/demo_assets/dude.jpg ADDED

Git LFS Details

  • SHA256: a2c1aab95534c8c1fd8611235f20a19b5078b67f47d6f0029f2e8b8d7c41ab6d
  • Pointer size: 131 Bytes
  • Size of remote file: 118 kB
ComfyUI-WanBerniniS2V_v2/model_patch.py ADDED
@@ -0,0 +1,226 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import inspect
2
+ import logging
3
+
4
+ import torch
5
+
6
+ import comfy.conds
7
+ import comfy.model_management
8
+ from comfy.ldm.wan.model import AudioInjector_WAN, WanModel_S2V
9
+ from comfy.model_base import WAN22_S2V
10
+
11
+
12
+ def _append_context_latents(self, x, kwargs):
13
+ context_latents = kwargs.get("context_latents", None)
14
+ if context_latents is None:
15
+ return x
16
+ for lat in context_latents:
17
+ cl = self.patch_embedding(lat.float().to(x.device)).to(x.dtype).flatten(2).transpose(1, 2)
18
+ x = torch.cat([x, cl], dim=1)
19
+ return x
20
+
21
+
22
+ def _patch_wan_model_s2v_forward():
23
+ if getattr(WanModel_S2V.forward_orig, "__wan_bernini_s2v_v2_patch__", False):
24
+ return
25
+
26
+ try:
27
+ source = inspect.getsource(WanModel_S2V.forward_orig)
28
+ except (OSError, TypeError):
29
+ source = ""
30
+ if "context_latents" in source and getattr(WanModel_S2V.forward_orig, "__wan_bernini_s2v_patch__", False):
31
+ WanModel_S2V.forward_orig.__wan_bernini_s2v_v2_patch__ = True
32
+ return
33
+
34
+ original = WanModel_S2V.forward_orig
35
+
36
+ def forward_orig(
37
+ self,
38
+ x,
39
+ t,
40
+ context,
41
+ audio_embed=None,
42
+ reference_latent=None,
43
+ control_video=None,
44
+ reference_motion=None,
45
+ clip_fea=None,
46
+ freqs=None,
47
+ transformer_options={},
48
+ **kwargs,
49
+ ):
50
+ if audio_embed is not None:
51
+ num_embeds = x.shape[-3] * 4
52
+ audio_emb_global, audio_emb = self.casual_audio_encoder(audio_embed[:, :, :, :num_embeds])
53
+ else:
54
+ audio_emb = None
55
+ audio_emb_global = None
56
+
57
+ bs, _, time, height, width = x.shape
58
+ x = self.patch_embedding(x.float()).to(x.dtype)
59
+ if control_video is not None:
60
+ x = x + self.cond_encoder(control_video)
61
+
62
+ if t.ndim == 1:
63
+ t = t.unsqueeze(1).repeat(1, x.shape[2])
64
+
65
+ grid_sizes = x.shape[2:]
66
+ x = x.flatten(2).transpose(1, 2)
67
+ seq_len = x.size(1)
68
+
69
+ cond_mask_weight = comfy.model_management.cast_to(self.trainable_cond_mask.weight, dtype=x.dtype, device=x.device).unsqueeze(1).unsqueeze(1)
70
+ x = x + cond_mask_weight[0]
71
+ x = _append_context_latents(self, x, kwargs)
72
+
73
+ if reference_latent is not None:
74
+ ref = self.patch_embedding(reference_latent.float()).to(x.dtype)
75
+ ref = ref.flatten(2).transpose(1, 2)
76
+ freqs_ref = self.rope_encode(reference_latent.shape[-3], reference_latent.shape[-2], reference_latent.shape[-1], t_start=max(30, time + 9), device=x.device, dtype=x.dtype)
77
+ ref = ref + cond_mask_weight[1]
78
+ x = torch.cat([x, ref], dim=1)
79
+ freqs = torch.cat([freqs, freqs_ref], dim=1)
80
+ t = torch.cat([t, torch.zeros((t.shape[0], reference_latent.shape[-3]), device=t.device, dtype=t.dtype)], dim=1)
81
+
82
+ if reference_motion is not None:
83
+ motion_encoded, freqs_motion = self.frame_packer(reference_motion, self)
84
+ motion_encoded = motion_encoded + cond_mask_weight[2]
85
+ x = torch.cat([x, motion_encoded], dim=1)
86
+ freqs = torch.cat([freqs, freqs_motion], dim=1)
87
+ t = torch.repeat_interleave(t, 2, dim=1)
88
+ t = torch.cat([t, torch.zeros((t.shape[0], 3), device=t.device, dtype=t.dtype)], dim=1)
89
+
90
+ from comfy.ldm.wan.model import sinusoidal_embedding_1d
91
+
92
+ e = self.time_embedding(
93
+ sinusoidal_embedding_1d(self.freq_dim, t.flatten()).to(dtype=x[0].dtype))
94
+ e = e.reshape(t.shape[0], -1, e.shape[-1])
95
+ e0 = self.time_projection(e).unflatten(2, (6, self.dim))
96
+
97
+ context = self.text_embedding(context)
98
+
99
+ patches_replace = transformer_options.get("patches_replace", {})
100
+ blocks_replace = patches_replace.get("dit", {})
101
+ transformer_options["total_blocks"] = len(self.blocks)
102
+ transformer_options["block_type"] = "double"
103
+ for i, block in enumerate(self.blocks):
104
+ transformer_options["block_index"] = i
105
+ if ("double_block", i) in blocks_replace:
106
+ def block_wrap(args):
107
+ out = {}
108
+ out["img"] = block(args["img"], context=args["txt"], e=args["vec"], freqs=args["pe"], transformer_options=args["transformer_options"])
109
+ return out
110
+ out = blocks_replace[("double_block", i)]({"img": x, "txt": context, "vec": e0, "pe": freqs, "transformer_options": transformer_options}, {"original_block": block_wrap})
111
+ x = out["img"]
112
+ else:
113
+ x = block(x, e=e0, freqs=freqs, context=context, transformer_options=transformer_options)
114
+ if audio_emb is not None:
115
+ inject_scale = kwargs.get("audio_inject_scale", 1.0)
116
+ if isinstance(inject_scale, torch.Tensor):
117
+ inject_scale = inject_scale.reshape(-1)[0].item()
118
+ x = self.audio_injector(
119
+ x, i, audio_emb, audio_emb_global, seq_len,
120
+ scale=inject_scale,
121
+ token_mask=kwargs.get("audio_inject_mask", None),
122
+ )
123
+ x = self.head(x, e)
124
+ x = self.unpatchify(x, grid_sizes)
125
+ return x
126
+
127
+ forward_orig.__wan_bernini_s2v_v2_patch__ = True
128
+ forward_orig.__wan_bernini_s2v_patch__ = True
129
+ forward_orig.__wan_bernini_s2v_original__ = original
130
+ WanModel_S2V.forward_orig = forward_orig
131
+
132
+
133
+ def _patch_audio_injector():
134
+ if getattr(AudioInjector_WAN.forward, "__wan_bernini_s2v_v2_masked_patch__", False):
135
+ return
136
+
137
+ original_forward = AudioInjector_WAN.forward
138
+
139
+ def forward(self, x, block_id, audio_emb, audio_emb_global, seq_len, scale=1.0, token_mask=None):
140
+ if token_mask is None:
141
+ return original_forward(self, x, block_id, audio_emb, audio_emb_global, seq_len, scale=scale)
142
+
143
+ audio_attn_id = self.injected_block_id.get(block_id, None)
144
+ if audio_attn_id is None:
145
+ return x
146
+
147
+ from einops import rearrange
148
+
149
+ num_frames = audio_emb.shape[1]
150
+ input_hidden_states = rearrange(x[:, :seq_len], "b (t n) c -> (b t) n c", t=num_frames)
151
+ if self.enable_adain and self.adain_mode == "attn_norm":
152
+ audio_emb_global = rearrange(audio_emb_global, "b t n c -> (b t) n c")
153
+ adain_hidden_states = self.injector_adain_layers[audio_attn_id](input_hidden_states, temb=audio_emb_global[:, 0])
154
+ attn_hidden_states = adain_hidden_states
155
+ else:
156
+ attn_hidden_states = self.injector_pre_norm_feat[audio_attn_id](input_hidden_states)
157
+
158
+ if audio_emb.dim() == 3:
159
+ attn_audio_emb = rearrange(audio_emb, "b t c -> (b t) 1 c", t=num_frames)
160
+ else:
161
+ attn_audio_emb = rearrange(audio_emb, "b t n c -> (b t) n c", t=num_frames)
162
+
163
+ residual_out = self.injector[audio_attn_id](x=attn_hidden_states, context=attn_audio_emb)
164
+ residual_out = rearrange(residual_out, "(b t) n c -> b (t n) c", t=num_frames)
165
+
166
+ if token_mask.ndim == 4:
167
+ token_mask = token_mask.flatten(1, 2)
168
+ if token_mask.shape[1] == residual_out.shape[1]:
169
+ residual_out = residual_out * token_mask.to(device=residual_out.device, dtype=residual_out.dtype)
170
+ else:
171
+ logging.warning(
172
+ "ComfyUI-WanBerniniS2V_v2: mask length %s does not match token count %s; using global audio injection",
173
+ token_mask.shape[1],
174
+ residual_out.shape[1],
175
+ )
176
+
177
+ x[:, :seq_len] = x[:, :seq_len] + residual_out * scale
178
+ return x
179
+
180
+ forward.__wan_bernini_s2v_v2_masked_patch__ = True
181
+ forward.__wan_bernini_s2v_masked_patch__ = True
182
+ forward.__wan_bernini_s2v_masked_original__ = original_forward
183
+ AudioInjector_WAN.forward = forward
184
+
185
+
186
+ def _patch_wan22_s2v_extra_conds():
187
+ if getattr(WAN22_S2V.extra_conds, "__wan_bernini_s2v_v2_masked_patch__", False):
188
+ return
189
+
190
+ original_extra_conds = WAN22_S2V.extra_conds
191
+
192
+ def extra_conds(self, **kwargs):
193
+ out = original_extra_conds(self, **kwargs)
194
+ audio_inject_mask = kwargs.get("audio_inject_mask", None)
195
+ if audio_inject_mask is not None:
196
+ out["audio_inject_mask"] = comfy.conds.CONDRegular(audio_inject_mask)
197
+ audio_inject_scale = kwargs.get("audio_inject_scale", None)
198
+ if audio_inject_scale is not None:
199
+ out["audio_inject_scale"] = comfy.conds.CONDRegular(torch.FloatTensor([audio_inject_scale]))
200
+ return out
201
+
202
+ extra_conds.__wan_bernini_s2v_v2_masked_patch__ = True
203
+ extra_conds.__wan_bernini_s2v_masked_patch__ = True
204
+ extra_conds.__wan_bernini_s2v_masked_original__ = original_extra_conds
205
+ WAN22_S2V.extra_conds = extra_conds
206
+
207
+ original_resize = WAN22_S2V.resize_cond_for_context_window
208
+
209
+ def resize_cond_for_context_window(self, cond_key, cond_value, window, x_in, device, retain_index_list=[]):
210
+ if cond_key == "audio_inject_mask":
211
+ mask = cond_value.cond
212
+ if mask.ndim == 4 and mask.shape[1] == x_in.shape[2]:
213
+ return cond_value._copy_with(window.get_tensor(mask, device, dim=1))
214
+ return original_resize(self, cond_key, cond_value, window, x_in, device, retain_index_list=retain_index_list)
215
+
216
+ resize_cond_for_context_window.__wan_bernini_s2v_v2_masked_patch__ = True
217
+ resize_cond_for_context_window.__wan_bernini_s2v_masked_patch__ = True
218
+ resize_cond_for_context_window.__wan_bernini_s2v_masked_original__ = original_resize
219
+ WAN22_S2V.resize_cond_for_context_window = resize_cond_for_context_window
220
+
221
+
222
+ def apply_model_patches():
223
+ _patch_wan_model_s2v_forward()
224
+ _patch_audio_injector()
225
+ _patch_wan22_s2v_extra_conds()
226
+ logging.info("ComfyUI-WanBerniniS2V_v2: applied Bernini S2V model patches")
ComfyUI-WanBerniniS2V_v2/nodes.py ADDED
@@ -0,0 +1,153 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import torch
2
+ from typing_extensions import override
3
+
4
+ import comfy.model_management
5
+ import comfy.utils
6
+ import node_helpers
7
+ from comfy_api.latest import ComfyExtension, io
8
+
9
+ from .audio_mask import build_timeline_audio_inject_mask
10
+ from .wan_audio import apply_timeline_audio_conditioning, resolve_timeline_segment_ranges
11
+
12
+
13
+ def _resize_long_edge(image, max_size, stride=16):
14
+ h, w = image.shape[1], image.shape[2]
15
+ scale = min(max_size / max(h, w), 1.0)
16
+ nh = max(stride, round(h * scale / stride) * stride)
17
+ nw = max(stride, round(w * scale / stride) * stride)
18
+ return comfy.utils.common_upscale(image[:, :, :, :3].movedim(-1, 1), nw, nh, "area", "disabled").movedim(1, -1)
19
+
20
+
21
+ def _build_context_latents(vae, width, height, length, source_video=None, reference_video=None, reference_images=None, ref_max_size=848):
22
+ context = []
23
+ if source_video is not None:
24
+ vid = comfy.utils.common_upscale(source_video[:length, :, :, :3].movedim(-1, 1), width, height, "area", "center").movedim(1, -1)
25
+ context.append(vae.encode(vid[:, :, :, :3]))
26
+
27
+ if reference_video is not None:
28
+ ref_vid = _resize_long_edge(reference_video[:length], ref_max_size)
29
+ context.append(vae.encode(ref_vid[:, :, :, :3]))
30
+
31
+ if reference_images:
32
+ for name in sorted(reference_images):
33
+ imgs = reference_images[name]
34
+ if imgs is None:
35
+ continue
36
+ for i in range(imgs.shape[0]):
37
+ img = _resize_long_edge(imgs[i:i + 1], ref_max_size)
38
+ context.append(vae.encode(img[:, :, :, :3]))
39
+ return context
40
+
41
+
42
+ class BerniniS2VConditioningV2(io.ComfyNode):
43
+ @classmethod
44
+ def define_schema(cls):
45
+ return io.Schema(
46
+ node_id="BerniniS2VConditioningV2",
47
+ display_name="Bernini S2V Conditioning v2",
48
+ category="model/conditioning/bernini",
49
+ description="Bernini in-context conditioning with masked S2V audio for one or two speakers. Requires a Wan 2.2 S2V grafted Bernini-R model. Paint speaker masks on the output frame; reference_image_0 maps to image0 in prompts.",
50
+ inputs=[
51
+ io.Conditioning.Input("positive"),
52
+ io.Conditioning.Input("negative"),
53
+ io.Vae.Input("vae"),
54
+ io.Int.Input("width", default=832, min=16, max=8192, step=16),
55
+ io.Int.Input("height", default=480, min=16, max=8192, step=16),
56
+ io.Int.Input("length", default=81, min=1, max=8192, step=4),
57
+ io.Int.Input("batch_size", default=1, min=1, max=4096),
58
+ io.AudioEncoderOutput.Input("audio_1"),
59
+ io.Mask.Input("mask_1", tooltip="White = speaker 1 lip-sync region on the output frame."),
60
+ io.AudioEncoderOutput.Input("audio_2", optional=True),
61
+ io.Mask.Input("mask_2", optional=True, tooltip="Required when audio_2 is connected."),
62
+ io.Int.Input("speaker_2_start_frame", default=-1, min=-1, max=8192, step=1,
63
+ tooltip="-1 auto-starts speaker 2 when speaker 1 audio ends."),
64
+ io.Image.Input("source_video", optional=True),
65
+ io.Image.Input("reference_video", optional=True),
66
+ io.Autogrow.Input("reference_images", optional=True,
67
+ template=io.Autogrow.TemplatePrefix(
68
+ input=io.Image.Input("reference_image"),
69
+ prefix="reference_image_", min=0, max=8)),
70
+ io.Int.Input("ref_max_size", default=848, min=16, max=8192, step=16, optional=True),
71
+ io.Int.Input("mask_crossfade_frames", default=4, min=0, max=64, step=1,
72
+ tooltip="Softens the mask handoff between speakers. 0 = hard cut."),
73
+ io.Float.Input("audio_inject_scale", default=1.0, min=0.0, max=10.0, step=0.01),
74
+ ],
75
+ outputs=[
76
+ io.Conditioning.Output(display_name="positive"),
77
+ io.Conditioning.Output(display_name="negative"),
78
+ io.Latent.Output(display_name="latent"),
79
+ ],
80
+ )
81
+
82
+ @classmethod
83
+ def execute(
84
+ cls,
85
+ positive,
86
+ negative,
87
+ vae,
88
+ width,
89
+ height,
90
+ length,
91
+ batch_size,
92
+ audio_1,
93
+ mask_1,
94
+ audio_2=None,
95
+ mask_2=None,
96
+ speaker_2_start_frame=-1,
97
+ source_video=None,
98
+ reference_video=None,
99
+ reference_images=None,
100
+ ref_max_size=848,
101
+ mask_crossfade_frames=4,
102
+ audio_inject_scale=1.0,
103
+ ) -> io.NodeOutput:
104
+ if audio_1 is None:
105
+ raise ValueError("Bernini S2V Conditioning v2 requires audio_1.")
106
+ if mask_1 is None:
107
+ raise ValueError("Bernini S2V Conditioning v2 requires mask_1.")
108
+ if audio_2 is not None and mask_2 is None:
109
+ raise ValueError("mask_2 is required when audio_2 is connected.")
110
+
111
+ segments = [{"audio_encoder_output": audio_1, "start_frame": 0, "mask_image": mask_1}]
112
+ if audio_2 is not None:
113
+ segments.append({
114
+ "audio_encoder_output": audio_2,
115
+ "start_frame": speaker_2_start_frame,
116
+ "mask_image": mask_2,
117
+ })
118
+
119
+ latent = torch.zeros(
120
+ [batch_size, 16, ((length - 1) // 4) + 1, height // 8, width // 8],
121
+ device=comfy.model_management.intermediate_device())
122
+
123
+ context = _build_context_latents(vae, width, height, length, source_video, reference_video, reference_images, ref_max_size)
124
+ if context:
125
+ positive = node_helpers.conditioning_set_values(positive, {"context_latents": context})
126
+ negative = node_helpers.conditioning_set_values(negative, {"context_latents": context})
127
+
128
+ positive, negative = apply_timeline_audio_conditioning(positive, negative, length, segments)
129
+ resolved_segments = resolve_timeline_segment_ranges(length, segments)
130
+ cond_values = {
131
+ "audio_inject_scale": audio_inject_scale,
132
+ "audio_inject_mask": build_timeline_audio_inject_mask(
133
+ width, height, length, resolved_segments,
134
+ crossfade_frames=mask_crossfade_frames,
135
+ device=comfy.model_management.intermediate_device(),
136
+ ),
137
+ }
138
+ positive = node_helpers.conditioning_set_values(positive, cond_values)
139
+ negative_values = dict(cond_values)
140
+ negative_values["audio_inject_mask"] = negative_values["audio_inject_mask"] * 0.0
141
+ negative = node_helpers.conditioning_set_values(negative, negative_values)
142
+
143
+ return io.NodeOutput(positive, negative, {"samples": latent})
144
+
145
+
146
+ class WanBerniniS2VV2Extension(ComfyExtension):
147
+ @override
148
+ async def get_node_list(self) -> list[type[io.ComfyNode]]:
149
+ return [BerniniS2VConditioningV2]
150
+
151
+
152
+ async def comfy_entrypoint() -> WanBerniniS2VV2Extension:
153
+ return WanBerniniS2VV2Extension()
ComfyUI-WanBerniniS2V_v2/wan_audio.py ADDED
@@ -0,0 +1,89 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import math
2
+
3
+ import torch
4
+
5
+ import node_helpers
6
+ from comfy_extras.nodes_wan import get_audio_embed_bucket_fps, linear_interpolation
7
+
8
+ WAN_AUDIO_INPUT_FPS = 50
9
+ WAN_AUDIO_VIDEO_RATE = 30
10
+ WAN_AUDIO_FPS = 16
11
+ WAN_AUDIO_SAMPLE_RATE = 16000
12
+
13
+
14
+ def _audio_feat(audio_encoder_output):
15
+ feat = torch.cat(audio_encoder_output["encoded_audio_all_layers"])
16
+ return linear_interpolation(feat, input_fps=WAN_AUDIO_INPUT_FPS, output_fps=WAN_AUDIO_VIDEO_RATE)
17
+
18
+
19
+ def audio_encoder_output_video_frames(audio_encoder_output, fps=WAN_AUDIO_FPS):
20
+ audio_samples = audio_encoder_output.get("audio_samples")
21
+ if audio_samples is not None:
22
+ return max(1, int(round(audio_samples / float(WAN_AUDIO_SAMPLE_RATE) * fps)))
23
+ feat = _audio_feat(audio_encoder_output)
24
+ return max(1, int(round(feat.shape[1] * fps / WAN_AUDIO_VIDEO_RATE)))
25
+
26
+
27
+ def _permute_audio_embed_bucket(audio_embed_bucket):
28
+ audio_embed_bucket = audio_embed_bucket.unsqueeze(0)
29
+ if len(audio_embed_bucket.shape) == 3:
30
+ return audio_embed_bucket.permute(0, 2, 1)
31
+ return audio_embed_bucket.permute(0, 2, 3, 1)
32
+
33
+
34
+ def build_timeline_audio_embed(length, segments):
35
+ latent_t = ((length - 1) // 4) + 1
36
+ batch_frames = latent_t * 4
37
+ total_feat_frames = int(math.ceil(batch_frames * WAN_AUDIO_VIDEO_RATE / WAN_AUDIO_FPS))
38
+
39
+ composite = None
40
+ cursor_auto = 0
41
+ for segment in segments:
42
+ feat = _audio_feat(segment["audio_encoder_output"])
43
+ if composite is None:
44
+ composite = torch.zeros(
45
+ feat.shape[0], total_feat_frames, feat.shape[2],
46
+ dtype=feat.dtype, device=feat.device)
47
+
48
+ start_frame = segment.get("start_frame", -1)
49
+ if start_frame is None or start_frame < 0:
50
+ start_frame = cursor_auto
51
+ else:
52
+ start_frame = int(start_frame)
53
+
54
+ start_feat = int(round(start_frame * WAN_AUDIO_VIDEO_RATE / WAN_AUDIO_FPS))
55
+ copy_len = min(feat.shape[1], total_feat_frames - start_feat)
56
+ if copy_len > 0 and start_feat < total_feat_frames:
57
+ composite[:, start_feat:start_feat + copy_len, :] = feat[:, :copy_len, :]
58
+
59
+ cursor_auto = start_frame + audio_encoder_output_video_frames(segment["audio_encoder_output"])
60
+
61
+ audio_embed_bucket, _ = get_audio_embed_bucket_fps(
62
+ composite, fps=WAN_AUDIO_FPS, batch_frames=batch_frames, m=0, video_rate=WAN_AUDIO_VIDEO_RATE)
63
+ audio_embed_bucket = _permute_audio_embed_bucket(audio_embed_bucket)
64
+ return audio_embed_bucket[:, :, :, :batch_frames]
65
+
66
+
67
+ def resolve_timeline_segment_ranges(length, segments):
68
+ batch_frames = (((length - 1) // 4) + 1) * 4
69
+ resolved = []
70
+ cursor_auto = 0
71
+ for segment in segments:
72
+ start_frame = segment.get("start_frame", -1)
73
+ if start_frame is None or start_frame < 0:
74
+ start_frame = cursor_auto
75
+ else:
76
+ start_frame = int(start_frame)
77
+ end_frame = min(batch_frames, start_frame + audio_encoder_output_video_frames(segment["audio_encoder_output"]))
78
+ resolved.append({**segment, "start_frame": start_frame, "end_frame": end_frame})
79
+ cursor_auto = end_frame
80
+ return resolved
81
+
82
+
83
+ def apply_timeline_audio_conditioning(positive, negative, length, segments):
84
+ audio_embed_bucket = build_timeline_audio_embed(length, segments)
85
+ if audio_embed_bucket is None or audio_embed_bucket.shape[3] <= 0:
86
+ return positive, negative
87
+ positive = node_helpers.conditioning_set_values(positive, {"audio_embed": audio_embed_bucket})
88
+ negative = node_helpers.conditioning_set_values(negative, {"audio_embed": audio_embed_bucket * 0.0})
89
+ return positive, negative