leeting770708 commited on
Commit
4501654
Β·
verified Β·
1 Parent(s): c1c6eed

Copy files from models/Comfy-Org/ace_step_1.5_ComfyUI_files

Browse files
.gitattributes CHANGED
@@ -33,5 +33,3 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
- model_structure.png filter=lfs diff=lfs merge=lfs -text
37
- teaser.png filter=lfs diff=lfs merge=lfs -text
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
README.md CHANGED
@@ -1,37 +1,39 @@
1
  ---
2
  license: apache-2.0
3
- pipeline_tag: video-to-video
 
 
 
 
4
  ---
5
 
6
- This repository contains the weights of [ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing](https://arxiv.org/abs/2506.21448).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7
 
8
- Project Page: https://thinksound-project.github.io/.
9
-
10
- Paper: https://huggingface.co/papers/2506.21448
11
-
12
- Github: https://github.com/FunAudioLLM/ThinkSound
13
-
14
- <img src="./teaser.png" alt="model_structure" style="zoom:20%;" />
15
-
16
- ## Abstract
17
- While end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like professionals in the creative industries, such generation requires sophisticated reasoning about items such as visual dynamics, acoustic environments, and temporal relationships. We present ThinkSound, a novel framework that leverages Chain-of-Thought (CoT) reasoning to enable stepwise, interactive audio generation and editing for videos. Our approach decomposes the process into three complementary stages: foundational foley generation that creates semantically coherent soundscapes, interactive object-centric refinement through precise user interactions, and targeted editing guided by natural language instructions. At each stage, a multimodal large language model generates contextually aligned CoT reasoning that guides a unified audio foundation model. Furthermore, we introduce AudioCoT, a comprehensive dataset with structured reasoning annotations that establishes connections between visual content, textual descriptions, and sound synthesis. Experiments demonstrate that ThinkSound achieves state-of-the-art performance in video-to-audio generation across both audio metrics and CoT metrics and excels in out-of-distribution Movie Gen Audio benchmark. The demo page is available at https://ThinkSound-Project.github.io.
18
-
19
- ## Model Overview
20
-
21
- <img src="./model_structure.png" alt="model_structure" style="zoom:40%;" />
22
-
23
- ## Citation
24
-
25
- If you find our work useful, please cite our paper:
26
 
27
- ```bibtex
28
- @misc{liu2025thinksoundchainofthoughtreasoningmultimodal,
29
- title={ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing},
30
- author={Huadai Liu and Jialei Wang and Kaicheng Luo and Wen Wang and Qian Chen and Zhou Zhao and Wei Xue},
31
- year={2025},
32
- eprint={2506.21448},
33
- archivePrefix={arXiv},
34
- primaryClass={eess.AS},
35
- url={https://arxiv.org/abs/2506.21448},
36
- }
37
- ```
 
1
  ---
2
  license: apache-2.0
3
+ tags:
4
+ - comfyui
5
+ - diffusion-single-file
6
+ base_model:
7
+ - ACE-Step/Ace-Step1.5
8
  ---
9
 
10
+ # ACE-Step 1.5
11
+
12
+ Repackaged model files for ComfyUI.
13
+
14
+ Original model repository: https://huggingface.co/ACE-Step/Ace-Step1.5
15
+
16
+ Place the files in the following folders:
17
+
18
+ ```
19
+ πŸ“‚ ComfyUI/
20
+ β”œβ”€β”€ πŸ“‚ models/
21
+ β”‚ β”œβ”€β”€ πŸ“‚ checkpoints/
22
+ β”‚ β”‚ └── ace_step_1.5_turbo_aio.safetensors
23
+ β”‚ β”œβ”€β”€ πŸ“‚ diffusion_models/
24
+ β”‚ β”‚ β”œβ”€β”€ acestep_v1.5_base.safetensors
25
+ β”‚ β”‚ β”œβ”€β”€ acestep_v1.5_turbo.safetensors
26
+ β”‚ β”‚ β”œβ”€β”€ acestep_v1.5_xl_base_bf16.safetensors
27
+ β”‚ β”‚ β”œβ”€β”€ acestep_v1.5_xl_sft_bf16.safetensors
28
+ β”‚ β”‚ └── acestep_v1.5_xl_turbo_bf16.safetensors
29
+ β”‚ β”œβ”€β”€ πŸ“‚ text_encoders/
30
+ β”‚ β”‚ β”œβ”€β”€ qwen_0.6b_ace15.safetensors
31
+ β”‚ β”‚ β”œβ”€β”€ qwen_1.7b_ace15.safetensors
32
+ β”‚ β”‚ └── qwen_4b_ace15.safetensors
33
+ β”‚ β”œβ”€β”€ πŸ“‚ vae/
34
+ β”‚ β”‚ └── ace_1.5_vae.safetensors
35
+ ```
36
 
37
+ ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
38
 
39
+ These are repackaged files to work with ComfyUI, original model repo is: https://huggingface.co/ACE-Step/Ace-Step1.5
 
 
 
 
 
 
 
 
 
 
checkpoints/ace_step_1.5_turbo_aio.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:67b0f43aa5c51c840bd0228e6a935d8ff416ec87e5df2fc0637da17a561252bc
3
+ size 10025478736
split_files/diffusion_models/acestep_v1.5_base.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4177f600501a6d4bd81cadaa0abac557ffd15c54e5c8cb52053cdb24a0844d6b
3
+ size 4787825604
split_files/diffusion_models/acestep_v1.5_turbo.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3f6e0797fad420a39bd33979eb6e840e30989e34a3794e843d23b60ec6e422d7
3
+ size 4787825604
split_files/diffusion_models/acestep_v1.5_xl_base_bf16.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:56bf816fc9a69a5f45635e867b2ad742e1e648eb51fadb7d124cb8332d2e0940
3
+ size 9974719930
split_files/diffusion_models/acestep_v1.5_xl_sft_bf16.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3c05ae268353b3540fb1fd7db4fd77ffbda9802ec641b624e15648e030ecf3ce
3
+ size 9974719930
split_files/diffusion_models/acestep_v1.5_xl_turbo_bf16.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:86a1afb0a1f711f0e3304ff65d874df3ae6783db683dcf982513fb9b6d14ae71
3
+ size 9974719892
split_files/text_encoders/qwen_0.6b_ace15.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fd4590c82153b8ddb67e15a2e7aaa8afa8b83a858c8a9b82a4831063156aa7a7
3
+ size 1191588248
split_files/text_encoders/qwen_1.7b_ace15.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ed63e9247d1f55f3ace04fa11e95b085fc82d459c82c5626f0b2e37b91ebd710
3
+ size 3708523360
split_files/text_encoders/qwen_4b_ace15.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ffe5ffb855086c2ab55e467e9859fb01894781020a0376484dd19de166b79873
3
+ size 8379154232
split_files/vae/ace_1.5_vae.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6de92e3a862acd287e08b024ac90f0783a8635451b728721a33ff03565bcb2bb
3
+ size 337431732