File size: 14,545 Bytes
ae83015
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
# Diffusers

<p align="left">
</p>

- [SD3/SD3.5](#jump1)
  - [模型介绍](#模型介绍)
  - [微调](#微调)
    - [环境搭建](#环境搭建)
    - [微调](#jump2)
    - [性能](#性能)
  - [推理](#推理)
    - [环境搭建及运行](#环境搭建及运行)
  - [环境变量声明](#环境变量声明)
- [引用](#引用)
  - [公网地址说明](#公网地址说明)

<a id="jump1"></a>

# Stable Diffusion 3 & Stable Diffusion 3.5

## 模型介绍

扩散模型(Diffusion Models)是一种生成模型,可生成各种各样的高分辨率图像。Diffusers 是 HuggingFace 发布的模型套件,是最先进的预训练扩散模型的首选库,用于生成图像,音频,甚至分子的3D结构。套件包含基于扩散模型的多种模型,提供了各种下游任务的训练与推理的实现。

- 参考实现:

  ```shell
  url=https://github.com/huggingface/diffusers
  commit_id=5f724735437d91ed05304da478f3b2022fe3f6fb
  ```

## 微调

### 环境搭建

【模型开发时推荐使用配套的环境版本】

请参考[安装指南](https://gitee.com/ascend/MindSpeed-MM/blob/master/docs/user-guide/installation.md)

1. 软件与驱动安装


    ```bash
    # python3.8
    conda create -n test python=3.8
    conda activate test

    # 安装 torch 和 torch_npu,注意要选择对应python版本、x86或arm的torch、torch_npu及apex包
    pip install torch-2.1.0-cp38-cp38m-manylinux2014_aarch64.whl 
    pip install torch_npu-2.1.0*-cp38-cp38m-linux_aarch64.whl
    
    # apex for Ascend 参考 https://gitee.com/ascend/apex
    # 建议从原仓编译安装

    # 将shell脚本中的环境变量路径修改为真实路径,下面为参考路径
    source /usr/local/Ascend/ascend-toolkit/set_env.sh
    ```

2. 克隆仓库到本地服务器

    ```shell
    git clone --branch 2.1.0 https://gitee.com/ascend/MindSpeed-MM.git
    ```

3. 模型搭建

    3.1 【下载 SD3/SD3.5 [GitHub参考实现](https://github.com/huggingface/diffusers) 或 [适配昇腾AI处理器的实现](https://gitee.com/ascend/ModelZoo-PyTorch.git) 或 在模型根目录下执行以下命令,安装模型对应PyTorch版本需要的依赖】

    ```shell
    git clone https://github.com/huggingface/diffusers.git
    cd diffusers
    git checkout 5f724735437d91ed05304da478f3b2022fe3f6fb
    cp -r ../MindSpeed-MM/examples/diffusers/sd3 ./sd3
    ```

    【主要代码路径】

    ```shell
    code_path=examples/dreambooth/
    ```

    3.2【安装其余依赖库】

    ```shell
    pip install -e .
    vim examples/dreambooth/requirements_sd3.txt #修改版本:torchvision==0.16.0, torch==2.1.0, accelerate==0.33.0, 添加deepspeed==0.15.2
    pip install -r examples/dreambooth/requirements_sd3.txt # 安装对应依赖
    ```

<a id="jump2"></a>

## 微调

1. 【准备微调数据集】

    用户需自行获取并解压[pokemon-blip-captions](https://huggingface.co/datasets/lambdalabs/pokemon-blip-captions/tree/main)数据集,并在以下启动shell脚本中将`dataset_name`参数设置为本地数据集的绝对路径

    ```shell
    vim sd3/finetune_sd3_dreambooth_deepspeed_**16.sh
    vim sd3/finetune_sd3_dreambooth_fp16.sh
    ```

    ```shell
    dataset_name="pokemon-blip-captions" # 数据集 路径
    ```

   - pokemon-blip-captions数据集格式如下:

    ```shell
    pokemon-blip-captions
    ├── dataset_infos.json
    ├── README.MD
    └── data
          └── train-001.parquet
    ```

    - 只包含图片的训练数据集,如非deepspeed脚本使用训练数据集dog:[下载地址](https://huggingface.co/datasets/diffusers/dog-example),在shell启动脚本中将`input_dir`参数设置为本地数据集绝对路径>

    ```shell
    input_dir="dog" # 数据集路径
    ```

    ```shell
    dog
    ├── alvan-nee-*****.jpeg
    ├── alvan-nee-*****.jpeg
    ```

    > **说明:**
    >该数据集的训练过程脚本只作为一种参考示例。
    >

2. 【配置 SD3/SD3.5 微调脚本】

    【SD3】
    用户可访问huggingface官网自行下载[sd3-medium模型](https://huggingface.co/stabilityai/stable-diffusion-3-medium-diffusers/tree/main) `model_name`模型

    ```bash
    export model_name="stabilityai/stable-diffusion-3-medium-diffusers" # 预训练模型路径
    ```

    【SD3.5】
    用户可访问huggingface官网自行下载[sd3.5-large模型](https://huggingface.co/stabilityai/stable-diffusion-3.5-large/tree/main) `model_name`模型

    ```bash
    export model_name="stabilityai/stable-diffusion-3.5-large" # 预训练模型路径
    ```

    获取对应的微调模型后,在以下shell启动脚本中将`model_name`参数设置为本地预训练模型绝对路径SD3与SD3.5为相同脚本

    ```shell
    scripts_path="./sd3" # 模型根目录(模型文件夹名称)
    model_name="stabilityai/stable-diffusion-3-medium-diffusers" # 预训练模型路径 (此为sd3)
    dataset_name="pokemon-blip-captions" 
    batch_size=4
    num_processors=8 # 卡数(为计算FPS使用,yaml文件里需同步修改)
    max_train_steps=2000
    mixed_precision="bf16" # 混精
    resolution=1024
    config_file="${scripts_path}/${mixed_precision}_accelerate_config.yaml"

    # accelerate launch --config_file ${config_file} \ 目录下
    --dataloader_num_workers=0 \ # 请基于系统配置与数据大小进行调整num workers
    ```

    数据集选择:如果选择默认[原仓数据集](https://huggingface.co/datasets/diffusers/dog-example),需修改两处`dataset_name`为`input_dir`:

    ```shell
    input_dir="dog"

    # accelerator 修改 --dataset_name=#dataset_name
    --instance_data_dir=$input_dir
    ```

    | 数据集 | 路径设置 | accelerate 设置 |
    |:----------:|:----------:|:----------:|
    | dog | input_dir="dog" | --instance_data_dir=$input_dir; --instance_prompt="A photo of sks dog" |
    | pokemon | dataset_name="pokemon-blip-captions" | --dataset_name=$dataset_name --caption_column="text"; --instance_prompt="A photo of pokemon" |

    修改`fp16_accelerate_config.yaml`的`deepspeed_config_file`的路径:

    ```shell
    vim sd3/fp16_accelerate_config.yaml
    # 修改:
    deepspeed_config_file: ./sd3/deepspeed_fp16.json # deepspeed JSON文件路径
    ```

3. 【Optional】Ubuntu系统需在`train_dreambooth_sd3.py`1705行附近 与 `train_dreambooth_lora_sd3.py`1861行附近 添加 `accelerator.print("")`

    ```shell
    vim examples/dreambooth/train_dreambooth_sd3.py
    # 或
    vim examples/dreambooth/train_dreambooth_lora_sd3.py
    ```

    如下:

    ```python
    if global_step >= args.max_train_steps: # 原代码
      break
    accelerator.print("") # 添加
    ```

4. 【如需保存checkpointing请修改代码】

    ```shell
    vim examples/dreambooth/train_dreambooth_sd3.py
    ```

    - 在文件上方的import栏增加`DistributedType``from accelerate import Accelerator`后 (30行附近)
    -`if accelerator.is_main_process`后增加 `or accelerator.distributed_type == DistributedType.DEEPSPEED`(dreambooth在1681行附近),并在`if args.checkpoints_total_limit is not None`后增加`and accelerator.is_main_process`

    ```python
    from accelerate import Accelerator, DistributedType
    # from accelerate import Accelerator # 原代码
    from accelerate.logging import get_logger # 原代码
     
    if accelerator.is_main_process or accelerator.distributed_type == DistributedType.DEEPSPEED:
    # if accelerator.is_main_process: # 原代码 1681/1833行附近
      if global_step % args.checkpointing_steps == 0:  # 原代码 不进行修改
        if args.checkpoints_total_limit is not None and accelerator.is_main_process: # 添加
    ```

5. 【修改文件】

    ```shell
    vim examples/dreambooth/train_dreambooth_sd3.py
    # 或
    vim examples/dreambooth/train_dreambooth_lora_sd3.py
    ```

    在log_validation里修改`pipeline = pipeline.to(accelerator.device)`,`train_dreambooth_sd3.py`在174行附近`train_dreambooth_lora_sd3.py`在198行附近

    ```python
    # 修改pipeline为:
    pipeline = pipeline.to(accelerator.device, dtype=torch_dtype)
    # pipeline = pipeline.to(accelerator.device) # 原代码
    ```

6. 【启动 SD3 微调脚本】

    本任务主要提供**混精fp16**和**混精bf16**dreambooth和dreambooth+lora的**8卡**训练脚本,使用与不使用**deepspeed**分布式训练。

    ```shell
    bash sd3/finetune_sd3_dreambooth_deepspeed_**16.sh #使用deepspeed,dreambooth微调 
    bash sd3/finetune_sd3_dreambooth_lora_deepspeed_fp16.sh #使用deepspeed,dreambooth微调 (sd3.5)
    bash sd3/finetune_sd3_dreambooth_fp16.sh #无使用deepspeed,dreambooth微调
    bash sd3/finetune_sd3_dreambooth_lora_fp16.sh #无使用deepspeed,dreambooth+lora微调 (sd3)
    ```

### 性能

#### 吞吐

SD3 在 **昇腾芯片****参考芯片** 上的性能对比:

| 芯片 | 卡数 |     任务     |  FPS  | batch_size | AMP_Type | Resolution | Torch_Version | deepspeed |
|:---:|:---:|:----------:|:-----:|:----------:|:---:|:---:|:---:|:---:|
| Atlas 900 A2 PODc | 8p | Dreambooth-全参微调  |   16.09 |     4      | bf16 | 1024 | 2.1 | ✔ |
| 竞品A | 8p | Dreambooth-全参微调  |  16.01 |     4      | bf16 | 1024 | 2.1 | ✔ |
| Atlas 900 A2 PODc | 8p | Dreambooth-全参微调 |  15.16 |     4      | fp16 | 1024 | 2.1 | ✔ |
| 竞品A | 8p | Dreambooth-全参微调 |   15.53 |     4      | fp16 | 1024 | 2.1 | ✔ |
| Atlas 900 A2 PODc |8p | Dreambooth-全参微调 | 3.11  | 1 | fp16 | 1024 | 2.1 | ✘ |
| 竞品A | 8p | Dreambooth-全参微调 | 3.71 | 1 | fp16 | 1024 | 2.1 | ✘ |
| Atlas 900 A2 PODc |8p | DreamBooth-LoRA | 108.8 | 8 | fp16 | 1024 | 2.1 | ✘ |
| 竞品A | 8p | DreamBooth-LoRA | 110.69 | 8 | fp16 | 1024 | 2.1 | ✘ |

SD3.5 在 **昇腾芯片****参考芯片** 上的性能对比:

| 芯片 | 卡数 |     任务     |  FPS  | batch_size | AMP_Type | Resolution | Torch_Version | deepspeed | gradient checkpointing |
|:---:|:---:|:----------:|:-----:|:----------:|:---:|:---:|:---:|:---:|:---:|
| Atlas 900 A2 PODc | 8p | Dreambooth-全参微调  |   26.24 |     8      | bf16 | 512 | 2.1 | ✔ | ✔ |
| 竞品A | 8p | Dreambooth-全参微调  |  28.33 |     8      | bf16 | 512 | 2.1 | ✔ | ✔ |
| Atlas 900 A2 PODc | 8p | Dreambooth-Lora |  47.93 |     8      | fp16 | 512 | 2.1 | ✔ | ✘ |
| 竞品A | 8p | Dreambooth-Lora |   47.95 |     8      | fp16 | 512 | 2.1 | ✔ | ✘ |

## 推理

### 环境搭建及运行

  **同微调对应章节**

 【运行推理的脚本】

  图生图推理脚本需先准备图片:[下载地址](https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png)
  修改推理脚本中预训练模型路径以及图生图推理脚本中的本地图片加载路径
  调用推理脚本

  ```shell
  cd sd3/ # # 进入sd3目录
  ```

【SD3/SD3.5模型推理】

```shell
vim infer_sd3_text2img.py # 进入运行T2I推理的Python文件
# 或
vim infer_sd3_img2img.py # 进入运行I2I推理的Python文件
```

  1. 修改路径

      ```python
      MODEL_PATH = "stabilityai/stable-diffusion-3.5-large"  # 路径可选择sd3/sd3.5模型权重 或 Dreambooth 微调后输出模型
      DTYPE = torch.float16 # 可选择混精模式
      ```

  2. 运行代码

      ```shell
      python infer_sd3_text2img.py  # 单卡推理,文生图
      python infer_sd3_img2img.py   # 单卡推理,图生图
      ```

  【lora微调SD3模型推理】

  ```shell
  vim infer_sd3_text2img_lora.py
  ```

  1. 修改路径

      ```python
      MODEL_PATH = "stabilityai/stable-diffusion-3.5-large"  # 路径可选择sd3/sd3.5模型权重 或 Dreambooth 微调后输出模型
      LORA_WEIGHTS = "./output/pytorch_lora_weights.safetensors"  # LoRA权重路径
      ```

  2. 运行代码

      ```shell
      python infer_sd3_text2img_lora.py
      ```

  【分布式推理】

  ```shell
  vim infer_sd3_text2img_distrib.py
  ```

- 修改模型权重路径 model_path为模型权重路径或微调后的权重路径
- 如lora微调 可将lora_weights修改为Lora权重路径

  ```python
  model_path = "stabilityai/stable-diffusion-3.5-large"  # 模型权重/微调权重路径
  lora_weights = "/pytorch_lora_weights.safetensors"  # Lora权重路径
  ```

- 启动分布式推理脚本
  
  - 因使用accelerate进行分布式推理,config可设置:`--num_processes=卡数``num_machines=机器数````shell
  accelerate launch --num_processes=4 infer_sd3_text2img_distrib.py # 单机四卡进行分布式推理
  ```

## 使用基线数据集进行评估

## 环境变量声明
ASCEND_SLOG_PRINT_TO_STDOUT: 是否开启日志打印, 0:关闭日志打屏,1:开启日志打屏  
ASCEND_GLOBAL_LOG_LEVEL: 设置应用类日志的日志级别及各模块日志级别,仅支持调试日志。0:对应DEBUG级别,1:对应INFO级别,2:对应WARNING级别,3:对应ERROR级别,4:对应NULL级别,不输出日志  
ASCEND_GLOBAL_EVENT_ENABLE: 设置应用类日志是否开启Event日志,0:关闭Event日志,1:开启Event日志  
TASK_QUEUE_ENABLE: 用于控制开启task_queue算子下发队列优化的等级,0:关闭,1:开启Level 1优化,2:开启Level 2优化  
COMBINED_ENABLE: 设置combined标志。设置为0表示关闭此功能;设置为1表示开启,用于优化非连续两个算子组合类场景  
HCCL_WHITELIST_DISABLE: 配置在使用HCCL时是否开启通信白名单,0:开启白名单,1:关闭白名单  
CPU_AFFINITY_CONF: 控制CPU端算子任务的处理器亲和性,即设定任务绑核,设置0或未设置:表示不启用绑核功能, 1:表示开启粗粒度绑核, 2:表示开启细粒度绑核  
HCCL_CONNECT_TIMEOUT:  用于限制不同设备之间socket建链过程的超时等待时间,需要配置为整数,取值范围[120,7200],默认值为120,单位s  
ACLNN_CACHE_LIMIT: 配置单算子执行API在Host侧缓存的算子信息条目个数  
TOKENIZERS_PARALLELISM: 用于控制Hugging Face的transformers库中的分词器(tokenizer)在多线程环境下的行为  
PYTORCH_NPU_ALLOC_CONF: 控制缓存分配器行为  
OMP_NUM_THREADS: 设置执行期间使用的线程数

## 引用

### 公网地址说明

代码涉及公网地址参考 [公网地址](https://gitee.com/ascend/MindSpeed-MM/blob/master/docs/public_address_statement.md)