| ## Training Code | |
| We can choose whether to use deepspeed or fsdp in z_image, which can save a lot of video memory | |
| . | |
| The metadata_control.json is a little different from normal json in Z-Image, you need to add a control_file_path, and [DWPose](https://github.com/IDEA-Research/DWPose) is suggested as tool to generate control file. | |
| ```json | |
| [ | |
| { | |
| "file_path": "train/00000002.jpg", | |
| "control_file_path": "control/00000002.jpg", | |
| "text": "A group of young men in suits and sunglasses are walking down a city street.", | |
| "type": "image" | |
| }, | |
| ..... | |
| ] | |
| ``` | |
| Some parameters in the sh file can be confusing, and they are explained in this document: | |
| - `enable_bucket` is used to enable bucket training. When enabled, the model does not crop the images at the center, but instead, it trains the entire images after grouping them into buckets based on resolution. | |
| - `random_hw_adapt` is used to enable automatic height and width scaling for images. When `random_hw_adapt` is enabled, the training images will have their height and width set to `image_sample_size` as the maximum and `512` as the minimum. | |
| - For example, when `random_hw_adapt` is enabled, `image_sample_size=1024`, the resolution of image inputs for training is `512x512` to `1024x1024` | |
| - `resume_from_checkpoint` is used to set the training should be resumed from a previous checkpoint. Use a path or `"latest"` to automatically select the last available checkpoint. | |
| When train model with multi machines, please set the params as follows: | |
| ```sh | |
| export MASTER_ADDR="your master address" | |
| export MASTER_PORT=10086 | |
| export WORLD_SIZE=1 # The number of machines | |
| export NUM_PROCESS=8 # The number of processes, such as WORLD_SIZE * 8 | |
| export RANK=0 # The rank of this machine | |
| accelerate launch --mixed_precision="bf16" --main_process_ip=$MASTER_ADDR --main_process_port=$MASTER_PORT --num_machines=$WORLD_SIZE --num_processes=$NUM_PROCESS --machine_rank=$RANK scripts/xxx/xxx.py | |
| ``` | |
| Without deepspeed: | |
| Training z_image without DeepSpeed may result in insufficient GPU memory. | |
| ```sh | |
| export MODEL_NAME="models/Diffusion_Transformer/Z-Image-Turbo" | |
| export DATASET_NAME="datasets/internal_datasets/" | |
| export DATASET_META_NAME="datasets/internal_datasets/metadata.json" | |
| # NCCL_IB_DISABLE=1 and NCCL_P2P_DISABLE=1 are used in multi nodes without RDMA. | |
| # export NCCL_IB_DISABLE=1 | |
| # export NCCL_P2P_DISABLE=1 | |
| NCCL_DEBUG=INFO | |
| accelerate launch --mixed_precision="bf16" scripts/z_image_fun/train_control.py \ | |
| --config_path="config/z_image/z_image_control.yaml" \ | |
| --pretrained_model_name_or_path=$MODEL_NAME \ | |
| --train_data_dir=$DATASET_NAME \ | |
| --train_data_meta=$DATASET_META_NAME \ | |
| --train_batch_size=1 \ | |
| --image_sample_size=1328 \ | |
| --gradient_accumulation_steps=1 \ | |
| --dataloader_num_workers=8 \ | |
| --num_train_epochs=100 \ | |
| --checkpointing_steps=50 \ | |
| --learning_rate=2e-05 \ | |
| --lr_scheduler="constant_with_warmup" \ | |
| --lr_warmup_steps=100 \ | |
| --seed=42 \ | |
| --output_dir="output_dir_z_image_control" \ | |
| --gradient_checkpointing \ | |
| --mixed_precision="bf16" \ | |
| --adam_weight_decay=3e-2 \ | |
| --adam_epsilon=1e-10 \ | |
| --vae_mini_batch=1 \ | |
| --max_grad_norm=0.05 \ | |
| --enable_bucket \ | |
| --uniform_sampling \ | |
| --transformer_path="models/Personalized_Model/Z-Image-Turbo-Fun-Controlnet-Union.safetensors" \ | |
| --trainable_modules "control" | |
| ``` | |
| With Deepspeed Zero-2: | |
| ```sh | |
| export MODEL_NAME="models/Diffusion_Transformer/Z-Image-Turbo" | |
| export DATASET_NAME="datasets/internal_datasets/" | |
| export DATASET_META_NAME="datasets/internal_datasets/metadata.json" | |
| # NCCL_IB_DISABLE=1 and NCCL_P2P_DISABLE=1 are used in multi nodes without RDMA. | |
| # export NCCL_IB_DISABLE=1 | |
| # export NCCL_P2P_DISABLE=1 | |
| NCCL_DEBUG=INFO | |
| accelerate launch --use_deepspeed --deepspeed_config_file config/zero_stage2_config.json --deepspeed_multinode_launcher standard scripts/z_image_fun/train_control.py \ | |
| --config_path="config/z_image/z_image_control.yaml" \ | |
| --pretrained_model_name_or_path=$MODEL_NAME \ | |
| --train_data_dir=$DATASET_NAME \ | |
| --train_data_meta=$DATASET_META_NAME \ | |
| --train_batch_size=1 \ | |
| --image_sample_size=1328 \ | |
| --gradient_accumulation_steps=1 \ | |
| --dataloader_num_workers=8 \ | |
| --num_train_epochs=100 \ | |
| --checkpointing_steps=50 \ | |
| --learning_rate=2e-05 \ | |
| --lr_scheduler="constant_with_warmup" \ | |
| --lr_warmup_steps=100 \ | |
| --seed=42 \ | |
| --output_dir="output_dir_z_image_control" \ | |
| --gradient_checkpointing \ | |
| --mixed_precision="bf16" \ | |
| --adam_weight_decay=3e-2 \ | |
| --adam_epsilon=1e-10 \ | |
| --vae_mini_batch=1 \ | |
| --max_grad_norm=0.05 \ | |
| --enable_bucket \ | |
| --uniform_sampling \ | |
| --transformer_path="models/Personalized_Model/Z-Image-Turbo-Fun-Controlnet-Union.safetensors" \ | |
| --trainable_modules "control" | |
| ``` | |
| With FSDP: | |
| ```sh | |
| export MODEL_NAME="models/Diffusion_Transformer/Z-Image-Turbo" | |
| export DATASET_NAME="datasets/internal_datasets/" | |
| export DATASET_META_NAME="datasets/internal_datasets/metadata.json" | |
| # NCCL_IB_DISABLE=1 and NCCL_P2P_DISABLE=1 are used in multi nodes without RDMA. | |
| # export NCCL_IB_DISABLE=1 | |
| # export NCCL_P2P_DISABLE=1 | |
| NCCL_DEBUG=INFO | |
| accelerate launch --mixed_precision="bf16" --use_fsdp --fsdp_auto_wrap_policy TRANSFORMER_BASED_WRAP --fsdp_transformer_layer_cls_to_wrap BaseZImageTransformerBlock,ZImageControlTransformerBlock --fsdp_sharding_strategy "FULL_SHARD" --fsdp_state_dict_type=SHARDED_STATE_DICT --fsdp_backward_prefetch "BACKWARD_PRE" --fsdp_cpu_ram_efficient_loading False scripts/z_image_fun/train_control.py \ | |
| --config_path="config/z_image/z_image_control.yaml" \ | |
| --pretrained_model_name_or_path=$MODEL_NAME \ | |
| --train_data_dir=$DATASET_NAME \ | |
| --train_data_meta=$DATASET_META_NAME \ | |
| --train_batch_size=1 \ | |
| --image_sample_size=1328 \ | |
| --gradient_accumulation_steps=1 \ | |
| --dataloader_num_workers=8 \ | |
| --num_train_epochs=100 \ | |
| --checkpointing_steps=50 \ | |
| --learning_rate=2e-05 \ | |
| --lr_scheduler="constant_with_warmup" \ | |
| --lr_warmup_steps=100 \ | |
| --seed=42 \ | |
| --output_dir="output_dir_z_image_control" \ | |
| --gradient_checkpointing \ | |
| --mixed_precision="bf16" \ | |
| --adam_weight_decay=3e-2 \ | |
| --adam_epsilon=1e-10 \ | |
| --vae_mini_batch=1 \ | |
| --max_grad_norm=0.05 \ | |
| --enable_bucket \ | |
| --uniform_sampling \ | |
| --transformer_path="models/Personalized_Model/Z-Image-Turbo-Fun-Controlnet-Union.safetensors" \ | |
| --trainable_modules "control" | |
| ``` |