Diff-Instruct with Diffused Reward: Towards Principled One-step Generator RL
Paper • 2605.24001 • Published
How to use Junyi-W/DIDR with Diffusers:
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("Junyi-W/DIDR", dtype=torch.bfloat16, device_map="cuda")
prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]Diff-Instruct with Diffused Reward: Towards Principled One-step Generator RL (NeurIPS 2026), Junyi Wu, Weijian Luo, Haoyang Zheng, Ruizhe Zhang, Guang Lin
Feel free to contact us if you have any questions about the paper!
Junyi Wu wu2393@purdue.edu
We can use the standard diffusers pipeline:
import torch
from diffusers import DiffusionPipeline, UNet2DConditionModel, LCMScheduler
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
base_model_id = "stabilityai/stable-diffusion-xl-base-1.0"
repo_name = "Junyi-W/DIDR"
ckpt_name = "didr_sdxl_1step_unet_fp16.safetensors"
# Load model.
unet = UNet2DConditionModel.from_config(UNet2DConditionModel.load_config(base_model_id, subfolder="unet")).to("cuda", torch.float16)
unet.load_state_dict(load_file(hf_hub_download(repo_name, ckpt_name), device="cuda"))
pipe = DiffusionPipeline.from_pretrained(base_model_id, unet=unet, torch_dtype=torch.float16, variant="fp16").to("cuda")
pipe.scheduler = LCMScheduler.from_config(pipe.scheduler.config)
prompt="a photo of a cat"
image=pipe(prompt=prompt, num_inference_steps=1, guidance_scale=0, timesteps=[399]).images[0]
import torch
from diffusers import DiffusionPipeline, UNet2DConditionModel, LCMScheduler
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
base_model_id = "stabilityai/stable-diffusion-xl-base-1.0"
repo_name = "Junyi-W/DIDR"
ckpt_name = "didr_longer_sdxl_1step_unet_fp16.safetensors"
# Load model.
unet = UNet2DConditionModel.from_config(UNet2DConditionModel.load_config(base_model_id, subfolder="unet")).to("cuda", torch.float16)
unet.load_state_dict(load_file(hf_hub_download(repo_name, ckpt_name), device="cuda"))
pipe = DiffusionPipeline.from_pretrained(base_model_id, unet=unet, torch_dtype=torch.float16, variant="fp16").to("cuda")
pipe.scheduler = LCMScheduler.from_config(pipe.scheduler.config)
prompt="a photo of a cat"
image=pipe(prompt=prompt, num_inference_steps=1, guidance_scale=0, timesteps=[399]).images[0]
import torch
from diffusers import ZImagePipeline, ZImageTransformer2DModel
base_model_id = "Tongyi-MAI/Z-Image-Turbo"
repo_name = "Junyi-W/DIDR"
# Load model.
transformer = ZImageTransformer2DModel.from_pretrained(repo_name, subfolder="zimage_transformer", torch_dtype=torch.bfloat16)
pipe = ZImagePipeline.from_pretrained(base_model_id, transformer=transformer, torch_dtype=torch.bfloat16).to("cuda")
prompt="A photorealistic portrait of a young woman in traditional Chinese red Hanfu, intricate golden embroidery, dramatic lighting, ultra high definition"
image=pipe(prompt=prompt, height=1024, width=1024, num_inference_steps=1, guidance_scale=0.0).images[0]
For more information, please refer to the code repository
DIDR is released under Creative Commons Attribution-NonCommercial 4.0 International License.
If you find DIDR useful or relevant to your research, please kindly cite our paper:
@inproceedings{wu2026didr,
title={Diff-Instruct with Diffused Reward: Towards Principled One-step Generator RL},
author={Wu, Junyi and Luo, Weijian and Zheng, Haoyang and Zhang, Ruizhe and Lin, Guang},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2026}
}
Our SDXL generators are initialized from DMD2 and our Z-Image generators from Z-Image-Turbo. We thank the authors for releasing their models.
Base model
Tongyi-MAI/Z-Image-Turbo