MEND-SD3.5M-CLIPScore

Official MEND: RL for Flow Models via Proximal Velocity Matching evaluation adapter, trained on CLIPScore for 100 updates. This is the paper's checkpoint-100, including when the training run continued beyond that update.

Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik.

Paper · Code · Collection · Project and blog · HF paper

MEND accepts proposed reward-gradient moves when their capped reward gain pays for a quadratic displacement price, then fits the resulting velocity targets. The unchanged sample is always a candidate. Training uses an EMA behavior policy, with no KL term or frozen reference model in the loss.

Generate with Diffusers

Install the MEND core package using the installation guide, or install Diffusers, PEFT, Transformers, Accelerate and a suitable CUDA PyTorch build in a uv environment. Accept the base model terms where required and authenticate with hf auth login. The adapter is public; base weights are downloaded separately.

import torch
from diffusers import StableDiffusion3Pipeline

pipe = StableDiffusion3Pipeline.from_pretrained(
    "stabilityai/stable-diffusion-3.5-medium", torch_dtype=torch.bfloat16,
)
pipe.load_lora_weights("shreshthsaini/MEND-SD3.5M-CLIPScore", weight_name="pytorch_lora_weights.safetensors")
pipe.enable_model_cpu_offload()
image = pipe(
    "a small blue book on a large red book",
    height=512, width=512, num_inference_steps=40, guidance_scale=1.0,
    generator=torch.Generator(device="cpu").manual_seed(0),
).images[0]
image.save("mend.png")

CPU offload trades speed for GPU memory. Direct Diffusers generation can produce different pixels from the paper's evaluation sampler. Use MEND's inference and evaluation guides to reproduce that protocol. No minimum VRAM is claimed.

Original PEFT adapter

python scripts/download_weights.py mend_clipscore downloads the original adapter without loading a model. Official aliases pin a Hub commit and check SHA256 hashes.

python scripts/generate.py --lora mend_clipscore \
  --prompt "a small blue book on a large red book" --seeds 0 \
  --model stabilityai/stable-diffusion-3.5-medium --guidance_scale 1.0 \
  --num_steps 40 --resolution 512 --batch_size 1 --out_dir outputs/mend_clipscore

Checkpoint and measurements

Setting Value
Base stabilityai/stable-diffusion-3.5-medium
Base revision used by the study b940f670f0eda2d07fbb75229e779da1ad11eb80
Training reward / updates CLIPScore / 100
Adapter Evaluation EMA policy, rank 32, original alpha 64
Sampling 512 pixels, 40 steps, guidance 1.0
Evaluation DrawBench, 200 prompts, 1000 images

The original study evaluation measured:

PickScore HPSv2.1 HPSv3 ImageReward CLIPScore Aesthetic
22.2658 0.2681 1.1494 0.9795 0.3143 5.2231

These are archived study measurements, rather than a fresh release evaluation. Reward gains vary by prompt and evaluator. SD3 Medium seed-1 and seed-2 releases are distinct training seeds. See the paper for aggregate comparisons and protocols.

Files and validation

adapter_config.json and adapter_model.safetensors preserve the original evaluation adapter. pytorch_lora_weights.safetensors provides its equivalent Diffusers format by absorbing the alpha/rank scale of 2 into LoRA B. Use one format at a time. Training optimizer states, the auxiliary behavior adapter and base weights are excluded.

release_manifest.json records sampling settings, original evaluation hashes and conversion validation. SHA256SUMS lists file hashes. All tensors are finite and every shape matches the full base transformer. Original PEFT and direct Diffusers loading passed on the meta architecture. All 191 actual adapter CPU forwards were bitwise equal after conversion; save/load preserved every tensor. Full GPU image generation and metric evaluation were not repeated for release.

License and citation

Powered by Stability AI. The SD3.5 Medium adapter retains the Stability AI Community License. A copy is included in LICENSE.md, with attribution in NOTICE. MEND code is Apache-2.0. Base models, reward models and datasets retain their own terms.

If you use MEND or these weights, please cite:

@misc{saini2026mend,
  title         = {{MEND}: {RL} For Flow Models via Proximal Velocity Matching},
  author        = {Saini, Shreshth and Birkbeck, Neil and Wang, Yilin and Adsumilli, Balu and Bovik, Alan C.},
  year          = {2026},
  eprint        = {2610.05954},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  doi           = {10.48550/arXiv.2610.05954},
  url           = {https://arxiv.org/abs/2610.05954}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for shreshthsaini/MEND-SD3.5M-CLIPScore

Adapter
(113)
this model

Space using shreshthsaini/MEND-SD3.5M-CLIPScore 1

Collection including shreshthsaini/MEND-SD3.5M-CLIPScore

Paper for shreshthsaini/MEND-SD3.5M-CLIPScore