Instructions to use Cuttle-fish-my/Rethinking-CFG-OPD-ckpts with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Cuttle-fish-my/Rethinking-CFG-OPD-ckpts with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
File size: 3,775 Bytes
4091d9d 5a35dd0 4091d9d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 | ---
license: apache-2.0
base_model: stabilityai/stable-diffusion-3.5-medium
library_name: peft
tags:
- lora
- on-policy-distillation
---
# Text-rendering distillation — PDM student (λ = 2)
The **Positive-Direction Matching** student from the text-rendering half of our study of
classifier-free guidance in on-policy distillation: a single OCR teacher distilled into a
student along the student's own denoising trajectories, on SD3.5-Medium.
```
pdm_lambda2.safetensors 71.7 MB LoRA over a frozen stabilityai/stable-diffusion-3.5-medium
adapter_config.json r=32, alpha=64, attention q/k/v + output projections
```
## What this setting is for
It is the **benign half** of the paper's "when does NBA emerge?" comparison, and the result is
that PDM and the naive objective perform the same.
Teacher and student here share one base model, one LoRA shape and one guidance scale, and the
teacher's negative branch is the plain unconditional branch — nothing the student cannot
represent (`c⁻_T = c⁻_S = ∅`). Under that shared negative conditioning both branch errors fall
together for every objective, so naive's under-identification never gets travelled. The
contrasting privileged-conditioning setting is where the two objectives separate.
## Results
OCR reward (%, ↑) on the 1 018-prompt held-out set, 40 denoising steps, trained at γ = 4.5:
| Method | γ=1 | γ=2 | γ=3 | γ=4 | γ=4.5 |
|---|---|---|---|---|---|
| Teacher | 75.24 | 92.91 | 94.16 | 94.43 | 94.09 |
| OPD (Naive) | 73.87 | 92.72 | 94.52 | 94.43 | 94.11 |
| **OPD (PDM, this)** | **74.48** | **93.73** | **93.99** | **94.02** | **94.38** |
Both students track the teacher across the sweep, including the drop at γ = 1 — that drop is
inherited from the teacher rather than introduced by distillation.
Reward is `R_OCR = 1 − min(d, |s|) / |s|`, where `s` is the target string, `ŝ` is what
PaddleOCR reads back, both lowercased with spaces stripped, and `d = 0` if `s` occurs inside `ŝ`
and `Lev(ŝ, s)` otherwise. Note the substring short-circuit: rendering the target inside otherwise
unrelated text still scores 1.0, so this is a legibility score rather than a plain edit distance.
## Training
| | |
|---|---|
| objective | PDM, `train.opd_loss = "pdm"`, `train.pdm_lambda = 2.0` |
| guidance | γ = 4.5, student and teacher |
| sampler | flow (ODE), `noise_level = 0` ⇒ mean matching; 10 steps train / 40 eval |
| partial trajectory | `timestep_fraction = 0.99` ⇒ K = 9 of 10 transitions supervised |
| optimiser | AdamW, constant lr 3e-4, grad clip 1.0, EMA decay 0.9 every 8 steps |
| batch | 8 GPUs × 3 prompts × 3 accumulated ⇒ 72 trajectories per update |
| steps | 1000 updates, seed 42 |
| teacher | upstream's released OCR teacher, [`quanhaol/DiffusionOPD`](https://huggingface.co/quanhaol/DiffusionOPD) |
The teacher is **not** included here — fetch it from upstream. (The PEFT checkpoint this was
extracted from carried a copy of it; we verified it is byte-identical to the released one and
dropped it.)
## Use
The code lives in the text-rendering repo, a fork of
[`ali-vilab/DiffusionOPD`](https://github.com/ali-vilab/DiffusionOPD):
```bash
torchrun --nproc_per_node=8 scripts/evaluation.py \
--checkpoint_path <dir containing lora/> \
--model_type sd3 --dataset ocr \
--guidance_scale 4.5 --num_inference_steps 40 --mixed_precision fp16 \
--output_dir evaluation_output/pdm_lambda2
```
`evaluation.py` expects the adapter at `<checkpoint_path>/lora/adapter_model.safetensors`, so
place `pdm_lambda2.safetensors` and `adapter_config.json` in a `lora/` subdirectory, renaming the
weights to `adapter_model.safetensors`.
## Licence
Apache-2.0. SD3.5-Medium and the upstream teacher carry their own terms.
|