File size: 10,151 Bytes
478cb8f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
# FLUX Redux + ControlNet Style-Composition β€” Final Report & Replication Guide

**Date:** August 2026 Β· **Pod:** RunPod H100 80GB, ComfyUI 0.27.0 (`plx1029/comfyui-qwen` image)
**Goal:** 1 style-reference image + 1 composition-reference image (+ optional text prompt) β†’
style-transferred, composition-locked output on FLUX.1-dev. Maximum quality.

Everything here also lives on HF `aleph65/ComfyUI` (workflows in `workflows/`, models in
`models/`, this archive in `redux_experiment_final/`).

---

## TL;DR β€” the final production stack

```
FLUX.1-dev bf16 (UNETLoader, weight_dtype=default)
β”œβ”€ Style img  β†’ sigclip_vision_384 β†’ CLIPVisionEncode(center) β†’ StyleModelApply
β”‚               (flux1-redux-dev, strength 0.35, strength_type=MULTIPLY β€” never attn_bias, see Β§Findings)
β”œβ”€ Prompt     β†’ handwritten per-run: target's content + style ref's aesthetic β†’ FluxGuidance 3.0
β”œβ”€ Comp img   β†’ ImageResize+ 896x1152 fill/crop β†’ DepthAnythingV2(vitl) β†’ Union-Pro-2.0
β”‚               via ControlNetApplySD3 (strength 0.7, start 0, end 0.8) β€” DEPTH ONLY, single CN
└─ KSampler euler/simple/32 steps/CFG 1.0/denoise 1.0, seed 777 for all test runs
   └─ optional HIRES: LatentUpscaleBy bislerp 1.5x β†’ 2nd pass denoise 0.30 (Redux kept, CN released)
```

A/B variant (`bfl-lora`): same graph but structure via `flux1-depth-dev-lora` **0.85** +
`InstructPixToPixConditioning`, **FluxGuidance 10.0** (BFL spec β€” the LoRA is distilled at
guidance 10). Competitive with the CN path since the v4 fixes; less tunable (no end_percent,
no stacking).

## The 7 workflow files (in `workflows/`)

| file | purpose |
|---|---|
| `flux-v2-redux-style-composition.json` | **The main one.** Style+composition, all fixes baked. Bypassed groups: canny stack, ReduxAdvanced alt, turbo-LoRA fast preview, HIRES |
| `flux-v2-redux-style-composition-bfl-lora.json` | BFL depth-LoRA A/B variant, spec settings |
| `flux-redux-fal-dev.json` | fal `flux/dev/redux` endpoint replica (768Γ—1024, empty prompt, strength 1.0 multiply, euler/simple/28, guidance 3.5). Doubles as simple 1-image workflow |
| `flux-redux-prompt.json` | 1 image + prompt with authority (attn_bias 0.5 β€” fine here, NO ControlNet in this graph) |
| `flux-redux-schnell.json` | Replicate `flux-redux-schnell` replica (4 steps, no guidance, no prompt; 1024Β² = their 1:1 default, set 896Γ—1152 for portrait) |
| `flux-redux-style-composition.json` (v1) | Historical β€” superseded by flux-v2 |
| `flux-redux-style-composition-bfl-lora.json` (v1) | Historical β€” superseded |

`workflows/wip/` = every draft revision; `workflows/original-old-workflows/` = the pre-project
XLabs/Union-v1 workflows this started from (retired: XLabs is legacy, Union v1 needs
`SetUnionControlNetType` which v2.0 must NOT have).

## Models (all mirrored on HF `aleph65/ComfyUI` under `models/`)

| file | dir | source | gated |
|---|---|---|---|
| `flux1-dev.safetensors` (bf16 23.8GB) | diffusion_models | black-forest-labs/FLUX.1-dev | yes |
| `flux1-schnell.safetensors` (bf16 23.8GB) | diffusion_models | black-forest-labs/FLUX.1-schnell | no |
| `flux1-redux-dev.safetensors` (129MB) | style_models | black-forest-labs/FLUX.1-Redux-dev | yes |
| `FLUX.1-dev-ControlNet-Union-Pro-2.0.safetensors` (4.28GB) | controlnet | Shakker-Labs (renamed from `diffusion_pytorch_model.safetensors`) | no |
| `flux1-depth-dev-lora.safetensors` (1.24GB) | loras | black-forest-labs/FLUX.1-Depth-dev-lora | yes |
| `t5xxl_fp16.safetensors` + `clip_l.safetensors` | text_encoders | Comfy-Org repack | no |
| `ae.safetensors` | vae | FLUX.1-dev repo | (yes) |
| `sigclip_vision_patch14_384.safetensors` | clip_vision | Comfy-Org/sigclip_vision_384 | no |
| `flux1-turbo-alpha.safetensors` | loras | alimama-creative (fast-preview group) | no |
| `depth_anything_v2_vitl.pth` | auto-downloads into `custom_nodes/comfyui_controlnet_aux/ckpts/` | | no |

**Custom node packs (3):** `comfyui_controlnet_aux` (Fannovel16 β€” DepthAnythingV2),
`ComfyUI_essentials` (cubiq β€” ImageResize+), `ComfyUI_AdvancedRefluxControl` (kaibioinfo β€”
ReduxAdvanced alt branch). Everything else is ComfyUI core.

## Replication on a fresh pod

1. Pod from `plx1029/comfyui-qwen` image (or any ComfyUI β‰₯0.3.8 β€” that's when native
   StyleModelApply strength landed).
2. `HUGGING_FACE_ACCESS_TOKEN` env var; accept licenses for FLUX.1-dev, FLUX.1-Redux-dev,
   FLUX.1-Depth-dev-lora on that account.
3. `./download_missing_models.sh flux-v2-redux-style-composition.json ...` β€” resolves the node
   packs from each node's `cnr_id`/`aux_id` and pulls all models from the mirror. (Or run
   `scripts/download_models.py` which pulls from original sources.)
4. Restart ComfyUI, load the workflow, drop images, go. Defaults are the calibrated ones.
5. Batch runs: `scripts/run_matrix_v4.py` queues via the API (`POST /prompt` with API-format
   graphs, poll `/history/<id>`). `scripts/prompts_v4.py` holds the handwritten prompt fragments.

## Findings (chronological, what actually mattered)

1. **Baseline worked but bodies deformed** (v1–v3): aesthetics great, frequent limb/torso
   deformities in both two-image paths. Single-image workflows always clean.
2. **ROOT CAUSE (code-verified, undocumented upstream): `attn_bias` + ControlNet is broken.**
   `StyleModelApply(attn_bias)` appends all 729 Redux tokens at FULL strength and stores the
   attenuation as an attention-bias mask. The base flux transformer honors the mask; the
   ControlNet branch never receives it (`ControlNetFlux.forward_orig` has no attn_mask param;
   not in its `extra_conds`). Base attends Redux @50%, CN attends @100% β†’ two conditioning
   streams steer different images every step β†’ ghosting/deformities. Also: any attention mask
   kicks flux off the flash-attention fast path.
   **Fix: `strength_type=multiply` at 0.25–0.35 whenever a ControlNet is present.** attn_bias
   remains fine in CN-free graphs. (8-run isolation diagnostic + two deep-research passes;
   see `comparison_sheets/_diag_sheet.jpg` β€” column d is the smoking gun.)
3. **BFL depth-LoRA wants its spec: guidance 10.0 + LoRA 0.85.** It is distilled at guidance 10;
   running 3–4 gives mush. Earlier sweeps that said "guidance 4 looks better" were contaminated
   by the attn_bias bug. The CN path conversely wants guidance ~3.0. **The two paths need
   different FluxGuidance values.**
4. **Prompts are the anatomy lever.** Flux resolves limb ambiguity from text. Empty prompt =
   worst; generic anatomy boilerplate = marginal; **handwritten per-run prompts describing the
   composition target's actual content fused with the style ref's actual aesthetic = best**.
   Depth maps are semantically ambiguous at limb crossings; the prompt disambiguates.
   Negative prompts are INERT at CFG 1.0 (comfy skips uncond entirely) β€” anatomy must be
   handled positively.
5. **Calibration verdicts** (30-run, `comparison_sheets/_C*.jpg`): multiply 0.35 > 0.25/0.30
   (style strength, all clean) > ReduxAdvanced ds3 (clean but weakens style); 896Γ—1152 fine
   despite Union-Pro-2.0's 512 training res (no penalty once the conditioning fight was fixed);
   CN end 0.4 loses composition, 0.6–0.8 both hold β†’ keep 0.8, use 0.6 for freedom; depth-only
   single CN (canny stack changes scenes β€” keep off unless contour lock wanted).
6. **Input preprocessing matters**: fill/crop everything to 896Γ—1152 in-graph (portrait 3:4
   bucket), depth map == latent to the pixel, zero padding (padding becomes literal image
   content in the latent-concat/bfl path). Save crops + depth maps for auditing.
7. **Input spec** (what to feed it): 3:4 portrait, 1152Γ—1536, sRGB 8-bit, no watermarks/text/
   borders, EXIF baked. Style refs: style-defining content in the vertical middle β€” sigclip
   center-crops to a square, top/bottom ~12% discarded. Composition refs: clear fg/bg depth
   separation.
8. **Hires two-pass** (1.5Γ— latent bislerp + denoise 0.30, Redux carried, CN released):
   validated win β€” real detail gain, aesthetic preserved. Ships bypassed.
9. **Known edge case, unsolved**: person-style Γ— empty-room-target (our rΓ—r4 combos) is a
   contradictory ask; results inconsistent even with interior prompts. Needs its own treatment.
10. **Redux mode compatibility notes**: scaled-fp8 flux checkpoints black-screen with Redux
    (ComfyUI #5849). XLabs CN stack is legacy/dormant. Union-Pro-**2.0** must NOT have
    `SetUnionControlNetType` (that's v1-only).

## Prompt-writing recipe (for new image pairs)

Write one sentence of **content** from the composition target (subject, pose, clothing,
setting β€” what the depth map encodes), one sentence of **style** from the style ref (lighting
type, palette, finish), then `"Sharp focus, natural proportions, coherent anatomy."`
If the target has no person, say so explicitly and describe the space instead.
Fragments for the 10 experiment inputs are in `scripts/prompts_v4.py`.

## What's in this archive

```
README.md                  ← this file
workflows/                 ← all 7 finals + wip/ history + original-old-workflows/
inputs/                    ← inputs_v1 (first test set) + inputs_v2 (curated 1152Γ—1536 3:4 set)
scripts/                   ← graph builders, all matrix/sweep/calibration runners, prompts, model downloader
research/                  ← FLUX-REDUX-CONTROLNET-RESEARCH.md (full research log w/ sources)
outputs_v4_final/          ← the final 84-run matrix: 36 style-comp (+hires/ +inputs/ audits),
                             36 bfl-lora, 12 singles, calib/ (30 calibration runs)
comparison_sheets/         ← every verdict sheet: _diag_sheet (the attn_bias smoking gun),
                             _A* (v2 sweeps), _C* (v4 calibration), _v2_vs_v3, _v3_vs_v4
claude_memory/             ← Claude Code memory files from this project
```

Run provenance: every test image used **seed 777**; matrices covered all 24 rΓ—t + 12 ordered
rΓ—r combos per two-image workflow. v1β†’v4 output sets remain in `/workspace/outputs*` on the pod
(not archived here except v4).

β€” Built with Claude Code (Fable 5), Aug 2026. It was fun, sir. 🫑