File size: 10,151 Bytes
478cb8f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 | # FLUX Redux + ControlNet Style-Composition β Final Report & Replication Guide
**Date:** August 2026 Β· **Pod:** RunPod H100 80GB, ComfyUI 0.27.0 (`plx1029/comfyui-qwen` image)
**Goal:** 1 style-reference image + 1 composition-reference image (+ optional text prompt) β
style-transferred, composition-locked output on FLUX.1-dev. Maximum quality.
Everything here also lives on HF `aleph65/ComfyUI` (workflows in `workflows/`, models in
`models/`, this archive in `redux_experiment_final/`).
---
## TL;DR β the final production stack
```
FLUX.1-dev bf16 (UNETLoader, weight_dtype=default)
ββ Style img β sigclip_vision_384 β CLIPVisionEncode(center) β StyleModelApply
β (flux1-redux-dev, strength 0.35, strength_type=MULTIPLY β never attn_bias, see Β§Findings)
ββ Prompt β handwritten per-run: target's content + style ref's aesthetic β FluxGuidance 3.0
ββ Comp img β ImageResize+ 896x1152 fill/crop β DepthAnythingV2(vitl) β Union-Pro-2.0
β via ControlNetApplySD3 (strength 0.7, start 0, end 0.8) β DEPTH ONLY, single CN
ββ KSampler euler/simple/32 steps/CFG 1.0/denoise 1.0, seed 777 for all test runs
ββ optional HIRES: LatentUpscaleBy bislerp 1.5x β 2nd pass denoise 0.30 (Redux kept, CN released)
```
A/B variant (`bfl-lora`): same graph but structure via `flux1-depth-dev-lora` **0.85** +
`InstructPixToPixConditioning`, **FluxGuidance 10.0** (BFL spec β the LoRA is distilled at
guidance 10). Competitive with the CN path since the v4 fixes; less tunable (no end_percent,
no stacking).
## The 7 workflow files (in `workflows/`)
| file | purpose |
|---|---|
| `flux-v2-redux-style-composition.json` | **The main one.** Style+composition, all fixes baked. Bypassed groups: canny stack, ReduxAdvanced alt, turbo-LoRA fast preview, HIRES |
| `flux-v2-redux-style-composition-bfl-lora.json` | BFL depth-LoRA A/B variant, spec settings |
| `flux-redux-fal-dev.json` | fal `flux/dev/redux` endpoint replica (768Γ1024, empty prompt, strength 1.0 multiply, euler/simple/28, guidance 3.5). Doubles as simple 1-image workflow |
| `flux-redux-prompt.json` | 1 image + prompt with authority (attn_bias 0.5 β fine here, NO ControlNet in this graph) |
| `flux-redux-schnell.json` | Replicate `flux-redux-schnell` replica (4 steps, no guidance, no prompt; 1024Β² = their 1:1 default, set 896Γ1152 for portrait) |
| `flux-redux-style-composition.json` (v1) | Historical β superseded by flux-v2 |
| `flux-redux-style-composition-bfl-lora.json` (v1) | Historical β superseded |
`workflows/wip/` = every draft revision; `workflows/original-old-workflows/` = the pre-project
XLabs/Union-v1 workflows this started from (retired: XLabs is legacy, Union v1 needs
`SetUnionControlNetType` which v2.0 must NOT have).
## Models (all mirrored on HF `aleph65/ComfyUI` under `models/`)
| file | dir | source | gated |
|---|---|---|---|
| `flux1-dev.safetensors` (bf16 23.8GB) | diffusion_models | black-forest-labs/FLUX.1-dev | yes |
| `flux1-schnell.safetensors` (bf16 23.8GB) | diffusion_models | black-forest-labs/FLUX.1-schnell | no |
| `flux1-redux-dev.safetensors` (129MB) | style_models | black-forest-labs/FLUX.1-Redux-dev | yes |
| `FLUX.1-dev-ControlNet-Union-Pro-2.0.safetensors` (4.28GB) | controlnet | Shakker-Labs (renamed from `diffusion_pytorch_model.safetensors`) | no |
| `flux1-depth-dev-lora.safetensors` (1.24GB) | loras | black-forest-labs/FLUX.1-Depth-dev-lora | yes |
| `t5xxl_fp16.safetensors` + `clip_l.safetensors` | text_encoders | Comfy-Org repack | no |
| `ae.safetensors` | vae | FLUX.1-dev repo | (yes) |
| `sigclip_vision_patch14_384.safetensors` | clip_vision | Comfy-Org/sigclip_vision_384 | no |
| `flux1-turbo-alpha.safetensors` | loras | alimama-creative (fast-preview group) | no |
| `depth_anything_v2_vitl.pth` | auto-downloads into `custom_nodes/comfyui_controlnet_aux/ckpts/` | | no |
**Custom node packs (3):** `comfyui_controlnet_aux` (Fannovel16 β DepthAnythingV2),
`ComfyUI_essentials` (cubiq β ImageResize+), `ComfyUI_AdvancedRefluxControl` (kaibioinfo β
ReduxAdvanced alt branch). Everything else is ComfyUI core.
## Replication on a fresh pod
1. Pod from `plx1029/comfyui-qwen` image (or any ComfyUI β₯0.3.8 β that's when native
StyleModelApply strength landed).
2. `HUGGING_FACE_ACCESS_TOKEN` env var; accept licenses for FLUX.1-dev, FLUX.1-Redux-dev,
FLUX.1-Depth-dev-lora on that account.
3. `./download_missing_models.sh flux-v2-redux-style-composition.json ...` β resolves the node
packs from each node's `cnr_id`/`aux_id` and pulls all models from the mirror. (Or run
`scripts/download_models.py` which pulls from original sources.)
4. Restart ComfyUI, load the workflow, drop images, go. Defaults are the calibrated ones.
5. Batch runs: `scripts/run_matrix_v4.py` queues via the API (`POST /prompt` with API-format
graphs, poll `/history/<id>`). `scripts/prompts_v4.py` holds the handwritten prompt fragments.
## Findings (chronological, what actually mattered)
1. **Baseline worked but bodies deformed** (v1βv3): aesthetics great, frequent limb/torso
deformities in both two-image paths. Single-image workflows always clean.
2. **ROOT CAUSE (code-verified, undocumented upstream): `attn_bias` + ControlNet is broken.**
`StyleModelApply(attn_bias)` appends all 729 Redux tokens at FULL strength and stores the
attenuation as an attention-bias mask. The base flux transformer honors the mask; the
ControlNet branch never receives it (`ControlNetFlux.forward_orig` has no attn_mask param;
not in its `extra_conds`). Base attends Redux @50%, CN attends @100% β two conditioning
streams steer different images every step β ghosting/deformities. Also: any attention mask
kicks flux off the flash-attention fast path.
**Fix: `strength_type=multiply` at 0.25β0.35 whenever a ControlNet is present.** attn_bias
remains fine in CN-free graphs. (8-run isolation diagnostic + two deep-research passes;
see `comparison_sheets/_diag_sheet.jpg` β column d is the smoking gun.)
3. **BFL depth-LoRA wants its spec: guidance 10.0 + LoRA 0.85.** It is distilled at guidance 10;
running 3β4 gives mush. Earlier sweeps that said "guidance 4 looks better" were contaminated
by the attn_bias bug. The CN path conversely wants guidance ~3.0. **The two paths need
different FluxGuidance values.**
4. **Prompts are the anatomy lever.** Flux resolves limb ambiguity from text. Empty prompt =
worst; generic anatomy boilerplate = marginal; **handwritten per-run prompts describing the
composition target's actual content fused with the style ref's actual aesthetic = best**.
Depth maps are semantically ambiguous at limb crossings; the prompt disambiguates.
Negative prompts are INERT at CFG 1.0 (comfy skips uncond entirely) β anatomy must be
handled positively.
5. **Calibration verdicts** (30-run, `comparison_sheets/_C*.jpg`): multiply 0.35 > 0.25/0.30
(style strength, all clean) > ReduxAdvanced ds3 (clean but weakens style); 896Γ1152 fine
despite Union-Pro-2.0's 512 training res (no penalty once the conditioning fight was fixed);
CN end 0.4 loses composition, 0.6β0.8 both hold β keep 0.8, use 0.6 for freedom; depth-only
single CN (canny stack changes scenes β keep off unless contour lock wanted).
6. **Input preprocessing matters**: fill/crop everything to 896Γ1152 in-graph (portrait 3:4
bucket), depth map == latent to the pixel, zero padding (padding becomes literal image
content in the latent-concat/bfl path). Save crops + depth maps for auditing.
7. **Input spec** (what to feed it): 3:4 portrait, 1152Γ1536, sRGB 8-bit, no watermarks/text/
borders, EXIF baked. Style refs: style-defining content in the vertical middle β sigclip
center-crops to a square, top/bottom ~12% discarded. Composition refs: clear fg/bg depth
separation.
8. **Hires two-pass** (1.5Γ latent bislerp + denoise 0.30, Redux carried, CN released):
validated win β real detail gain, aesthetic preserved. Ships bypassed.
9. **Known edge case, unsolved**: person-style Γ empty-room-target (our rΓr4 combos) is a
contradictory ask; results inconsistent even with interior prompts. Needs its own treatment.
10. **Redux mode compatibility notes**: scaled-fp8 flux checkpoints black-screen with Redux
(ComfyUI #5849). XLabs CN stack is legacy/dormant. Union-Pro-**2.0** must NOT have
`SetUnionControlNetType` (that's v1-only).
## Prompt-writing recipe (for new image pairs)
Write one sentence of **content** from the composition target (subject, pose, clothing,
setting β what the depth map encodes), one sentence of **style** from the style ref (lighting
type, palette, finish), then `"Sharp focus, natural proportions, coherent anatomy."`
If the target has no person, say so explicitly and describe the space instead.
Fragments for the 10 experiment inputs are in `scripts/prompts_v4.py`.
## What's in this archive
```
README.md β this file
workflows/ β all 7 finals + wip/ history + original-old-workflows/
inputs/ β inputs_v1 (first test set) + inputs_v2 (curated 1152Γ1536 3:4 set)
scripts/ β graph builders, all matrix/sweep/calibration runners, prompts, model downloader
research/ β FLUX-REDUX-CONTROLNET-RESEARCH.md (full research log w/ sources)
outputs_v4_final/ β the final 84-run matrix: 36 style-comp (+hires/ +inputs/ audits),
36 bfl-lora, 12 singles, calib/ (30 calibration runs)
comparison_sheets/ β every verdict sheet: _diag_sheet (the attn_bias smoking gun),
_A* (v2 sweeps), _C* (v4 calibration), _v2_vs_v3, _v3_vs_v4
claude_memory/ β Claude Code memory files from this project
```
Run provenance: every test image used **seed 777**; matrices covered all 24 rΓt + 12 ordered
rΓr combos per two-image workflow. v1βv4 output sets remain in `/workspace/outputs*` on the pod
(not archived here except v4).
β Built with Claude Code (Fable 5), Aug 2026. It was fun, sir. π«‘
|