aleph65's picture
Archive flux-redux experiment: final writeup, workflows, inputs, scripts, research, v4 outputs, comparison sheets, Claude memories
478cb8f verified
|
Raw
History Blame Contribute Delete
10.2 kB
# FLUX Redux + ControlNet Style-Composition β€” Final Report & Replication Guide
**Date:** August 2026 Β· **Pod:** RunPod H100 80GB, ComfyUI 0.27.0 (`plx1029/comfyui-qwen` image)
**Goal:** 1 style-reference image + 1 composition-reference image (+ optional text prompt) β†’
style-transferred, composition-locked output on FLUX.1-dev. Maximum quality.
Everything here also lives on HF `aleph65/ComfyUI` (workflows in `workflows/`, models in
`models/`, this archive in `redux_experiment_final/`).
---
## TL;DR β€” the final production stack
```
FLUX.1-dev bf16 (UNETLoader, weight_dtype=default)
β”œβ”€ Style img β†’ sigclip_vision_384 β†’ CLIPVisionEncode(center) β†’ StyleModelApply
β”‚ (flux1-redux-dev, strength 0.35, strength_type=MULTIPLY β€” never attn_bias, see Β§Findings)
β”œβ”€ Prompt β†’ handwritten per-run: target's content + style ref's aesthetic β†’ FluxGuidance 3.0
β”œβ”€ Comp img β†’ ImageResize+ 896x1152 fill/crop β†’ DepthAnythingV2(vitl) β†’ Union-Pro-2.0
β”‚ via ControlNetApplySD3 (strength 0.7, start 0, end 0.8) β€” DEPTH ONLY, single CN
└─ KSampler euler/simple/32 steps/CFG 1.0/denoise 1.0, seed 777 for all test runs
└─ optional HIRES: LatentUpscaleBy bislerp 1.5x β†’ 2nd pass denoise 0.30 (Redux kept, CN released)
```
A/B variant (`bfl-lora`): same graph but structure via `flux1-depth-dev-lora` **0.85** +
`InstructPixToPixConditioning`, **FluxGuidance 10.0** (BFL spec β€” the LoRA is distilled at
guidance 10). Competitive with the CN path since the v4 fixes; less tunable (no end_percent,
no stacking).
## The 7 workflow files (in `workflows/`)
| file | purpose |
|---|---|
| `flux-v2-redux-style-composition.json` | **The main one.** Style+composition, all fixes baked. Bypassed groups: canny stack, ReduxAdvanced alt, turbo-LoRA fast preview, HIRES |
| `flux-v2-redux-style-composition-bfl-lora.json` | BFL depth-LoRA A/B variant, spec settings |
| `flux-redux-fal-dev.json` | fal `flux/dev/redux` endpoint replica (768Γ—1024, empty prompt, strength 1.0 multiply, euler/simple/28, guidance 3.5). Doubles as simple 1-image workflow |
| `flux-redux-prompt.json` | 1 image + prompt with authority (attn_bias 0.5 β€” fine here, NO ControlNet in this graph) |
| `flux-redux-schnell.json` | Replicate `flux-redux-schnell` replica (4 steps, no guidance, no prompt; 1024Β² = their 1:1 default, set 896Γ—1152 for portrait) |
| `flux-redux-style-composition.json` (v1) | Historical β€” superseded by flux-v2 |
| `flux-redux-style-composition-bfl-lora.json` (v1) | Historical β€” superseded |
`workflows/wip/` = every draft revision; `workflows/original-old-workflows/` = the pre-project
XLabs/Union-v1 workflows this started from (retired: XLabs is legacy, Union v1 needs
`SetUnionControlNetType` which v2.0 must NOT have).
## Models (all mirrored on HF `aleph65/ComfyUI` under `models/`)
| file | dir | source | gated |
|---|---|---|---|
| `flux1-dev.safetensors` (bf16 23.8GB) | diffusion_models | black-forest-labs/FLUX.1-dev | yes |
| `flux1-schnell.safetensors` (bf16 23.8GB) | diffusion_models | black-forest-labs/FLUX.1-schnell | no |
| `flux1-redux-dev.safetensors` (129MB) | style_models | black-forest-labs/FLUX.1-Redux-dev | yes |
| `FLUX.1-dev-ControlNet-Union-Pro-2.0.safetensors` (4.28GB) | controlnet | Shakker-Labs (renamed from `diffusion_pytorch_model.safetensors`) | no |
| `flux1-depth-dev-lora.safetensors` (1.24GB) | loras | black-forest-labs/FLUX.1-Depth-dev-lora | yes |
| `t5xxl_fp16.safetensors` + `clip_l.safetensors` | text_encoders | Comfy-Org repack | no |
| `ae.safetensors` | vae | FLUX.1-dev repo | (yes) |
| `sigclip_vision_patch14_384.safetensors` | clip_vision | Comfy-Org/sigclip_vision_384 | no |
| `flux1-turbo-alpha.safetensors` | loras | alimama-creative (fast-preview group) | no |
| `depth_anything_v2_vitl.pth` | auto-downloads into `custom_nodes/comfyui_controlnet_aux/ckpts/` | | no |
**Custom node packs (3):** `comfyui_controlnet_aux` (Fannovel16 β€” DepthAnythingV2),
`ComfyUI_essentials` (cubiq β€” ImageResize+), `ComfyUI_AdvancedRefluxControl` (kaibioinfo β€”
ReduxAdvanced alt branch). Everything else is ComfyUI core.
## Replication on a fresh pod
1. Pod from `plx1029/comfyui-qwen` image (or any ComfyUI β‰₯0.3.8 β€” that's when native
StyleModelApply strength landed).
2. `HUGGING_FACE_ACCESS_TOKEN` env var; accept licenses for FLUX.1-dev, FLUX.1-Redux-dev,
FLUX.1-Depth-dev-lora on that account.
3. `./download_missing_models.sh flux-v2-redux-style-composition.json ...` β€” resolves the node
packs from each node's `cnr_id`/`aux_id` and pulls all models from the mirror. (Or run
`scripts/download_models.py` which pulls from original sources.)
4. Restart ComfyUI, load the workflow, drop images, go. Defaults are the calibrated ones.
5. Batch runs: `scripts/run_matrix_v4.py` queues via the API (`POST /prompt` with API-format
graphs, poll `/history/<id>`). `scripts/prompts_v4.py` holds the handwritten prompt fragments.
## Findings (chronological, what actually mattered)
1. **Baseline worked but bodies deformed** (v1–v3): aesthetics great, frequent limb/torso
deformities in both two-image paths. Single-image workflows always clean.
2. **ROOT CAUSE (code-verified, undocumented upstream): `attn_bias` + ControlNet is broken.**
`StyleModelApply(attn_bias)` appends all 729 Redux tokens at FULL strength and stores the
attenuation as an attention-bias mask. The base flux transformer honors the mask; the
ControlNet branch never receives it (`ControlNetFlux.forward_orig` has no attn_mask param;
not in its `extra_conds`). Base attends Redux @50%, CN attends @100% β†’ two conditioning
streams steer different images every step β†’ ghosting/deformities. Also: any attention mask
kicks flux off the flash-attention fast path.
**Fix: `strength_type=multiply` at 0.25–0.35 whenever a ControlNet is present.** attn_bias
remains fine in CN-free graphs. (8-run isolation diagnostic + two deep-research passes;
see `comparison_sheets/_diag_sheet.jpg` β€” column d is the smoking gun.)
3. **BFL depth-LoRA wants its spec: guidance 10.0 + LoRA 0.85.** It is distilled at guidance 10;
running 3–4 gives mush. Earlier sweeps that said "guidance 4 looks better" were contaminated
by the attn_bias bug. The CN path conversely wants guidance ~3.0. **The two paths need
different FluxGuidance values.**
4. **Prompts are the anatomy lever.** Flux resolves limb ambiguity from text. Empty prompt =
worst; generic anatomy boilerplate = marginal; **handwritten per-run prompts describing the
composition target's actual content fused with the style ref's actual aesthetic = best**.
Depth maps are semantically ambiguous at limb crossings; the prompt disambiguates.
Negative prompts are INERT at CFG 1.0 (comfy skips uncond entirely) β€” anatomy must be
handled positively.
5. **Calibration verdicts** (30-run, `comparison_sheets/_C*.jpg`): multiply 0.35 > 0.25/0.30
(style strength, all clean) > ReduxAdvanced ds3 (clean but weakens style); 896Γ—1152 fine
despite Union-Pro-2.0's 512 training res (no penalty once the conditioning fight was fixed);
CN end 0.4 loses composition, 0.6–0.8 both hold β†’ keep 0.8, use 0.6 for freedom; depth-only
single CN (canny stack changes scenes β€” keep off unless contour lock wanted).
6. **Input preprocessing matters**: fill/crop everything to 896Γ—1152 in-graph (portrait 3:4
bucket), depth map == latent to the pixel, zero padding (padding becomes literal image
content in the latent-concat/bfl path). Save crops + depth maps for auditing.
7. **Input spec** (what to feed it): 3:4 portrait, 1152Γ—1536, sRGB 8-bit, no watermarks/text/
borders, EXIF baked. Style refs: style-defining content in the vertical middle β€” sigclip
center-crops to a square, top/bottom ~12% discarded. Composition refs: clear fg/bg depth
separation.
8. **Hires two-pass** (1.5Γ— latent bislerp + denoise 0.30, Redux carried, CN released):
validated win β€” real detail gain, aesthetic preserved. Ships bypassed.
9. **Known edge case, unsolved**: person-style Γ— empty-room-target (our rΓ—r4 combos) is a
contradictory ask; results inconsistent even with interior prompts. Needs its own treatment.
10. **Redux mode compatibility notes**: scaled-fp8 flux checkpoints black-screen with Redux
(ComfyUI #5849). XLabs CN stack is legacy/dormant. Union-Pro-**2.0** must NOT have
`SetUnionControlNetType` (that's v1-only).
## Prompt-writing recipe (for new image pairs)
Write one sentence of **content** from the composition target (subject, pose, clothing,
setting β€” what the depth map encodes), one sentence of **style** from the style ref (lighting
type, palette, finish), then `"Sharp focus, natural proportions, coherent anatomy."`
If the target has no person, say so explicitly and describe the space instead.
Fragments for the 10 experiment inputs are in `scripts/prompts_v4.py`.
## What's in this archive
```
README.md ← this file
workflows/ ← all 7 finals + wip/ history + original-old-workflows/
inputs/ ← inputs_v1 (first test set) + inputs_v2 (curated 1152Γ—1536 3:4 set)
scripts/ ← graph builders, all matrix/sweep/calibration runners, prompts, model downloader
research/ ← FLUX-REDUX-CONTROLNET-RESEARCH.md (full research log w/ sources)
outputs_v4_final/ ← the final 84-run matrix: 36 style-comp (+hires/ +inputs/ audits),
36 bfl-lora, 12 singles, calib/ (30 calibration runs)
comparison_sheets/ ← every verdict sheet: _diag_sheet (the attn_bias smoking gun),
_A* (v2 sweeps), _C* (v4 calibration), _v2_vs_v3, _v3_vs_v4
claude_memory/ ← Claude Code memory files from this project
```
Run provenance: every test image used **seed 777**; matrices covered all 24 rΓ—t + 12 ordered
rΓ—r combos per two-image workflow. v1β†’v4 output sets remain in `/workspace/outputs*` on the pod
(not archived here except v4).
β€” Built with Claude Code (Fable 5), Aug 2026. It was fun, sir. 🫑