# Open-weight image models ≤32B for the Weiner 4-koma pipeline Research date: 2026-06-12. Method: deep-research workflow (5 search angles, 22 sources, 108 claims extracted, 18 adversarially verified 3-0). Claims marked *(unverified)* had their verification cut off by a rate limit — they come from official/primary sources but were not independently confirmed. ## Use-case requirements Generate a 2×2-grid 4-panel comic in ONE image (claymation/3D-render style, recurring sausage-sensei mascot). Speech-bubble text composited programmatically afterward (in-model text = bonus only). Script comes from a separate LLM → prompt adherence for multi-panel scene layout is the critical capability. Character consistency within the grid (free, by construction) and across generations (LoRA and/or image-reference conditioning). Must run on HF ZeroGPU and be license-clean for a public hackathon Space. ## Infrastructure correction (verified 3-0) ZeroGPU is no longer H200/70GB: it now runs **RTX Pro 6000 Blackwell** — `large` (default) = half GPU / **48GB VRAM** at 1× quota, `xlarge` = full GPU / 96GB at 2× quota. Max `@spaces.GPU(duration=...)` remains effectively ~120s per call. Source: https://huggingface.co/docs/hub/spaces-zerogpu ## Quality bar - Nano Banana 2 (Gemini 3.1 Flash Image Preview): Elo 1258, #4 on Artificial Analysis arena (verified). - **HiDream-O1-Image-1.5: Elo 1264, #3 — ABOVE Nano Banana 2** (verified). Open-weight availability NOT confirmed (HiDream-I1 was MIT). **Watch item: if weights are open, test it.** ## Candidates ### Qwen-Image-Edit-2509 / 2511 (~20B MMDiT + Qwen2.5-VL encoder) - License: **Apache 2.0** (verified). - **Multi-image reference editing** (person+person, person+product, person+scene) — condition generation on mascot reference images; character lock WITHOUT training (verified). - 2509 explicitly improved identity/character preservation across pose/style changes (verified). - 2511 (unverified, official Qwen blog): further character-consistency gains; bakes popular community LoRAs into the base model. - Known best-in-class in-model text rendering (bonus: maybe legible bubbles for free). - 20B bf16 is tight-but-fine on 48GB with standard offload/quantization; official ZeroGPU Spaces exist; diffusers support mature. ### FLUX.2-klein 4B (+ 9B) - 4B: **Apache 2.0** (verified). 9B: FLUX Non-Commercial License (verified) → avoid for public Space. - ~13GB bf16; distilled 4-step variant = seconds per image (verified). - **LoRA in ~1hr on a single RTX 4090 (<24GB) via ostris/ai-toolkit**, diffusers-loadable (verified). - Unified t2i + editing + multi-reference per BFL blog/card *(unverified — votes lost to rate limit)*. - We already have klein infra (tiny-army sidecar Space + klein starter repo with LoRA tooling). ### Z-Image / Z-Image-Turbo (Tongyi, 6B) - **Apache 2.0** (verified). 6B family (verified). - Turbo: 8 DiT forward passes, sub-second on H800, fits 16GB consumer VRAM (verified). - Already integrated in tiny-army's local image path. - No verified reference/editing variant or comic-layout evidence; distilled few-step models typically trade away complex-prompt adherence → risk for 4-panel layouts. ### FLUX.2-dev (~32B + Mistral Small 3.1 24B encoder) - At/over the param cap; bf16 >80GB, ~62GB offloaded, ~20GB at 4-bit (verified). - Native multi-reference up to 10 images (verified) — best built-in character conditioning. - FLUX dev line is non-commercial-licensed → **ruled out** for a public hackathon Space. ### SD 3.5 / HiDream-I1 / CogView4 - Superseded on quality/adherence by the above within the same size envelope; no verified advantage for this use case. Not shortlisted. ## Ranking for this use case 1. **Qwen-Image-Edit-2511 (or 2509) — the comic engine.** Apache 2.0, the only verified ≤32B model with multi-image character-reference conditioning (mascot consistency with zero training: feed a Weiner reference sheet each call), strongest prompt adherence reputation for complex multi-panel scenes, and text rendering good enough that in-model bubbles might work. Cost: heaviest of the shortlist; needs the 48GB slice + offload and ~tens of seconds per comic (within 120s). 2. **FLUX.2-klein 4B — the on-theme Build Small pick.** Apache 2.0, trivially fast (4-step distilled), cheap LoRA pipeline already in-house → Weiner LoRA locks the mascot. Lower ceiling on per-panel scene complexity than 20B Qwen; the fine-tune narrows the style/character gap. 3. **Z-Image-Turbo 6B — speed fallback.** Apache 2.0, near-instant, already integrated; use if quota pressure dominates and layouts prove simple enough. ## Recommended architecture - **A/B test first** (same 10 scripts through Qwen-Edit-2511-with-reference-sheet and klein-4B-with/without-LoRA; blank-bubble prompting + programmatic text composite). - Production default: Qwen-Image-Edit-2511 + Weiner reference sheet (no training needed). - Featured small-model tier: klein 4B + Weiner LoRA (trained on a reference corpus we can bootstrap with the bigger model) — the "fine-tuned the 4B open model on our mascot" story. - Watch item: HiDream-O1-Image-1.5 weights — if open, benchmark immediately.