# Sozai — Architecture Sozai is a single-file React SPA (`index.html` / `landing.html`, compiled in the browser with Babel) backed by one FastAPI + Gradio process (`app.py`). The heavy models run on Hugging Face **ZeroGPU**, where CUDA may only be initialized inside a `@spaces.GPU` worker. This doc focuses on the **img2img "Develop"** path (FLUX.2-klein-4B + a watercolour scene LoRA), which is the most involved flow. --- ## 1. System architecture ```mermaid graph TD subgraph Browser["Browser — single-file SPA (React + Babel, no build step)"] UI["index.html / landing.html
DetailsForm · PhotoStack · DarkroomDeveloper
MapLibre 2D · Cesium 3D · DevelopScroll (scrapbook)"] end subgraph App["app.py — FastAPI + Gradio (one CPU process)"] API["HTTP / API layer
/api/config · /api/img2img · /api/nsfw
/api/transcribe · rooms (WebSocket + SSE)"] GW["Model gateways (lazy, cached)
autocaption · ASR · NSFW · img2img"] TRACE["Phoenix tracing (in-process)"] end subgraph GPU["ZeroGPU worker (ephemeral, per @spaces.GPU call)"] W["loads model from cache, runs on A10G/H200
CUDA inits ONLY here
pickle boundary: only plain data crosses in/out"] end Hub["HF Hub
FLUX.2-klein-4B · MiniCPM-V · NeMo ASR
Falconsai NSFW · scene LoRA (in repo)"] Ext["Browser-side services
Stadia watercolor tiles · Cesium ion · OSM/Nominatim geocode"] UI -->|fetch JSON / SSE / WS| API UI -->|/assets/* maps to public/| App UI -.->|tiles / geocode| Ext API --> GW GW -->|"@spaces.GPU"| W App -->|huggingface_hub download in MAIN process| Hub W -->|read from disk cache| Hub ``` **Invariant:** CUDA only ever initializes inside a `@spaces.GPU` worker. The main process stays CPU-only, which is why model **downloads happen in the main process** (it can write the cache) and the worker only **reads** them. --- ## 2. Develop (img2img) function flow Two lanes = two processes, and the **pickle boundary** between them is the key design constraint. See the box below if that term is new. > ### What "pickle boundary" means > > `@spaces.GPU` doesn't just call your function — it runs it in a **separate > process** (the GPU worker). Separate processes **don't share memory**: each has > its own private world of objects, and one can't reach into the other's. So to > run a worker function with arguments, the data has to be **copied across** — and > Python copies objects across processes by **pickling** them. > > **Pickling** is Python's word for *serialization*: flattening a live in-memory > object into a stream of bytes, so it can travel (to a file or another process) > and be **unpickled** — rebuilt — on the far side. Think *freeze-dry → ship → > rehydrate*. Every call across the boundary does this round trip: > > 1. main process **pickles** the arguments → bytes > 2. bytes are sent to the worker > 3. worker **unpickles** them → live objects, runs the function > 4. worker **pickles** the return value → bytes, sends back > 5. main **unpickles** the result > > The catch: **not everything can be pickled.** Plain data (numbers, strings, > lists, dicts, a `PIL.Image`, a numpy array) pickles fine. A loaded ML pipeline > does **not** — it holds CUDA tensors, open handles, and live modules that can't > be flattened, and at ~14 GB you'd never want to ship it every call anyway. > > That single fact shapes the design: > - the **photo** (`PIL.Image`) is small and picklable → it crosses *in* as the > argument to `_img2img_infer(src)`, and the developed PIL crosses back *out*. > - the **pipeline** can't (and shouldn't) cross → so it is **built inside the > worker** and kept in module state (`_img2img_state["pipe"]`), fetched locally > on each call rather than passed in. > > One line to remember: **across the worker boundary, send the *data*, never the > *model*.** ```mermaid sequenceDiagram autonumber participant B as Browser (HangingPrint) participant M as Main process (app.py, CPU) participant T as BG download thread (main) participant W as GPU worker (@spaces.GPU) participant H as HF Hub / disk cache Note over M: app start / first GET /api/config M->>M: _img2img_prep() — import torch+diffusers, resolve lora_path M->>T: _img2img_start_bg_download() T->>H: snapshot_download (xet OFF, retry+resume) ~24 GB H-->>T: weights written to disk cache T-->>M: _img2img_state.snapshot_ready = true Note over B: user opens Darkroom, each print mounts B->>M: POST /api/img2img {data_url, name} M->>M: _img2img_prep() (cached) — decode data_url to PIL src M->>W: _img2img_infer(src) (PIL pickled in) W->>W: _ensure_img2img_pipe() alt snapshot_ready == false W-->>M: raise "model still downloading" M-->>B: 500 (frontend keeps the original photo) else weights ready W->>H: Flux2KleinPipeline.from_pretrained(cache, local_files_only, accelerate) W->>W: peft load_lora_weights + set_adapters(["WTRCLR8"]) (isolated, falls back to base) W->>W: _img2img_square(src) to 1024x1024 W->>W: pipe(image=sq, steps=8, guidance) to PIL out W-->>M: PIL out (pickled back) M->>M: _pil_to_data_url(out) M-->>B: {available, image} B->>B: store developedSrc, shake-reveal swaps the element end ``` ### Function reference | # | Function | Process | In | Out | Relies on | |---|----------|---------|----|-----|-----------| | 1 | `_img2img_prep()` | main | env (LoRA path/repo) | `available` bool; caches `lora_path` | `torch`, `diffusers (Flux2KleinPipeline)` import | | 2 | `_img2img_start_bg_download()` | main (daemon thread) | repo id | writes ~24 GB to cache; sets `snapshot_ready` | `huggingface_hub.snapshot_download` (xet off, retry+resume) | | 3 | `img2img_rest(request)` | main | `{data_url, name}` | `{available, image}` | gate = `_img2img_prep()`; base64 decode | | 4 | `_img2img_infer(src)` | **GPU worker** | PIL `src` | PIL out | `@spaces.GPU`; calls 5/6/7 | | 5 | `_ensure_img2img_pipe()` | GPU worker | (module state) | cached `pipe` (+ `lora_loaded`) | `from_pretrained` + `accelerate`; `peft` for the LoRA; gated on `snapshot_ready` | | 6 | `_img2img_square(img)` | GPU worker | PIL | 1024² RGB PIL | EXIF transpose + centre crop | | 7 | `pipe(...)` | GPU worker | squared PIL + prompt | PIL out | `torch` (8-step flow sampler), VAE, transformer | | 8 | `_pil_to_data_url(img)` | main | PIL | png data URL | base64 | --- ## 3. Once cached — the warm path After the one-time download (2) and first build (5), each function has an early "already done?" guard, so the call collapses to *decode → cached pipe → 8-step sample → encode*. ```mermaid flowchart TD A["POST /api/img2img"] --> B["decode data_url to PIL"] B --> C["_img2img_infer @spaces.GPU"] C --> D{"pipe cached
in worker?"} D -->|"yes (WARM)"| E["return cached pipe"] D -->|"no (COLD)"| F{"weights on disk?"} F -->|yes| G["from_pretrained + peft LoRA
(rebuild, NO download)"] F -->|no| H["wait on bg download (~19 min)"] E --> I["_img2img_square to 1024"] G --> I I --> J["pipe(image, 8 steps) to PIL"] J --> K["_pil_to_data_url to {image}"] ``` ### Cache layers and their lifetimes | Layer | Where | Set by | Survives | |-------|-------|--------|----------| | 1. weights on disk | `~/.cache/huggingface` (~24 GB) | step 2 download | container life (lost on rebuild/restart — **ephemeral**) | | 2. `snapshot_ready` | `_img2img_state` flag (RAM) | step 2 download | process life | | 3. built pipe | `_img2img_state["pipe"]` + ~14 GB GPU memory | step 5 build | the GPU worker's warm window | ### What makes it cold again | Event | layer 1 | layer 2 | layer 3 | next call | |-------|---------|---------|---------|-----------| | another warm request | kept | kept | kept | ~12 s (pure inference) | | GPU idle timeout (ZeroGPU) | kept | kept | dropped | ~30 s (rebuild pipe from disk, no download) | | Space restart / sleep | wiped* | reset | reset | full re-download (~19 min) | | git push / factory rebuild | wiped | reset | reset | full re-download (~19 min) | \* the ephemeral disk is usually wiped on restart, sending you back to step 2. **Persistent storage** (a `/data` volume) would keep layer 1 across restarts, so future rebuilds become "build + warm" instead of "build + 19-min download." ### The ZeroGPU nuance `@spaces.GPU` runs the function in a forked GPU worker, and the GPU is attached only for the call's duration. `_img2img_state["pipe"]` is set *inside the worker*, so whether it survives to the next call depends on ZeroGPU keeping that worker (and its GPU memory) warm. Empirically it does, for a window — hence ~12 s repeat calls — and after an idle stretch ZeroGPU releases the GPU, the pipe is gone, and the next call rebuilds (but does **not** re-download, since layer 1 on disk is intact). The download only repeats when the **container** is replaced. --- ## 4. FLUX.2-klein dependency stack | Package | Role | |---------|------| | `diffusers` (git main) | provides `Flux2KleinPipeline` (loader + denoising loop); klein is too new for a release | | `torch` | GPU compute — the flow-matching sampler | | `transformers` | the text-encoder sub-models diffusers pulls in | | `accelerate` | `low_cpu_mem_usage` meta-device loading (avoids OOM on the ~24 GB) | | `safetensors` | weight format (model shards **and** the LoRA) | | `peft` | injects LoRA adapters — `load_lora_weights()` raises without it | | `huggingface_hub` | downloads + resolves the cache | | `hf_xet` | HF chunked download backend — **disabled** here (crawled / failed writes in the worker) | | `spaces` | the `@spaces.GPU` decorator | The base model makes a generic watercolour from a prompt; the 92 MB LoRA is a set of low-rank weight deltas trained on the target style, and `peft` is what loads them so the `WTRCLR8` trigger token activates the trained look.