# Sozai — Architecture
Sozai is a single-file React SPA (`index.html` / `landing.html`, compiled in the
browser with Babel) backed by one FastAPI + Gradio process (`app.py`). The heavy
models run on Hugging Face **ZeroGPU**, where CUDA may only be initialized inside
a `@spaces.GPU` worker. This doc focuses on the **img2img "Develop"** path
(FLUX.2-klein-4B + a watercolour scene LoRA), which is the most involved flow.
---
## 1. System architecture
```mermaid
graph TD
subgraph Browser["Browser — single-file SPA (React + Babel, no build step)"]
UI["index.html / landing.html
DetailsForm · PhotoStack · DarkroomDeveloper
MapLibre 2D · Cesium 3D · DevelopScroll (scrapbook)"]
end
subgraph App["app.py — FastAPI + Gradio (one CPU process)"]
API["HTTP / API layer
/api/config · /api/img2img · /api/nsfw
/api/transcribe · rooms (WebSocket + SSE)"]
GW["Model gateways (lazy, cached)
autocaption · ASR · NSFW · img2img"]
TRACE["Phoenix tracing (in-process)"]
end
subgraph GPU["ZeroGPU worker (ephemeral, per @spaces.GPU call)"]
W["loads model from cache, runs on A10G/H200
CUDA inits ONLY here
pickle boundary: only plain data crosses in/out"]
end
Hub["HF Hub
FLUX.2-klein-4B · MiniCPM-V · NeMo ASR
Falconsai NSFW · scene LoRA (in repo)"]
Ext["Browser-side services
Stadia watercolor tiles · Cesium ion · OSM/Nominatim geocode"]
UI -->|fetch JSON / SSE / WS| API
UI -->|/assets/* maps to public/| App
UI -.->|tiles / geocode| Ext
API --> GW
GW -->|"@spaces.GPU"| W
App -->|huggingface_hub download in MAIN process| Hub
W -->|read from disk cache| Hub
```
**Invariant:** CUDA only ever initializes inside a `@spaces.GPU` worker. The main
process stays CPU-only, which is why model **downloads happen in the main
process** (it can write the cache) and the worker only **reads** them.
---
## 2. Develop (img2img) function flow
Two lanes = two processes, and the **pickle boundary** between them is the key
design constraint. See the box below if that term is new.
> ### What "pickle boundary" means
>
> `@spaces.GPU` doesn't just call your function — it runs it in a **separate
> process** (the GPU worker). Separate processes **don't share memory**: each has
> its own private world of objects, and one can't reach into the other's. So to
> run a worker function with arguments, the data has to be **copied across** — and
> Python copies objects across processes by **pickling** them.
>
> **Pickling** is Python's word for *serialization*: flattening a live in-memory
> object into a stream of bytes, so it can travel (to a file or another process)
> and be **unpickled** — rebuilt — on the far side. Think *freeze-dry → ship →
> rehydrate*. Every call across the boundary does this round trip:
>
> 1. main process **pickles** the arguments → bytes
> 2. bytes are sent to the worker
> 3. worker **unpickles** them → live objects, runs the function
> 4. worker **pickles** the return value → bytes, sends back
> 5. main **unpickles** the result
>
> The catch: **not everything can be pickled.** Plain data (numbers, strings,
> lists, dicts, a `PIL.Image`, a numpy array) pickles fine. A loaded ML pipeline
> does **not** — it holds CUDA tensors, open handles, and live modules that can't
> be flattened, and at ~14 GB you'd never want to ship it every call anyway.
>
> That single fact shapes the design:
> - the **photo** (`PIL.Image`) is small and picklable → it crosses *in* as the
> argument to `_img2img_infer(src)`, and the developed PIL crosses back *out*.
> - the **pipeline** can't (and shouldn't) cross → so it is **built inside the
> worker** and kept in module state (`_img2img_state["pipe"]`), fetched locally
> on each call rather than passed in.
>
> One line to remember: **across the worker boundary, send the *data*, never the
> *model*.**
```mermaid
sequenceDiagram
autonumber
participant B as Browser (HangingPrint)
participant M as Main process (app.py, CPU)
participant T as BG download thread (main)
participant W as GPU worker (@spaces.GPU)
participant H as HF Hub / disk cache
Note over M: app start / first GET /api/config
M->>M: _img2img_prep() — import torch+diffusers, resolve lora_path
M->>T: _img2img_start_bg_download()
T->>H: snapshot_download (xet OFF, retry+resume) ~24 GB
H-->>T: weights written to disk cache
T-->>M: _img2img_state.snapshot_ready = true
Note over B: user opens Darkroom, each print mounts
B->>M: POST /api/img2img {data_url, name}
M->>M: _img2img_prep() (cached) — decode data_url to PIL src
M->>W: _img2img_infer(src) (PIL pickled in)
W->>W: _ensure_img2img_pipe()
alt snapshot_ready == false
W-->>M: raise "model still downloading"
M-->>B: 500 (frontend keeps the original photo)
else weights ready
W->>H: Flux2KleinPipeline.from_pretrained(cache, local_files_only, accelerate)
W->>W: peft load_lora_weights + set_adapters(["WTRCLR8"]) (isolated, falls back to base)
W->>W: _img2img_square(src) to 1024x1024
W->>W: pipe(image=sq, steps=8, guidance) to PIL out
W-->>M: PIL out (pickled back)
M->>M: _pil_to_data_url(out)
M-->>B: {available, image}
B->>B: store developedSrc, shake-reveal swaps the
element
end
```
### Function reference
| # | Function | Process | In | Out | Relies on |
|---|----------|---------|----|-----|-----------|
| 1 | `_img2img_prep()` | main | env (LoRA path/repo) | `available` bool; caches `lora_path` | `torch`, `diffusers (Flux2KleinPipeline)` import |
| 2 | `_img2img_start_bg_download()` | main (daemon thread) | repo id | writes ~24 GB to cache; sets `snapshot_ready` | `huggingface_hub.snapshot_download` (xet off, retry+resume) |
| 3 | `img2img_rest(request)` | main | `{data_url, name}` | `{available, image}` | gate = `_img2img_prep()`; base64 decode |
| 4 | `_img2img_infer(src)` | **GPU worker** | PIL `src` | PIL out | `@spaces.GPU`; calls 5/6/7 |
| 5 | `_ensure_img2img_pipe()` | GPU worker | (module state) | cached `pipe` (+ `lora_loaded`) | `from_pretrained` + `accelerate`; `peft` for the LoRA; gated on `snapshot_ready` |
| 6 | `_img2img_square(img)` | GPU worker | PIL | 1024² RGB PIL | EXIF transpose + centre crop |
| 7 | `pipe(...)` | GPU worker | squared PIL + prompt | PIL out | `torch` (8-step flow sampler), VAE, transformer |
| 8 | `_pil_to_data_url(img)` | main | PIL | png data URL | base64 |
---
## 3. Once cached — the warm path
After the one-time download (2) and first build (5), each function has an
early "already done?" guard, so the call collapses to *decode → cached pipe →
8-step sample → encode*.
```mermaid
flowchart TD
A["POST /api/img2img"] --> B["decode data_url to PIL"]
B --> C["_img2img_infer @spaces.GPU"]
C --> D{"pipe cached
in worker?"}
D -->|"yes (WARM)"| E["return cached pipe"]
D -->|"no (COLD)"| F{"weights on disk?"}
F -->|yes| G["from_pretrained + peft LoRA
(rebuild, NO download)"]
F -->|no| H["wait on bg download (~19 min)"]
E --> I["_img2img_square to 1024"]
G --> I
I --> J["pipe(image, 8 steps) to PIL"]
J --> K["_pil_to_data_url to {image}"]
```
### Cache layers and their lifetimes
| Layer | Where | Set by | Survives |
|-------|-------|--------|----------|
| 1. weights on disk | `~/.cache/huggingface` (~24 GB) | step 2 download | container life (lost on rebuild/restart — **ephemeral**) |
| 2. `snapshot_ready` | `_img2img_state` flag (RAM) | step 2 download | process life |
| 3. built pipe | `_img2img_state["pipe"]` + ~14 GB GPU memory | step 5 build | the GPU worker's warm window |
### What makes it cold again
| Event | layer 1 | layer 2 | layer 3 | next call |
|-------|---------|---------|---------|-----------|
| another warm request | kept | kept | kept | ~12 s (pure inference) |
| GPU idle timeout (ZeroGPU) | kept | kept | dropped | ~30 s (rebuild pipe from disk, no download) |
| Space restart / sleep | wiped* | reset | reset | full re-download (~19 min) |
| git push / factory rebuild | wiped | reset | reset | full re-download (~19 min) |
\* the ephemeral disk is usually wiped on restart, sending you back to step 2.
**Persistent storage** (a `/data` volume) would keep layer 1 across restarts, so
future rebuilds become "build + warm" instead of "build + 19-min download."
### The ZeroGPU nuance
`@spaces.GPU` runs the function in a forked GPU worker, and the GPU is attached
only for the call's duration. `_img2img_state["pipe"]` is set *inside the worker*,
so whether it survives to the next call depends on ZeroGPU keeping that worker
(and its GPU memory) warm. Empirically it does, for a window — hence ~12 s repeat
calls — and after an idle stretch ZeroGPU releases the GPU, the pipe is gone, and
the next call rebuilds (but does **not** re-download, since layer 1 on disk is
intact). The download only repeats when the **container** is replaced.
---
## 4. FLUX.2-klein dependency stack
| Package | Role |
|---------|------|
| `diffusers` (git main) | provides `Flux2KleinPipeline` (loader + denoising loop); klein is too new for a release |
| `torch` | GPU compute — the flow-matching sampler |
| `transformers` | the text-encoder sub-models diffusers pulls in |
| `accelerate` | `low_cpu_mem_usage` meta-device loading (avoids OOM on the ~24 GB) |
| `safetensors` | weight format (model shards **and** the LoRA) |
| `peft` | injects LoRA adapters — `load_lora_weights()` raises without it |
| `huggingface_hub` | downloads + resolves the cache |
| `hf_xet` | HF chunked download backend — **disabled** here (crawled / failed writes in the worker) |
| `spaces` | the `@spaces.GPU` decorator |
The base model makes a generic watercolour from a prompt; the 92 MB LoRA is a set
of low-rank weight deltas trained on the target style, and `peft` is what loads
them so the `WTRCLR8` trigger token activates the trained look.