File size: 9,904 Bytes
be2de62
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
# Sozai β€” Architecture

Sozai is a single-file React SPA (`index.html` / `landing.html`, compiled in the
browser with Babel) backed by one FastAPI + Gradio process (`app.py`). The heavy
models run on Hugging Face **ZeroGPU**, where CUDA may only be initialized inside
a `@spaces.GPU` worker. This doc focuses on the **img2img "Develop"** path
(FLUX.2-klein-4B + a watercolour scene LoRA), which is the most involved flow.

---

## 1. System architecture

```mermaid
graph TD
  subgraph Browser["Browser β€” single-file SPA (React + Babel, no build step)"]
    UI["index.html / landing.html<br/>DetailsForm Β· PhotoStack Β· DarkroomDeveloper<br/>MapLibre 2D Β· Cesium 3D Β· DevelopScroll (scrapbook)"]
  end

  subgraph App["app.py β€” FastAPI + Gradio (one CPU process)"]
    API["HTTP / API layer<br/>/api/config Β· /api/img2img Β· /api/nsfw<br/>/api/transcribe Β· rooms (WebSocket + SSE)"]
    GW["Model gateways (lazy, cached)<br/>autocaption Β· ASR Β· NSFW Β· img2img"]
    TRACE["Phoenix tracing (in-process)"]
  end

  subgraph GPU["ZeroGPU worker (ephemeral, per @spaces.GPU call)"]
    W["loads model from cache, runs on A10G/H200<br/>CUDA inits ONLY here<br/>pickle boundary: only plain data crosses in/out"]
  end

  Hub["HF Hub<br/>FLUX.2-klein-4B Β· MiniCPM-V Β· NeMo ASR<br/>Falconsai NSFW Β· scene LoRA (in repo)"]
  Ext["Browser-side services<br/>Stadia watercolor tiles Β· Cesium ion Β· OSM/Nominatim geocode"]

  UI -->|fetch JSON / SSE / WS| API
  UI -->|/assets/* maps to public/| App
  UI -.->|tiles / geocode| Ext
  API --> GW
  GW -->|"@spaces.GPU"| W
  App -->|huggingface_hub download in MAIN process| Hub
  W -->|read from disk cache| Hub
```

**Invariant:** CUDA only ever initializes inside a `@spaces.GPU` worker. The main
process stays CPU-only, which is why model **downloads happen in the main
process** (it can write the cache) and the worker only **reads** them.

---

## 2. Develop (img2img) function flow

Two lanes = two processes, and the **pickle boundary** between them is the key
design constraint. See the box below if that term is new.

> ### What "pickle boundary" means
>
> `@spaces.GPU` doesn't just call your function β€” it runs it in a **separate
> process** (the GPU worker). Separate processes **don't share memory**: each has
> its own private world of objects, and one can't reach into the other's. So to
> run a worker function with arguments, the data has to be **copied across** β€” and
> Python copies objects across processes by **pickling** them.
>
> **Pickling** is Python's word for *serialization*: flattening a live in-memory
> object into a stream of bytes, so it can travel (to a file or another process)
> and be **unpickled** β€” rebuilt β€” on the far side. Think *freeze-dry β†’ ship β†’
> rehydrate*. Every call across the boundary does this round trip:
>
> 1. main process **pickles** the arguments β†’ bytes
> 2. bytes are sent to the worker
> 3. worker **unpickles** them β†’ live objects, runs the function
> 4. worker **pickles** the return value β†’ bytes, sends back
> 5. main **unpickles** the result
>
> The catch: **not everything can be pickled.** Plain data (numbers, strings,
> lists, dicts, a `PIL.Image`, a numpy array) pickles fine. A loaded ML pipeline
> does **not** β€” it holds CUDA tensors, open handles, and live modules that can't
> be flattened, and at ~14 GB you'd never want to ship it every call anyway.
>
> That single fact shapes the design:
> - the **photo** (`PIL.Image`) is small and picklable β†’ it crosses *in* as the
>   argument to `_img2img_infer(src)`, and the developed PIL crosses back *out*.
> - the **pipeline** can't (and shouldn't) cross β†’ so it is **built inside the
>   worker** and kept in module state (`_img2img_state["pipe"]`), fetched locally
>   on each call rather than passed in.
>
> One line to remember: **across the worker boundary, send the *data*, never the
> *model*.**

```mermaid
sequenceDiagram
  autonumber
  participant B as Browser (HangingPrint)
  participant M as Main process (app.py, CPU)
  participant T as BG download thread (main)
  participant W as GPU worker (@spaces.GPU)
  participant H as HF Hub / disk cache

  Note over M: app start / first GET /api/config
  M->>M: _img2img_prep() β€” import torch+diffusers, resolve lora_path
  M->>T: _img2img_start_bg_download()
  T->>H: snapshot_download (xet OFF, retry+resume) ~24 GB
  H-->>T: weights written to disk cache
  T-->>M: _img2img_state.snapshot_ready = true

  Note over B: user opens Darkroom, each print mounts
  B->>M: POST /api/img2img {data_url, name}
  M->>M: _img2img_prep() (cached) β€” decode data_url to PIL src
  M->>W: _img2img_infer(src)  (PIL pickled in)
  W->>W: _ensure_img2img_pipe()
  alt snapshot_ready == false
    W-->>M: raise "model still downloading"
    M-->>B: 500 (frontend keeps the original photo)
  else weights ready
    W->>H: Flux2KleinPipeline.from_pretrained(cache, local_files_only, accelerate)
    W->>W: peft load_lora_weights + set_adapters(["WTRCLR8"])  (isolated, falls back to base)
    W->>W: _img2img_square(src) to 1024x1024
    W->>W: pipe(image=sq, steps=8, guidance) to PIL out
    W-->>M: PIL out (pickled back)
    M->>M: _pil_to_data_url(out)
    M-->>B: {available, image}
    B->>B: store developedSrc, shake-reveal swaps the <img> element
  end
```

### Function reference

| # | Function | Process | In | Out | Relies on |
|---|----------|---------|----|-----|-----------|
| 1 | `_img2img_prep()` | main | env (LoRA path/repo) | `available` bool; caches `lora_path` | `torch`, `diffusers (Flux2KleinPipeline)` import |
| 2 | `_img2img_start_bg_download()` | main (daemon thread) | repo id | writes ~24 GB to cache; sets `snapshot_ready` | `huggingface_hub.snapshot_download` (xet off, retry+resume) |
| 3 | `img2img_rest(request)` | main | `{data_url, name}` | `{available, image}` | gate = `_img2img_prep()`; base64 decode |
| 4 | `_img2img_infer(src)` | **GPU worker** | PIL `src` | PIL out | `@spaces.GPU`; calls 5/6/7 |
| 5 | `_ensure_img2img_pipe()` | GPU worker | (module state) | cached `pipe` (+ `lora_loaded`) | `from_pretrained` + `accelerate`; `peft` for the LoRA; gated on `snapshot_ready` |
| 6 | `_img2img_square(img)` | GPU worker | PIL | 1024Β² RGB PIL | EXIF transpose + centre crop |
| 7 | `pipe(...)` | GPU worker | squared PIL + prompt | PIL out | `torch` (8-step flow sampler), VAE, transformer |
| 8 | `_pil_to_data_url(img)` | main | PIL | png data URL | base64 |

---

## 3. Once cached β€” the warm path

After the one-time download (2) and first build (5), each function has an
early "already done?" guard, so the call collapses to *decode β†’ cached pipe β†’
8-step sample β†’ encode*.

```mermaid
flowchart TD
  A["POST /api/img2img"] --> B["decode data_url to PIL"]
  B --> C["_img2img_infer @spaces.GPU"]
  C --> D{"pipe cached<br/>in worker?"}
  D -->|"yes (WARM)"| E["return cached pipe"]
  D -->|"no (COLD)"| F{"weights on disk?"}
  F -->|yes| G["from_pretrained + peft LoRA<br/>(rebuild, NO download)"]
  F -->|no| H["wait on bg download (~19 min)"]
  E --> I["_img2img_square to 1024"]
  G --> I
  I --> J["pipe(image, 8 steps) to PIL"]
  J --> K["_pil_to_data_url to {image}"]
```

### Cache layers and their lifetimes

| Layer | Where | Set by | Survives |
|-------|-------|--------|----------|
| 1. weights on disk | `~/.cache/huggingface` (~24 GB) | step 2 download | container life (lost on rebuild/restart β€” **ephemeral**) |
| 2. `snapshot_ready` | `_img2img_state` flag (RAM) | step 2 download | process life |
| 3. built pipe | `_img2img_state["pipe"]` + ~14 GB GPU memory | step 5 build | the GPU worker's warm window |

### What makes it cold again

| Event | layer 1 | layer 2 | layer 3 | next call |
|-------|---------|---------|---------|-----------|
| another warm request | kept | kept | kept | ~12 s (pure inference) |
| GPU idle timeout (ZeroGPU) | kept | kept | dropped | ~30 s (rebuild pipe from disk, no download) |
| Space restart / sleep | wiped* | reset | reset | full re-download (~19 min) |
| git push / factory rebuild | wiped | reset | reset | full re-download (~19 min) |

\* the ephemeral disk is usually wiped on restart, sending you back to step 2.
**Persistent storage** (a `/data` volume) would keep layer 1 across restarts, so
future rebuilds become "build + warm" instead of "build + 19-min download."

### The ZeroGPU nuance

`@spaces.GPU` runs the function in a forked GPU worker, and the GPU is attached
only for the call's duration. `_img2img_state["pipe"]` is set *inside the worker*,
so whether it survives to the next call depends on ZeroGPU keeping that worker
(and its GPU memory) warm. Empirically it does, for a window β€” hence ~12 s repeat
calls β€” and after an idle stretch ZeroGPU releases the GPU, the pipe is gone, and
the next call rebuilds (but does **not** re-download, since layer 1 on disk is
intact). The download only repeats when the **container** is replaced.

---

## 4. FLUX.2-klein dependency stack

| Package | Role |
|---------|------|
| `diffusers` (git main) | provides `Flux2KleinPipeline` (loader + denoising loop); klein is too new for a release |
| `torch` | GPU compute β€” the flow-matching sampler |
| `transformers` | the text-encoder sub-models diffusers pulls in |
| `accelerate` | `low_cpu_mem_usage` meta-device loading (avoids OOM on the ~24 GB) |
| `safetensors` | weight format (model shards **and** the LoRA) |
| `peft` | injects LoRA adapters β€” `load_lora_weights()` raises without it |
| `huggingface_hub` | downloads + resolves the cache |
| `hf_xet` | HF chunked download backend β€” **disabled** here (crawled / failed writes in the worker) |
| `spaces` | the `@spaces.GPU` decorator |

The base model makes a generic watercolour from a prompt; the 92 MB LoRA is a set
of low-rank weight deltas trained on the target style, and `peft` is what loads
them so the `WTRCLR8` trigger token activates the trained look.