# Building Sozai: one process, one page, and a watercolour darkroom
> A developer's tour of how Sozai actually works — from the moment a visitor
> hits the landing page to a shared, real-time watercolour scrapbook.
>
> This is the **flow-oriented** companion to [`architecture.md`](architecture.md).
> Where that doc drills into the FLUX.2 img2img pipeline and the ZeroGPU pickle
> boundary, this one walks the **whole website** end to end.
---
## What Sozai is
Sozai (素材, *"material / source"*) is a collaborative photo-journaling app. You
drop in your photos, it captions and titles them for you by *looking* at them,
"develops" them into watercolour prints in a darkroom animation, and lets you pin
them onto a 2D or 3D map of where they happened — alone, or live with friends in
a shared room.
The whole thing is deliberately small in moving parts:
- **One frontend**, shipped as two hand-written HTML files (`landing.html` +
`index.html`) that compile React in the browser with Babel. No build step, no
bundler, no `node_modules`.
- **One backend**, a single `app.py` (~3k lines) running **FastAPI + Gradio** in
one CPU process.
- **Heavy models on demand**, run on Hugging Face **ZeroGPU**, where CUDA may
only be touched inside an ephemeral `@spaces.GPU` worker.
That's the entire system. Everything below is how those three pieces talk.
---
## The 10,000-foot view
```mermaid
flowchart LR
subgraph Browser["Browser — single-file React SPA, no build step"]
SPA["landing.html · index.html
screens · MapLibre 2D · Cesium 3D · Collab over SSE"]
end
subgraph Server["app.py — FastAPI + Gradio, one CPU process"]
HTTP["routes / /app /room/ID
REST + Gradio API + SSE rooms"]
GW["lazy model gateways
nsfw · autocaption · ASR · img2img"]
DB["SQLite — rooms · participants · room_state"]
PH["Phoenix tracing, in-process"]
end
subgraph GPU["ZeroGPU worker — @spaces.GPU, ephemeral"]
W["CUDA inits ONLY here
MiniCPM-V · NeMo ASR · NSFW · FLUX.2-klein"]
end
Hub["HF Hub — weights cache"]
Ext["Browser-side: Stadia tiles · Cesium ion · OSM geocode"]
SPA -->|fetch / SSE / Gradio| HTTP
SPA -.->|tiles / geocode| Ext
HTTP --> GW
HTTP --> DB
GW -->|"@spaces.GPU"| W
GW --> PH
Server -->|download in MAIN process| Hub
W -->|read disk cache| Hub
```
The single invariant worth burning into memory: **CUDA only ever initializes
inside a `@spaces.GPU` worker.** The main process stays CPU-only, which is why
model *downloads* happen in the main process (it owns the cache) while the worker
only *reads* them. The rest of the design falls out of that one rule.
---
## Part 1 — Two HTML files, three routes
There is no client-side router framework. The server hands out exactly three HTML
responses:
| Route | Serves | Purpose |
|-------|--------|---------|
| `GET /` | `landing.html` | the splash: make a room / join by code / go solo |
| `GET /app` | `index.html` | the app, **solo** (no room) |
| `GET /room/{room_id}` | `index.html` | the app, **in a room** (same file) |
`/app` and `/room/{id}` return the *identical* document. The app decides whether
it's collaborative purely by reading `location.pathname` at runtime:
```js
function roomId() {
const m = location.pathname.match(/^\/room\/([^\/?#]+)/);
return m ? decodeURIComponent(m[1]) : null;
}
```
If there's a room id in the URL, collaboration turns on. If not, every collab
call quietly no-ops. One codebase, two modes, zero branching at the route level.
One backend detail worth calling out: before `index.html` is served, the server
**inlines MapLibre GL** directly into the page (`_inline_maplibre`), replacing a
`` marker with the cached library source. This sidesteps
every CDN/MIME/proxy issue that can break a map on Hugging Face Spaces; if the
library can't be cached, the marker is left alone and an in-page CDN loader takes
over.
---
## Part 2 — The screen state machine
Once `index.html` boots, the entire app is a `` that renders one ``,
and `` is a tiny state machine keyed off a single string, `screen`:
```mermaid
flowchart TD
start([Visitor]) --> L["Landing — landing.html
GET /"]
L -->|make a room| MR["POST /api/rooms
creates share code + room_id"]
L -->|join by code| JC["POST /api/join-by-code
resolves code to room_id"]
L -->|continue without a room| SOLO["GET /app"]
MR --> ROOM["GET /room/ID"]
JC --> ROOM
SOLO --> BOOT["index.html SPA boots
React + Babel in the browser"]
ROOM --> BOOT
BOOT --> CFG["GET /api/config
feature flags + tokens"]
BOOT --> CONNECT{in a /room/ URL?}
CONNECT -->|yes| RT["Collab.connect — SSE
RoomPersistence loads saved scrapbook"]
CONNECT -->|no| SOLOMODE["solo: collab no-ops, nothing persists"]
BOOT --> SCREENS
subgraph SCREENS["Root — single screen state machine"]
ADD["screen = add
AddToAlbum + DetailsForm
upload · EXIF · NSFW · auto caption/title/tags · mic"]
DARK["screen = darkroom
DarkroomDeveloper
baths animation then shake-to-reveal"]
MAP["screen = map
MapOfMyTime — scrapbook · MapLibre 2D · Cesium 3D"]
PETS["screen = pets
PetsScreen"]
TRACE["screen = trace
TraceScreen — embedded Phoenix"]
end
ADD -->|all photos added| DARK
DARK -->|reveal all / skip| MAP
ADD -.-> PETS
ADD -.-> TRACE
PETS -.-> ADD
TRACE -.-> ADD
MAP -->|export| PDF["printable scrapbook
browser print, no upload"]
```
The mainline journey is **add → darkroom → map**. `pets` (a customizable desktop
companion) and `trace` (an embedded Phoenix LLM-trace viewer) are side branches
reachable from the top bar. The whole thing is driven by:
```js
function Page({ screen, setScreen, ... }) {
if (screen === "add") return setScreen("darkroom")} ... />;
if (screen === "darkroom") return setScreen("map")} onSkip={() => setScreen("map")} ... />;
if (screen === "pets") return setScreen("add")} />;
if (screen === "trace") return setScreen("add")} />;
return ; // "map"
}
```
In a room, screen changes are **broadcast** — peers who have "Follow" enabled get
mirrored to the same screen (more on that in Part 6).
### A note on the source of truth
Photos live in two stores, and the split is deliberate:
- **`stackPhotos`** (React state in ``) — the working list while you're in
the **add** flow. Each photo carries its own `title/caption/date/time/lat/lon`,
so the form fields can never desync from the photo list.
- **`window.SozaiPhotos`** — a small global pub/sub store that the **map** reads
from (pins, transforms, etc.).
Crucially, `SozaiPhotos` is **never** persisted to `localStorage`. Solo mode is
intentionally ephemeral (refresh = empty), and room photos live on the *server*
(`/api/rooms/{id}/state`), pulled in on join. Keeping a local copy would leak old
or deleted photos into a room you later joined — so the in-memory store plus the
server-side room state are the only two homes a photo ever has.
---
## Part 3 — Getting photos in (the `add` screen)
`AddToAlbum` + `DetailsForm` is where the AI earns its keep. When you add a
photo, several things happen, all backed by the same lazy-model pattern:
1. **NSFW gate** — `POST /api/nsfw` runs `Falconsai/nsfw_image_detection`. If the
model or its deps are missing, the gate is **open** (score `0.0`) so a missing
optional dependency can never block an upload. Threshold is `0.80` by default.
2. **Auto caption / title / tags** — each "✨ Auto" button calls the Gradio API
endpoint `autocaption(kind=...)`. One function serves all three: `kind`
selects the instruction, reply directive, generation length, and trace-span
name. The model (`MiniCPM-V`, a vision-language model) **looks at the pixels**
— the filename and alt text are only used as a fallback hint when no image can
be resolved.
3. **Voice notes** — the mic button records audio and `POST /api/transcribe`
turns it into text via NVIDIA **NeMo** streaming ASR, so you can dictate a
caption instead of typing.
Every model call is wrapped in `traced(...)`, which opens a Phoenix span with the
model id, prompt, invocation params, and output — so the **trace** screen shows a
live timeline of every LLM call the app makes. On any failure, autocaption
returns a lightweight fallback derived from the alt text, so the button *always*
does something.
---
## Part 4 — The Darkroom (img2img watercolour develop)
When you've added your photos, **add → darkroom** drops you into `DarkroomDeveloper`,
which is the most theatrical part of the app. It has two phases:
- **`baths`** — a top-down view of developing trays; the prints "soak" while the
backend turns each photo into a watercolour.
- **`hang`** — the prints hang on a line, and you **shake each one** (drag it) to
reveal the developed watercolour, complete with water-ripple SVG filters and a
safelight flicker.
Under the hood, each print posts to `POST /api/img2img`, which runs
**FLUX.2-klein-4B** plus a watercolour-scene **LoRA** (trigger token `WTRCLR8`)
inside a `@spaces.GPU` worker. This is the flow the original
[`architecture.md`](architecture.md) covers in depth — including the **pickle
boundary** (send the *data*, never the *model*), the lazy ~24 GB background
download, the warm/cold cache layers, and why the first develop is slow but
repeats land in ~12 s. If you only read one section of that doc, read §2–3.
The frontend is forgiving here too: if the model is still downloading, the API
returns a 500 and the print simply keeps the original photo. The shake-to-reveal
just swaps the `
` once a developed version is available.
---
## Part 5 — The Map of My Time
**darkroom → map** lands on `MapOfMyTime`, which has three view variants the user
can switch between (and persists in `localStorage`):
| Variant | Component | Tech |
|---------|-----------|------|
| `scrapbook` | `DevelopScroll` | a scrolling, scrapbook-style layout |
| 2D map | `WatercolorMap` | **MapLibre GL** + Stadia watercolor raster tiles |
| 3D globe | `CesiumMap` | **Cesium ion** (only if a token is configured) |
Photos with EXIF GPS auto-pin; everything else sits in a draft **tray** you can
drag onto the map. Each pinned photo carries its own canvas transform (position,
rotation, scale), and the whole arrangement is part of the room state that gets
synced and saved. From here you can **export** the scrapbook — `exportScrapbook`
builds a clean, self-contained printable HTML document and opens the browser's
print dialog. No server round-trip, so it works offline and never uploads your
images.
Browser-side services (Stadia tiles, Cesium ion, OSM/Nominatim geocoding) are
called **directly from the browser** — they never round-trip through `app.py`.
`/api/config` hands the frontend the tokens (e.g. Cesium ion) it needs so they
can be rotated server-side in one place.
---
## Part 6 — Real-time collaboration
This is the part that makes "build memories *together*" literal. The transport is
deliberately humble: **Server-Sent Events inbound + plain POST outbound**, not
WebSockets. Why? Hugging Face Spaces' proxy drops custom WebSocket upgrades, so
the app uses the same SSE+POST channel Gradio itself uses on Spaces. EventSource
also auto-reconnects on its own, so there's no manual reconnect loop.
```mermaid
sequenceDiagram
autonumber
participant A as Guest A — host
participant B as Guest B — joiner
participant S as app.py — Manager
participant DB as SQLite
A->>S: POST /api/rooms
S->>DB: insert room + share code
S-->>A: room_path /room/ID
A->>S: GET /api/rooms/ID/stream — SSE
S-->>A: init — you + participants
B->>S: POST /api/join-by-code
S-->>B: room_path /room/ID
B->>S: GET /api/rooms/ID/stream — SSE
S->>A: join_request — admit or decline
S-->>B: join_pending
A->>S: POST /send admit_join
S-->>B: init + presence_update
S->>A: presence_update
Note over A,B: live edits flow both ways
B->>S: POST /send object_updated · photo_pinned · cursor
S->>A: broadcast, excluding the sender
A->>S: POST /send screen_changed
S->>B: broadcast — followers mirror the view
Note over S,DB: durable layer, debounced
A->>S: PUT /api/rooms/ID/state
S->>DB: upsert room_state, versioned
B->>S: GET /api/rooms/ID/state on join
S-->>B: saved scrapbook
```
A few design choices worth highlighting:
- **Host approval.** A joiner is *pending* until the host (the owner if connected,
else the earliest-connected participant, so approval never deadlocks) admits
them. Every SSE connection gets a queue up front — keyed by `session_id` — so a
pending joiner can still receive `join_pending` / `join_declined` before they're
a full participant.
- **Presence & "what are you touching."** Beyond cursors, the app emits throttled
`interaction` pings on button presses, field focus, and typing (captured
app-wide in the capture phase), so collaborators literally *see* labels float
near each other's cursors — "🔘 create room", "✏️ caption".
- **Follow vs. roam.** Screen and map-variant changes are broadcast; a per-user
"Following" toggle decides whether you mirror a peer's view or roam
independently. Accessibility prefs live in `localStorage` and are *never*
broadcast — everyone keeps their own a11y settings while everything else syncs.
- **No echo loop.** The shared store never re-broadcasts events it receives, so
applying a peer's `photo_moved` doesn't bounce back out.
- **Two layers of state.** Live edits flow over SSE for immediacy;
`RoomPersistence` is the *durable* layer — a debounced (1.2s), version-tracked
`PUT /api/rooms/{id}/state` that lets a returning user (or a late joiner) load
the scrapbook exactly as it was left. On a version conflict, the client adopts
the server's newer state.
- **Notifications without a server.** Chimes are synthesized in-browser with the
Web Audio API (themed note sequences), because the Python `chime` library would
only make sound on the machine running Python — useless when every participant
needs to hear it in their own browser.
---
## Part 7 — The AI plumbing
Four models, four endpoints, one pattern. Every model is loaded **lazily** and
**cached**, runs inference inside a `@spaces.GPU` worker, and degrades gracefully
if it's missing.
```mermaid
flowchart LR
subgraph BROWSER["Browser"]
U1["upload a photo"]
U2["edit a field, hit Auto"]
U3["hold the mic"]
U4["enter the Darkroom"]
end
subgraph MAIN["app.py main process — CPU only"]
NSFW["/api/nsfw"]
CAP["autocaption — Gradio API"]
ASR["/api/transcribe"]
I2I["/api/img2img"]
TR["traced -> Phoenix span"]
DL["bg snapshot_download
writes HF cache"]
end
subgraph GPU["ZeroGPU worker — @spaces.GPU, CUDA only here"]
GN["_nsfw_infer — Falconsai"]
GC["_caption_infer — MiniCPM-V"]
GA["_asr_run — NeMo"]
GI["_img2img_infer — FLUX.2-klein + LoRA"]
end
U1 --> NSFW --> GN
U2 --> CAP --> GC
U3 --> ASR --> GA
U4 --> I2I --> GI
CAP --> TR
ASR --> TR
I2I --> DL
GI -. reads cache .-> DL
```
| Model | Job | Endpoint | Worker fn |
|-------|-----|----------|-----------|
| `Falconsai/nsfw_image_detection` | upload safety gate | `POST /api/nsfw` | `_nsfw_infer` |
| `MiniCPM-V` (GGUF, Q4_K_M) | caption / title / tags | `autocaption` (Gradio) | `_caption_infer` |
| NVIDIA NeMo streaming ASR | voice → text | `POST /api/transcribe` | `_asr_run` |
| `FLUX.2-klein-4B` + scene LoRA | photo → watercolour | `POST /api/img2img` | `_img2img_infer` |
`/api/config` reports which of these are available at runtime, so the UI can show
the active caption model, hide Cesium when there's no ion token, etc. **Phoenix**
runs in-process and is proxied at `/phoenix` (and embedded in the `trace` screen),
so every captioning/ASR call is observable without running a separate service.
---
## Part 8 — Lifecycle & housekeeping
Rooms are cheap and self-cleaning. State lives in SQLite (`rooms`,
`participants`, `room_state`) and a background daemon sweeps every 5 minutes:
1. Rooms idle > **1 hour** are flagged inactive (can't be joined).
2. Rooms idle past **1h10m** are deleted outright (keyed on `last_activity`, so
the DB tracks live usage, not creation time).
3. A hard cap deletes anything older than **24 hours** regardless.
4. Orphaned `participants` / `room_state` rows are cascaded away, and an hourly
`PRAGMA incremental_vacuum` returns freed pages to the OS so the heavy
`room_state` blobs don't bloat the file.
The cleanup is keyed on activity, not age, which is exactly what you want for
ephemeral collaborative rooms: an actively used room stays alive; an abandoned one
disappears within the grace window.
---
## The shape of it
If you remember four things about Sozai's architecture:
1. **One page, one process, three routes.** `/` is the splash; `/app` and
`/room/{id}` are the *same* SPA in solo vs. collaborative mode.
2. **The screen is a string.** `add → darkroom → map` is the whole user journey,
driven by one piece of state.
3. **CUDA lives only in the worker.** Downloads happen in the CPU main process;
inference happens behind `@spaces.GPU`; only plain data crosses the boundary.
4. **Live + durable are two layers.** SSE+POST for immediacy, a debounced
versioned `room_state` PUT for "load it back exactly as we left it."
Everything else — the watercolour develop, the floating interaction labels, the
synthesized chimes, the 3D globe — is detail hung on that frame.
*For the deep dive on the FLUX.2 img2img pipeline, the pickle boundary, and the
ZeroGPU cache lifetimes, see [`architecture.md`](architecture.md).*