indic-heritage-studio / docs /architecture.md
Dev2506's picture
Add files using upload-large-folder tool
15d68eb verified
|
Raw
History Blame Contribute Delete
20.7 kB
# Indic Heritage Studio v2 β€” System Architecture
## High-level diagram
```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ USER INTERFACE (Gradio) β”‚
β”‚ β”‚
│ Tab 1: Text→Image Tab 2: Style Transfer Tab 3: Image→Video │
β”‚ Tab 4: ControlNet Tab 5: Inpainting Tab 6: Batch (multi-GPU) β”‚
β”‚ + AI Style Advisor widget (calls Qwen via free AMD API) β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ AGENT LAYER (free AMD Model API) β”‚
β”‚ β”‚
β”‚ StyleAdvisor β†’ recommends heritage style (JSON) β”‚
β”‚ PromptEngineer β†’ enriches prompt with style keywords (SDXL-aware) β”‚
β”‚ Critic β†’ scores output, suggests regeneration β”‚
β”‚ β”‚
β”‚ LLM: Qwen3.6-35B-A3B (fallback: DeepSeek-V4-Flash) β”‚
β”‚ Endpoint: https://developer.amd.com.cn/radeon/api/v1 β”‚
β”‚ Cost: $0 β€” does NOT burn Radeon GPU credits β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚ (only when generation is needed)
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ CORE GPU LAYER β€” MULTI-GPU (8 Γ— NVIDIA 80GB dev / AMD Radeon final) β”‚
β”‚ β”‚
β”‚ GPU 0: SDXL 1.0 + DreamShaper-XL turbo (T2I + Inpainting) β”‚
β”‚ GPU 1: SDXL + IP-Adapter XL (Style transfer) β”‚
β”‚ GPU 2: Stable Video Diffusion XT 1.1 (Image β†’ Video) β”‚
β”‚ GPU 3: SDXL + ControlNet (Canny/Depth/Pose)(Composition control) β”‚
β”‚ GPUs 4-7: 4 parallel batch workers (Batch processing) β”‚
β”‚ β”‚
β”‚ Per-style LoRAs (5 Γ— ~150 MB) loaded on demand by all SDXL pipelines β”‚
β”‚ Hardware: 8 Γ— NVIDIA A100/H100 80GB (dev) β€” AMD Radeon Cloud (final demo) β”‚
β”‚ Stack: PyTorch 2.4.1 + ROCm 6.2 + Diffusers 0.30 + PEFT 0.12 β”‚
β”‚ Cost: NVIDIA dev = $0; AMD final demo = ~2 of 10 credits β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```
## Rule compliance
The Track 1 Rules & Conditions specify:
- *"It is not allowed to rely solely on closed-source online APIs for core functions."*
- *"At least one key inference process must run locally on AMD Radeon GPU."*
Indic Heritage Studio v2 complies on every axis:
- **All 6 core pipelines** (T2I, style transfer, I2V, ControlNet, inpainting, batch) run 100% on AMD Radeon GPU during the final demo + benchmark.
- **Agent layer** (Qwen/DeepSeek) is *optional UX enrichment* β€” the studio works fully even if the AMD Model API is down. Each agent has a deterministic fallback (`agents/*.py`).
- **Per-style LoRAs** are trained locally on NVIDIA hardware (free) and loaded at inference time on AMD Radeon β€” the LoRA weights are committed to the repo, so AMD doesn't need to retrain anything.
## Module dependency graph
```
config/settings.py ─────────────┐
config/styles.py ────────────────
β”œβ”€β”€> agents/base.py ──> agents/style_advisor.py
β”‚ β”œβ”€> agents/prompt_engineer.py
β”‚ └─> agents/critic.py
β”‚
β”œβ”€β”€> core/text_to_image.py ──> uses agents/prompt_engineer + per-style LoRA
β”œβ”€β”€> core/style_transfer.py ──> uses IP-Adapter XL + per-style LoRA
β”œβ”€β”€> core/image_to_video.py ──> Stable Video Diffusion
β”œβ”€β”€> core/controlnet.py ──> Canny / Depth / OpenPose
β”œβ”€β”€> core/inpainting.py ──> SDXL Inpainting
└──> core/batch_processor.py ──> uses core/text_to_image OR core/style_transfer
β”‚
└──> ui/gradio_app.py
β”‚
└──> app.py
training/prepare_dataset.py ──> assets/datasets/<style>/
training/train_lora.py ──> assets/loras/<style>.safetensors
β”‚
└──> loaded by core/text_to_image, core/style_transfer, core/inpainting
utils/gpu_utils.py ──> multi-GPU device management, shard_workload, VRAMGuard
```
## Data flow β€” Text β†’ Image (with LoRA)
```
User enters: "a young woman reading under a banyan tree"
β”‚
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ (Optional) StyleAdvisor.recommend(prompt) β”‚
β”‚ β†’ "madhubani" (free AMD Qwen call) β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ PromptEngineer.enrich(prompt, madhubani_spec) β”‚
β”‚ β†’ "a young woman reading under a banyan tree, β”‚
β”‚ madhubani painting style, dense geometric patterns, β”‚
β”‚ double-lined borders, fine linework…" β”‚
β”‚ + builds negative prompt from style spec β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ TextToImagePipeline._maybe_swap_lora(madhubani_style) β”‚
β”‚ β†’ loads assets/loras/madhubani.safetensors β”‚
β”‚ β†’ fuses at scale=0.9 (style.lora_scale) β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ TextToImagePipeline.generate(enriched_prompt, style) β”‚
β”‚ β†’ StableDiffusionXLPipeline.from_pretrained("Lykon/…") β”‚
β”‚ β†’ pipe(prompt, negative, steps=25, guidance=7.0, β”‚
β”‚ width=1024, height=1024) β”‚
β”‚ β†’ returns PIL.Image β”‚
β”‚ β”‚
β”‚ BURNS ~3 GPU-SECONDS on NVIDIA / ~5 on AMD Radeon β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ (Optional) Critic.evaluate(image, style, prompt) β”‚
β”‚ β†’ {"style_fidelity": 8, "composition": 7, "tech": 9, β”‚
β”‚ "overall": 8.0, "should_regenerate": false} β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β–Ό
Return image to user
```
## Multi-GPU architecture
```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ 8 Γ— NVIDIA 80GB GPUs β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ GPU 0 GPU 1 GPU 2 GPU 3 β”‚
β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”β”‚
β”‚ β”‚ SDXL base β”‚ β”‚ SDXL base β”‚ β”‚ SVD-XT 1.1 β”‚ β”‚SDXL+ β”‚β”‚
β”‚ β”‚ + LoRA β”‚ β”‚ + IP-Adapt β”‚ β”‚ (25-frame β”‚ β”‚CtrlNetβ”‚β”‚
β”‚ β”‚ T2I + β”‚ β”‚ XL β”‚ β”‚ 1024Γ—576) β”‚ β”‚(Cannyβ”‚β”‚
β”‚ β”‚ Inpainting β”‚ β”‚ Style β”‚ β”‚ I2V β”‚ β”‚/Depthβ”‚β”‚
β”‚ β”‚ β”‚ β”‚ Transfer β”‚ β”‚ β”‚ β”‚/Pose)β”‚β”‚
β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”˜β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ GPU 4 GPU 5 GPU 6 GPU 7 β”‚
β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”β”‚
β”‚ β”‚ Batch β”‚ β”‚ Batch β”‚ β”‚ Batch β”‚ β”‚Batch β”‚β”‚
β”‚ β”‚ worker 0 β”‚ β”‚ worker 1 β”‚ β”‚ worker 2 β”‚ β”‚work 3β”‚β”‚
β”‚ β”‚ (own SDXL β”‚ β”‚ (own SDXL β”‚ β”‚ (own SDXL β”‚ β”‚(own β”‚β”‚
β”‚ β”‚ pipeline) β”‚ β”‚ pipeline) β”‚ β”‚ pipeline) β”‚ β”‚ SDXL)β”‚β”‚
β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”˜β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```
- **Pipelines 0-3 are resident** (loaded once, served forever). No model thrash.
- **Batch workers 4-7** spawn on demand via `multiprocessing.spawn`, each with its own SDXL pipeline copy. They consume jobs from a shared `multiprocessing.Queue`.
- On **AMD Radeon Cloud (single GPU)**: all pipelines share GPU 0 and load/unload on demand. The code auto-detects `device_count` and adjusts the strategy β€” no flags to flip.
## Per-style LoRA training
```
Heritage art images (raw, ~40 per style)
β”‚
β–Ό
training/prepare_dataset.py
- resize to 1024Γ—1024
- generate captions from style.prompt_tags
- write metadata.jsonl
β”‚
β–Ό
assets/datasets/<style>/
- <style>_001.jpg
- <style>_001.txt (caption)
- …
- metadata.jsonl
β”‚
β–Ό
training/train_lora.py
- SDXL UNet + LoRA (rank=32, alpha=32, dropout=0.05)
- target_modules: to_q, to_k, to_v, to_out.0, proj_in, proj_out
- AdamW 8-bit, lr=1e-4, cosine schedule, 800 steps
- mixed precision (bf16 on A100/H100, fp16 fallback)
- batch_size=1, grad_accum=4 (effective batch=4)
- ~30 min per style on a single GPU
β”‚
β–Ό
assets/loras/<style>.safetensors (~150 MB per style)
β”‚
β–Ό
Loaded at inference time by:
- core/text_to_image.py (style.lora_scale)
- core/style_transfer.py (style.lora_scale, IP-Adapter overlay)
- core/inpainting.py (style.lora_scale)
```
## SDXL + ControlNet flow
```
User supplies conditioning image (e.g. a sketch, a photo of a pose)
β”‚
β–Ό
ControlNetPipeline.detect(image, condition_type)
- "canny" β†’ CannyDetector (edge map)
- "depth" β†’ MidasDetector (depth map)
- "openpose" β†’ OpenposeDetector (body keypoint skeleton)
β”‚
β–Ό
conditioning_image (PIL, single channel)
β”‚
β–Ό
ControlNetPipeline.generate(conditioning_image, prompt, style)
- SDXL + ControlNet loaded on GPU 3
- Optional LoRA overlay for the chosen style
- controlnet_conditioning_scale = 0.8 (user-tunable)
β”‚
β–Ό
Generated heritage art that follows the conditioning composition
```
Use cases:
- **Tanjore painting**: OpenPose skeleton enforces frontal symmetrical deity pose.
- **Mughal miniature**: Canny edge map of a sketch preserves composition while applying Mughal brushwork.
- **Warli tarpa dance**: OpenPose of dancers preserves the spiral composition.
## SDXL inpainting flow
```
User supplies:
- source image (e.g. a damaged heritage painting scan, or a modern photo)
- mask (white = inpaint, black = keep)
- prompt ("ornate floral border with peacock motifs")
- style choice (e.g. madhubani)
β”‚
β–Ό
InpaintingPipeline.inpaint(image, mask, prompt, style)
- SDXL Inpainting checkpoint loaded on GPU 0
- Optional LoRA overlay
- strength = 1.0 (full replacement of masked region)
- 30 steps
β”‚
β–Ό
Output: original image with masked region replaced by heritage-style art
```
Use cases:
- **Heritage restoration**: mask damaged sections of a real Tanjore painting, regenerate in Tanjore style.
- **Creative compositing**: photorealistic portrait + Tanjore-style background.
- **Style framing**: keep a modern subject, replace the background with a Pattachitra-style scene.
## Stable Video Diffusion flow
```
User supplies: a still image (e.g. an SDXL-generated Madhubani scene)
β”‚
β–Ό
ImageToVideoPipeline.generate(image, style)
- SVD-XT-1.1 loaded on GPU 2
- Per-style motion tuning via style.svd_motion_bucket:
- tanjore: 80 (devotional icon β€” barely any motion)
- mughal: 90 (court scene β€” minimal motion, dignified)
- pattachitra: 100 (scroll painting β€” very subtle motion)
- madhubani: 110 (subtle ritual-like motion)
- warli: 180 (tarpa dance is dynamic β€” more motion)
- num_frames = 25, fps = 8 β†’ ~3-second video
- resolution = 1024Γ—576 (SVD native)
- decode_chunk_size = 8 (VRAM-friendly)
β”‚
β–Ό
25 PIL frames β†’ utils/video_utils.frames_to_mp4 β†’ MP4 file
```
## ROCm optimization layer (applied on AMD instance)
| Optimization | Where applied | Why |
|---|---|---|
| `torch_dtype=float16` | `config/settings.py` | Mandatory for ROCm speed |
| `ATTN_PRECISION=fp16` env var | `scripts/day1_setup.sh` | Avoids fp32 attention fallback |
| `HSA_OVERRIDE_GFX_VERSION` auto-set | `scripts/day1_setup.sh` | Some Radeon GPUs need explicit gfx version |
| SDPA attention (built-in) | `diffusers` auto-detects | Replaces xformers (CUDA-only) |
| `enable_attention_slicing("auto")` | All SDXL pipelines | Reduces VRAM peak |
| `enable_vae_slicing()` | All SDXL pipelines + SVD | Critical for SVD on 16 GB VRAM |
| `torch.inference_mode()` | All `generate()` methods | Disables autograd overhead |
| `HF_HUB_OFFLINE=1` (after download) | `.env` | Avoids network calls on every launch |
| `HF_HUB_ENABLE_HF_TRANSFER=1` | `.env` | Parallel HF downloads |
| Pipeline reuse (singleton) | `ui/gradio_app.py` | Load once, serve many requests |
| Multi-GPU pipeline pinning | `config/settings.py` | All 4 pipelines stay resident β€” zero thrash |
| Per-style LoRA hot-swap | `core/text_to_image.py` etc. | LoRA swap = ~3 sec, no full reload |
## Credit-budget architecture
The 10-credit AMD budget dictates the architecture:
1. **Agent layer is free** β†’ all UX work happens without burning credits
2. **Pipelines are lazy-loaded singletons** β†’ first call costs 1 load minute; subsequent calls are pure inference
3. **`scripts/generate_demo_outputs.py` runs ONCE on NVIDIA** (Week 2, free) β†’ pre-bakes all 62 demo assets. The AMD demo recording session only re-runs a handful of commands live for the camera; the actual outputs already exist as fallback.
4. **Multi-GPU on NVIDIA = ~4Γ— faster than single-GPU AMD** β†’ dev iteration is essentially free; AMD is only used when the rules require it.
## Failure modes & fallbacks
| Failure | Fallback | Where |
|---|---|---|
| AMD Model API down | Deterministic prompt enrichment | `agents/prompt_engineer.py:_heuristic_enrich` |
| Style Advisor returns bad JSON | Default to "madhubani" | `agents/style_advisor.py:_heuristic_recommend` |
| Per-style LoRA missing | Prompt-only style guidance | `core/text_to_image.py:_maybe_swap_lora` |
| Reference style image missing | Palette-fallback synthetic image | `core/style_transfer.py:_load_style_reference` |
| SVD OOM on AMD 16GB | Reduce `num_frames` to 14, `decode_chunk_size` to 4 | CLI flag `--frames 14` |
| ControlNet preprocessor download fails | Hard error (graceful β€” user must rerun) | `core/controlnet.py:_load_processor` |
| `/workspace` not persistent on AMD | Re-download models via `scripts/download_models.py` | `scripts/day1_setup.sh` |
| Multi-GPU spawn fails on AMD (1 GPU) | Falls back to single-process sequential | `core/batch_processor.py` |
## Performance characteristics
Expected on 8 Γ— NVIDIA 80GB:
| Pipeline | Latency (1 image) | Throughput (4-GPU batch) | Peak VRAM |
|---|---|---|---|
| T2I (SDXL, 1024Β², 25 steps, +LoRA) | ~3.5 s | ~60 img/min | ~14 GB/GPU |
| Style transfer (SDXL+IP-Adapter, 1024Β², 30 steps) | ~5 s | ~45 img/min | ~16 GB/GPU |
| Image β†’ Video (SVD, 25 frames, 1024Γ—576) | ~25 s | n/a (single-GPU) | ~22 GB |
| ControlNet (SDXL+Canny, 1024Β², 30 steps) | ~5 s | ~45 img/min | ~18 GB/GPU |
| Inpainting (SDXL Inpaint, 1024Β², 30 steps) | ~4.5 s | ~50 img/min | ~15 GB/GPU |
Expected on AMD Radeon Cloud (single GPU, ~16 GB VRAM):
| Pipeline | Latency (1 image) | Peak VRAM |
|---|---|---|
| T2I | ~7 s | ~10 GB |
| Style transfer | ~10 s | ~12 GB |
| Image β†’ Video (25f) | ~60 s | ~14 GB |
| ControlNet | ~10 s | ~13 GB |
| Inpainting | ~9 s | ~11 GB |
The AMD numbers are slower but well within the demo recording time budget (60 minutes total, including 5 live generation runs).