indic-heritage-studio / docs /architecture.md
Dev2506's picture
Add files using upload-large-folder tool
15d68eb verified
|
Raw
History Blame Contribute Delete
20.7 kB

Indic Heritage Studio v2 β€” System Architecture

High-level diagram

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                          USER INTERFACE (Gradio)                              β”‚
β”‚                                                                               β”‚
│  Tab 1: Text→Image   Tab 2: Style Transfer   Tab 3: Image→Video              │
β”‚  Tab 4: ControlNet   Tab 5: Inpainting        Tab 6: Batch (multi-GPU)       β”‚
β”‚  + AI Style Advisor widget (calls Qwen via free AMD API)                     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                   β”‚
                                   β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    AGENT LAYER (free AMD Model API)                            β”‚
β”‚                                                                               β”‚
β”‚  StyleAdvisor       β†’ recommends heritage style (JSON)                        β”‚
β”‚  PromptEngineer     β†’ enriches prompt with style keywords (SDXL-aware)        β”‚
β”‚  Critic             β†’ scores output, suggests regeneration                    β”‚
β”‚                                                                               β”‚
β”‚  LLM: Qwen3.6-35B-A3B (fallback: DeepSeek-V4-Flash)                          β”‚
β”‚  Endpoint: https://developer.amd.com.cn/radeon/api/v1                        β”‚
β”‚  Cost: $0 β€” does NOT burn Radeon GPU credits                                  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                   β”‚ (only when generation is needed)
                                   β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚          CORE GPU LAYER β€” MULTI-GPU (8 Γ— NVIDIA 80GB dev / AMD Radeon final)  β”‚
β”‚                                                                               β”‚
β”‚  GPU 0: SDXL 1.0 + DreamShaper-XL turbo    (T2I + Inpainting)                β”‚
β”‚  GPU 1: SDXL + IP-Adapter XL                (Style transfer)                  β”‚
β”‚  GPU 2: Stable Video Diffusion XT 1.1       (Image β†’ Video)                  β”‚
β”‚  GPU 3: SDXL + ControlNet (Canny/Depth/Pose)(Composition control)            β”‚
β”‚  GPUs 4-7: 4 parallel batch workers         (Batch processing)               β”‚
β”‚                                                                               β”‚
β”‚  Per-style LoRAs (5 Γ— ~150 MB) loaded on demand by all SDXL pipelines         β”‚
β”‚  Hardware: 8 Γ— NVIDIA A100/H100 80GB (dev) β€” AMD Radeon Cloud (final demo)   β”‚
β”‚  Stack:    PyTorch 2.4.1 + ROCm 6.2 + Diffusers 0.30 + PEFT 0.12             β”‚
β”‚  Cost:     NVIDIA dev = $0; AMD final demo = ~2 of 10 credits                β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Rule compliance

The Track 1 Rules & Conditions specify:

  • "It is not allowed to rely solely on closed-source online APIs for core functions."
  • "At least one key inference process must run locally on AMD Radeon GPU."

Indic Heritage Studio v2 complies on every axis:

  • All 6 core pipelines (T2I, style transfer, I2V, ControlNet, inpainting, batch) run 100% on AMD Radeon GPU during the final demo + benchmark.
  • Agent layer (Qwen/DeepSeek) is optional UX enrichment β€” the studio works fully even if the AMD Model API is down. Each agent has a deterministic fallback (agents/*.py).
  • Per-style LoRAs are trained locally on NVIDIA hardware (free) and loaded at inference time on AMD Radeon β€” the LoRA weights are committed to the repo, so AMD doesn't need to retrain anything.

Module dependency graph

config/settings.py ─────────────┐
config/styles.py ────────────────
                                β”œβ”€β”€> agents/base.py ──> agents/style_advisor.py
                                β”‚                    β”œβ”€> agents/prompt_engineer.py
                                β”‚                    └─> agents/critic.py
                                β”‚
                                β”œβ”€β”€> core/text_to_image.py    ──> uses agents/prompt_engineer + per-style LoRA
                                β”œβ”€β”€> core/style_transfer.py   ──> uses IP-Adapter XL + per-style LoRA
                                β”œβ”€β”€> core/image_to_video.py   ──> Stable Video Diffusion
                                β”œβ”€β”€> core/controlnet.py       ──> Canny / Depth / OpenPose
                                β”œβ”€β”€> core/inpainting.py       ──> SDXL Inpainting
                                └──> core/batch_processor.py  ──> uses core/text_to_image OR core/style_transfer
                                                            β”‚
                                                            └──> ui/gradio_app.py
                                                                 β”‚
                                                                 └──> app.py

training/prepare_dataset.py ──> assets/datasets/<style>/
training/train_lora.py      ──> assets/loras/<style>.safetensors
                                            β”‚
                                            └──> loaded by core/text_to_image, core/style_transfer, core/inpainting

utils/gpu_utils.py           ──> multi-GPU device management, shard_workload, VRAMGuard

Data flow β€” Text β†’ Image (with LoRA)

User enters: "a young woman reading under a banyan tree"
                    β”‚
                    β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ (Optional) StyleAdvisor.recommend(prompt)                   β”‚
β”‚   β†’ "madhubani" (free AMD Qwen call)                        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                    β”‚
                    β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ PromptEngineer.enrich(prompt, madhubani_spec)               β”‚
β”‚   β†’ "a young woman reading under a banyan tree,             β”‚
β”‚       madhubani painting style, dense geometric patterns,   β”‚
β”‚       double-lined borders, fine linework…"                 β”‚
β”‚   + builds negative prompt from style spec                  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                    β”‚
                    β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ TextToImagePipeline._maybe_swap_lora(madhubani_style)       β”‚
β”‚   β†’ loads assets/loras/madhubani.safetensors                β”‚
β”‚   β†’ fuses at scale=0.9 (style.lora_scale)                   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                    β”‚
                    β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ TextToImagePipeline.generate(enriched_prompt, style)        β”‚
β”‚   β†’ StableDiffusionXLPipeline.from_pretrained("Lykon/…")    β”‚
β”‚   β†’ pipe(prompt, negative, steps=25, guidance=7.0,          β”‚
β”‚          width=1024, height=1024)                           β”‚
β”‚   β†’ returns PIL.Image                                        β”‚
β”‚                                                              β”‚
β”‚   BURNS ~3 GPU-SECONDS on NVIDIA / ~5 on AMD Radeon         β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                    β”‚
                    β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ (Optional) Critic.evaluate(image, style, prompt)            β”‚
β”‚   β†’ {"style_fidelity": 8, "composition": 7, "tech": 9,      β”‚
β”‚       "overall": 8.0, "should_regenerate": false}           β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                    β”‚
                    β–Ό
                 Return image to user

Multi-GPU architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                       8 Γ— NVIDIA 80GB GPUs                       β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚  GPU 0              GPU 1              GPU 2              GPU 3 β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”β”‚
β”‚  β”‚ SDXL base  β”‚    β”‚ SDXL base  β”‚    β”‚ SVD-XT 1.1 β”‚    β”‚SDXL+ β”‚β”‚
β”‚  β”‚ + LoRA     β”‚    β”‚ + IP-Adapt β”‚    β”‚ (25-frame  β”‚    β”‚CtrlNetβ”‚β”‚
β”‚  β”‚ T2I +      β”‚    β”‚ XL         β”‚    β”‚  1024Γ—576) β”‚    β”‚(Cannyβ”‚β”‚
β”‚  β”‚ Inpainting β”‚    β”‚ Style      β”‚    β”‚ I2V        β”‚    β”‚/Depthβ”‚β”‚
β”‚  β”‚            β”‚    β”‚ Transfer   β”‚    β”‚            β”‚    β”‚/Pose)β”‚β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”˜β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚  GPU 4              GPU 5              GPU 6              GPU 7 β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”β”‚
β”‚  β”‚ Batch      β”‚    β”‚ Batch      β”‚    β”‚ Batch      β”‚    β”‚Batch β”‚β”‚
β”‚  β”‚ worker 0   β”‚    β”‚ worker 1   β”‚    β”‚ worker 2   β”‚    β”‚work 3β”‚β”‚
β”‚  β”‚ (own SDXL  β”‚    β”‚ (own SDXL  β”‚    β”‚ (own SDXL  β”‚    β”‚(own  β”‚β”‚
β”‚  β”‚  pipeline) β”‚    β”‚  pipeline) β”‚    β”‚  pipeline) β”‚    β”‚ SDXL)β”‚β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”˜β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
  • Pipelines 0-3 are resident (loaded once, served forever). No model thrash.
  • Batch workers 4-7 spawn on demand via multiprocessing.spawn, each with its own SDXL pipeline copy. They consume jobs from a shared multiprocessing.Queue.
  • On AMD Radeon Cloud (single GPU): all pipelines share GPU 0 and load/unload on demand. The code auto-detects device_count and adjusts the strategy β€” no flags to flip.

Per-style LoRA training

Heritage art images (raw, ~40 per style)
        β”‚
        β–Ό
training/prepare_dataset.py
  - resize to 1024Γ—1024
  - generate captions from style.prompt_tags
  - write metadata.jsonl
        β”‚
        β–Ό
assets/datasets/<style>/
  - <style>_001.jpg
  - <style>_001.txt  (caption)
  - …
  - metadata.jsonl
        β”‚
        β–Ό
training/train_lora.py
  - SDXL UNet + LoRA (rank=32, alpha=32, dropout=0.05)
  - target_modules: to_q, to_k, to_v, to_out.0, proj_in, proj_out
  - AdamW 8-bit, lr=1e-4, cosine schedule, 800 steps
  - mixed precision (bf16 on A100/H100, fp16 fallback)
  - batch_size=1, grad_accum=4 (effective batch=4)
  - ~30 min per style on a single GPU
        β”‚
        β–Ό
assets/loras/<style>.safetensors  (~150 MB per style)
        β”‚
        β–Ό
Loaded at inference time by:
  - core/text_to_image.py (style.lora_scale)
  - core/style_transfer.py (style.lora_scale, IP-Adapter overlay)
  - core/inpainting.py (style.lora_scale)

SDXL + ControlNet flow

User supplies conditioning image (e.g. a sketch, a photo of a pose)
        β”‚
        β–Ό
ControlNetPipeline.detect(image, condition_type)
  - "canny"    β†’ CannyDetector    (edge map)
  - "depth"    β†’ MidasDetector    (depth map)
  - "openpose" β†’ OpenposeDetector (body keypoint skeleton)
        β”‚
        β–Ό
conditioning_image (PIL, single channel)
        β”‚
        β–Ό
ControlNetPipeline.generate(conditioning_image, prompt, style)
  - SDXL + ControlNet loaded on GPU 3
  - Optional LoRA overlay for the chosen style
  - controlnet_conditioning_scale = 0.8 (user-tunable)
        β”‚
        β–Ό
Generated heritage art that follows the conditioning composition

Use cases:

  • Tanjore painting: OpenPose skeleton enforces frontal symmetrical deity pose.
  • Mughal miniature: Canny edge map of a sketch preserves composition while applying Mughal brushwork.
  • Warli tarpa dance: OpenPose of dancers preserves the spiral composition.

SDXL inpainting flow

User supplies:
  - source image (e.g. a damaged heritage painting scan, or a modern photo)
  - mask (white = inpaint, black = keep)
  - prompt ("ornate floral border with peacock motifs")
  - style choice (e.g. madhubani)
        β”‚
        β–Ό
InpaintingPipeline.inpaint(image, mask, prompt, style)
  - SDXL Inpainting checkpoint loaded on GPU 0
  - Optional LoRA overlay
  - strength = 1.0 (full replacement of masked region)
  - 30 steps
        β”‚
        β–Ό
Output: original image with masked region replaced by heritage-style art

Use cases:

  • Heritage restoration: mask damaged sections of a real Tanjore painting, regenerate in Tanjore style.
  • Creative compositing: photorealistic portrait + Tanjore-style background.
  • Style framing: keep a modern subject, replace the background with a Pattachitra-style scene.

Stable Video Diffusion flow

User supplies: a still image (e.g. an SDXL-generated Madhubani scene)
        β”‚
        β–Ό
ImageToVideoPipeline.generate(image, style)
  - SVD-XT-1.1 loaded on GPU 2
  - Per-style motion tuning via style.svd_motion_bucket:
      - tanjore: 80  (devotional icon β€” barely any motion)
      - mughal:  90  (court scene β€” minimal motion, dignified)
      - pattachitra: 100 (scroll painting β€” very subtle motion)
      - madhubani: 110 (subtle ritual-like motion)
      - warli:   180 (tarpa dance is dynamic β€” more motion)
  - num_frames = 25, fps = 8 β†’ ~3-second video
  - resolution = 1024Γ—576 (SVD native)
  - decode_chunk_size = 8 (VRAM-friendly)
        β”‚
        β–Ό
25 PIL frames β†’ utils/video_utils.frames_to_mp4 β†’ MP4 file

ROCm optimization layer (applied on AMD instance)

Optimization Where applied Why
torch_dtype=float16 config/settings.py Mandatory for ROCm speed
ATTN_PRECISION=fp16 env var scripts/day1_setup.sh Avoids fp32 attention fallback
HSA_OVERRIDE_GFX_VERSION auto-set scripts/day1_setup.sh Some Radeon GPUs need explicit gfx version
SDPA attention (built-in) diffusers auto-detects Replaces xformers (CUDA-only)
enable_attention_slicing("auto") All SDXL pipelines Reduces VRAM peak
enable_vae_slicing() All SDXL pipelines + SVD Critical for SVD on 16 GB VRAM
torch.inference_mode() All generate() methods Disables autograd overhead
HF_HUB_OFFLINE=1 (after download) .env Avoids network calls on every launch
HF_HUB_ENABLE_HF_TRANSFER=1 .env Parallel HF downloads
Pipeline reuse (singleton) ui/gradio_app.py Load once, serve many requests
Multi-GPU pipeline pinning config/settings.py All 4 pipelines stay resident β€” zero thrash
Per-style LoRA hot-swap core/text_to_image.py etc. LoRA swap = ~3 sec, no full reload

Credit-budget architecture

The 10-credit AMD budget dictates the architecture:

  1. Agent layer is free β†’ all UX work happens without burning credits
  2. Pipelines are lazy-loaded singletons β†’ first call costs 1 load minute; subsequent calls are pure inference
  3. scripts/generate_demo_outputs.py runs ONCE on NVIDIA (Week 2, free) β†’ pre-bakes all 62 demo assets. The AMD demo recording session only re-runs a handful of commands live for the camera; the actual outputs already exist as fallback.
  4. Multi-GPU on NVIDIA = ~4Γ— faster than single-GPU AMD β†’ dev iteration is essentially free; AMD is only used when the rules require it.

Failure modes & fallbacks

Failure Fallback Where
AMD Model API down Deterministic prompt enrichment agents/prompt_engineer.py:_heuristic_enrich
Style Advisor returns bad JSON Default to "madhubani" agents/style_advisor.py:_heuristic_recommend
Per-style LoRA missing Prompt-only style guidance core/text_to_image.py:_maybe_swap_lora
Reference style image missing Palette-fallback synthetic image core/style_transfer.py:_load_style_reference
SVD OOM on AMD 16GB Reduce num_frames to 14, decode_chunk_size to 4 CLI flag --frames 14
ControlNet preprocessor download fails Hard error (graceful β€” user must rerun) core/controlnet.py:_load_processor
/workspace not persistent on AMD Re-download models via scripts/download_models.py scripts/day1_setup.sh
Multi-GPU spawn fails on AMD (1 GPU) Falls back to single-process sequential core/batch_processor.py

Performance characteristics

Expected on 8 Γ— NVIDIA 80GB:

Pipeline Latency (1 image) Throughput (4-GPU batch) Peak VRAM
T2I (SDXL, 1024Β², 25 steps, +LoRA) ~3.5 s ~60 img/min ~14 GB/GPU
Style transfer (SDXL+IP-Adapter, 1024Β², 30 steps) ~5 s ~45 img/min ~16 GB/GPU
Image β†’ Video (SVD, 25 frames, 1024Γ—576) ~25 s n/a (single-GPU) ~22 GB
ControlNet (SDXL+Canny, 1024Β², 30 steps) ~5 s ~45 img/min ~18 GB/GPU
Inpainting (SDXL Inpaint, 1024Β², 30 steps) ~4.5 s ~50 img/min ~15 GB/GPU

Expected on AMD Radeon Cloud (single GPU, ~16 GB VRAM):

Pipeline Latency (1 image) Peak VRAM
T2I ~7 s ~10 GB
Style transfer ~10 s ~12 GB
Image β†’ Video (25f) ~60 s ~14 GB
ControlNet ~10 s ~13 GB
Inpainting ~9 s ~11 GB

The AMD numbers are slower but well within the demo recording time budget (60 minutes total, including 5 live generation runs).