# Pipeline Architecture ## Directory Layout ``` spatial-pipeline/ │ ├── app.py ← Gradio UI entry point ├── pipeline.py ← SpatialPipeline orchestrator │ ├── configs/ │ ├── __init__.py │ └── model_registry.py ← All supported models, per stage │ ├── models/ │ ├── __init__.py ← build_loader() factory │ ├── base_loader.py ← Abstract BaseLoader interface │ ├── background_removal.py ← BiRefNetLoader (matte + crop + composite) │ ├── depth_estimation.py ← TransformersDepthLoader │ └── reconstruction.py ← Open3D / GaussianSplat / DepthSplat loaders │ ├── utils/ │ ├── __init__.py │ ├── device.py ← CUDA/MPS/CPU detection │ ├── image_utils.py ← PIL/numpy/torch conversions, PLY I/O │ └── hf_utils.py ← HF Hub validation, search │ ├── outputs/ ← Generated images, depth maps, PLY files ├── requirements.txt ├── packages.txt ← apt packages for HF Spaces └── README.md ← HF Space card ``` ## Stage Data Flow ``` [User-supplied Image] │ ▼ ┌─────────────────────────────────────────┐ │ Stage 1: Background Removal │ │ BiRefNetLoader │ │ • AutoModelForImageSegmentation │ │ • Supports BiRefNet / BiRefNet Lite / │ │ BiRefNet Portrait / custom │ │ • Recipe borrowed from TripoSplat: │ │ predict alpha → erode edge → │ │ crop to bbox (+margin) → composite │ │ Output: PIL Image (RGB, composited), │ │ RGBA preview, alpha matte │ └─────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────┐ │ Stage 2: Image → Depth │ │ TransformersDepthLoader │ │ • HF pipeline("depth-estimation") │ │ • Depth Anything V2 / DPT / MiDaS │ │ Output: depth_raw, depth_normalised, │ │ depth_colourmap, depth_uint16 │ └─────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────┐ │ Stage 3: RGBD → 3D │ │ Open3DReconstructionLoader │ │ Back-project → coloured PLY, │ │ masked by the Stage 1 alpha matte │ │ GaussianSplatLoader │ │ Convert point cloud to 3DGS init PLY │ │ (compatible with graphdeco training) │ │ DepthSplatLoader (GPU only) │ │ Feed-forward Gaussian prediction │ │ Output: .ply file │ └─────────────────────────────────────────┘ ``` Stage 1's alpha matte is threaded all the way to Stage 3: `rgbd_to_pointcloud()` uses it (thresholded at 0.5) to decide which back-projected points are subject vs. background, instead of the old "reject near-black pixels" brightness heuristic. If Stage 1 is disabled for a run, reconstruction falls back to that brightness heuristic automatically. ## Swapping Models ### Via Gradio UI Each stage has a dropdown. Select **✏️ Custom HF model ID / URL**, type a model ID, click **Validate**, then run. ### Via Python API ```python from configs import get_config_by_display_name from pipeline import SpatialPipeline from PIL import Image pl = SpatialPipeline() result = pl.run( input_image = Image.open("photo.jpg"), bgremove_config = get_config_by_display_name("background_removal", "BiRefNet (general, MIT) — best quality"), depth_config = get_config_by_display_name("depth_estimation", "DPT-Large (Intel)"), recon_config = get_config_by_display_name("reconstruction", "Gaussian Splat Scaffold (CPU initialiser)"), ) ``` ### Adding a new model 1. Open `configs/model_registry.py` 2. Append a `ModelConfig` to the relevant list 3. If a new loader class is needed, add it in `models/` and register it in `models/__init__.py` ## Output Formats | File | Format | Use | |------|--------|-----| | `outputs/` (in-memory, shown in UI) | PNG RGBA | Background-removed preview | | `outputs/depth_colourmap.png` | PNG RGB | Visualisation only | | `outputs/depth_16bit.png` | PNG 16-bit greyscale | Metric-preserving depth | | `outputs/pointcloud.ply` | Binary PLY | Open3D, MeshLab, CloudCompare | | `outputs/gaussian_scaffold.ply` | Binary PLY (3DGS layout) | gaussian-splatting training | ## Running Locally ```bash # Install system deps (Ubuntu/Debian) sudo apt install libgl1-mesa-glx libglib2.0-0 # Python env python -m venv .venv && source .venv/bin/activate pip install -r requirements.txt # Launch python app.py # → http://localhost:7860 ``` ## Deploying to HuggingFace Spaces 1. Create a new Space at huggingface.co/new-space - SDK: Gradio - Hardware: **ZeroGPU** (recommended — free dynamic GPU access; requires a PRO, Team, or Enterprise account to host, see [Spaces ZeroGPU docs](https://huggingface.co/docs/hub/spaces-zerogpu)). CPU Basic still works for testing, but Stage 2/3 models — and DepthSplat in particular, which raises `RuntimeError` on CPU — run far slower or not at all. 2. `git clone https://huggingface.co/spaces//` 3. Copy all files from this repo into the cloned directory 4. `git add . && git commit -m "init" && git push` 5. **Selecting ZeroGPU in step 1 only reserves the hardware tier — you still need to explicitly pick "ZeroGPU" from the Hardware dropdown in the Space's Settings tab** (or via `huggingface_hub`'s `request_space_hardware`) after the Space is created; it isn't set from files in this repo. `app.py` already imports `spaces` and decorates the pipeline entry point with `@spaces.GPU`, which is a no-op until that hardware tier is actually selected. HF Spaces will automatically install `packages.txt` (apt) then `requirements.txt` (pip) and launch `app.py`. ## Credit The background-removal stage's model family (BiRefNet) and pre-processing recipe (predict → erode → crop-to-bbox → composite) are borrowed from [VAST-AI/TripoSplat](https://huggingface.co/spaces/VAST-AI/TripoSplat)'s own foreground-matting stage. This project loads BiRefNet via the standard `transformers.AutoModelForImageSegmentation` API rather than TripoSplat's from-scratch Swin-L port, since that keeps it consistent with the rest of this codebase's "swap any HF model id" loader pattern and doesn't depend on TripoSplat's bundled, non-separable checkpoint.