gaussian_studio / README.md
dgarch424's picture
Upload 21 files
728fc83 verified
|
Raw
History Blame Contribute Delete
2.88 kB
---
title: Image Depth 3D Gaussian Splatting Pipeline
emoji: 🌐
colorFrom: indigo
colorTo: purple
sdk: gradio
sdk_version: 4.44.0
python_version: "3.12"
app_file: app.py
pinned: false
license: apache-2.0
tags:
- background-removal
- depth-estimation
- gaussian-splatting
- 3d-reconstruction
- image-segmentation
- transformers
secrets:
- HF_TOKEN
---
# 🌐 Image → Depth → 3DGS Pipeline
A modular, end-to-end pipeline that chains three stages on a **user-supplied image**:
1. **Background Removal** — Matte the subject out with [BiRefNet](https://huggingface.co/ZhengPeng7/BiRefNet) (MIT), crop to it, and composite it onto a clean background before anything else runs
2. **Depth Estimation** — Predict a dense depth map from the background-removed image (Depth Anything v2 / MiDaS / DPT)
3. **3D Gaussian Splatting init** — Convert the RGBD data into a point cloud / Gaussian splat scaffold, using the background-removal matte to cleanly separate subject from background points
Each stage uses a **swappable model** selected from a dropdown or custom HuggingFace model ID.
## Why background removal first?
Monocular depth models are noticeably better behaved on an isolated subject than on a
photo with a busy background — there's no sky/horizon to confuse the depth ordering, and
the reconstruction stage can use the real segmentation matte (instead of a "is this pixel
dark" brightness guess) to decide which back-projected points belong to the subject.
The background-removal stage borrows both its model family and its pre-processing recipe
(predict alpha → erode the matte edge → crop to the subject's bounding box with a margin →
composite onto a solid colour) from [VAST-AI/TripoSplat](https://huggingface.co/spaces/VAST-AI/TripoSplat),
which uses BiRefNet for the same purpose ahead of its own Gaussian-splat generator. Rather
than re-host TripoSplat's from-scratch Swin-L port and its bundled checkpoint, this project
loads BiRefNet the standard way, via `transformers.AutoModelForImageSegmentation`, so it
drops into the same "swap any HF model id" pattern as the depth and reconstruction stages.
## Models
| Stage | Default | Alternatives |
|-------|---------|-------------|
| Background Removal | `ZhengPeng7/BiRefNet` | BiRefNet Lite, BiRefNet Portrait, custom |
| Depth Estimation | `depth-anything/Depth-Anything-V2-Small-hf` | MiDaS v3, DPT-Large, ZoeDepth |
| 3D Reconstruction | Point cloud from RGBD | Gaussian splat scaffold |
All background-removal presets are MIT-licensed and trained on the open DIS5K dataset.
## Usage
1. Upload an image
2. Select models for each stage (or paste a HuggingFace model ID)
3. Run the full pipeline or individual stages
4. Download outputs: background-removed image, depth map, point cloud PLY
## Local Setup
```bash
pip install -r requirements.txt
python app.py
```