--- license: apache-2.0 language: - en pipeline_tag: image-to-image library_name: pytorch tags: - image-to-image - pose-transfer - pose-guided - human-pose - keypoints - diffusion - dit - rectified-flow - vtuber - anime - windows - cuda --- # VTM-Elf 0.01 Pose-conditioned character generation for VTuber pipelines. **Status:** beta · **Platform:** Windows + NVIDIA CUDA only --- ## Architecture - **Dual-stream DiT-30M** (pose + identity in one backbone) - Pose from keypoints; identity from a reference image - Rectified flow + SD VAE @ **768×768** - Checkpoint: `VTM-elf.pt-000060000` (~436 MB, step 60k) --- ## Measured latency | Hardware | Resolution | Observed rate | |---|---|---| | **RTX 5060 Ti** | 768×768 | **~7 FPS** | --- ## Data Training used a small private character set: - ~**1,500** characters - ~**16–30** images per character Generation quality in this release is limited mainly by **model capacity (DiT-30M)** and **dataset scale**. --- ## Getting started ### Requirements - **Windows** - **NVIDIA GPU + CUDA** - Other OS / AMD / CPU: **not supported** in this release ### Download ```bash hf download sinBoo1/VTM-Elf-0.01 VTM-elf.pt-000060000 --local-dir ./VTM-Elf-0.01 ``` ```python from huggingface_hub import hf_hub_download ckpt = hf_hub_download( repo_id="sinBoo1/VTM-Elf-0.01", filename="VTM-elf.pt-000060000", ) ``` ### Files | File | Description | |---|---| | `VTM-elf.pt-000060000` | DiT-30M weights (step 60k) | | `README.md` | This card | Runtime code for live camera → keypoints → inference is separate from this weight repo. --- ## Limitations - Beta: soft detail, identity drift, and pose errors are common - Small data + DiT-30M capacity constrain image quality - ~7 FPS on RTX 5060 Ti at 768 - Windows + CUDA only - **Framing:** torso-up only (roughly head → mid-torso). Legs / most of the waist are not supported - **Hands:** not supported - **Character types not supported:** realistic humans; non-humanoid / furries - **Accessories:** glasses and hats generally work; most other accessories are not supported --- ## Intended use Research and prototyping of VTuber pose → image pipelines (live drive, pose retarget experiments). Not a finished production renderer.