VTM-1.5.1.pt = the release checkpoint (10-step teacher, 1 or 2 steps); card: how it works, download chart, commercial-use column, license: other
Browse files- README.md +19 -11
- VTM-1.5.1.pt +2 -2
- media/model-downloads.png +0 -0
- media/pipeline.png +0 -0
README.md
CHANGED
|
@@ -1,5 +1,7 @@
|
|
| 1 |
---
|
| 2 |
-
license:
|
|
|
|
|
|
|
| 3 |
language:
|
| 4 |
- en
|
| 5 |
pipeline_tag: image-to-image
|
|
@@ -19,7 +21,7 @@ tags:
|
|
| 19 |
- cuda
|
| 20 |
---
|
| 21 |
|
| 22 |
-
# VTM-
|
| 23 |
|
| 24 |
**Turn one anime picture into a live VTuber.** A keypoint-driven DiT draws your character in whatever pose and expression your face gives it, frame by frame.
|
| 25 |
|
|
@@ -40,6 +42,10 @@ These are the weights behind **[VTM Spark](https://github.com/sin-boo/VTM-Spark)
|
|
| 40 |
|
| 41 |
<sub>Every frame on the right was drawn by `VTM-1.5.1.pt` from the one picture on the left. It was driven through VTM Spark's live tracking pipeline (head turn, nod and tilt, blinks, eye direction, mouth shapes), fed with scripted iPhone-style motion instead of a real face.</sub>
|
| 42 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 43 |
---
|
| 44 |
|
| 45 |
## Use it
|
|
@@ -77,7 +83,9 @@ VTM Spark ships this blueprint and a ready-to-paste image-AI prompt in `characte
|
|
| 77 |
| `openseeface/*` | 21 MB | OpenSeeFace webcam face tracking models |
|
| 78 |
| `media/*` | | Demo images for this card |
|
| 79 |
|
| 80 |
-
VTM Spark places them under `models/dit/`, `models/trackers/` and `vendor/tools/openseeface/models/`.
|
|
|
|
|
|
|
| 81 |
|
| 82 |
---
|
| 83 |
|
|
@@ -144,14 +152,14 @@ Live VTubing and research on pose → image pipelines (live drive, pose retarget
|
|
| 144 |
|
| 145 |
## Licences
|
| 146 |
|
| 147 |
-
Apache License 2.0 covers **our** training work: `VTM-1.5.1.pt` and our tracker fine-tunes. It does **not** re-license anyone else's weights, and some files here start from third-party weights with their own terms:
|
| 148 |
|
| 149 |
-
| File | Licence |
|
| 150 |
-
|---|---|
|
| 151 |
-
| `VTM-1.5.1.pt` | Apache-2.0 (ours) |
|
| 152 |
-
| `trackers/animeseg_hair3.pt` | Our fine-tune of Mask2Former ADE20k weights, which Meta licenses **CC BY-NC 4.0
|
| 153 |
-
| `trackers/iris_pose.pt`, `trackers/dwpose_v2.pt` | Our fine-tunes of Ultralytics YOLO-pose pretrained weights, which Ultralytics licenses **AGPL-3.0** |
|
| 154 |
-
| `trackers/pose_landmarker_lite.task` | Apache-2.0 (Google / MediaPipe) |
|
| 155 |
-
| `openseeface/*` | BSD 2-Clause ([emilianavt/OpenSeeFace](https://github.com/emilianavt/OpenSeeFace)) |
|
| 156 |
|
| 157 |
Full inventory and sources: [THIRD_PARTY_NOTICES.md](https://github.com/sin-boo/VTM-Spark/blob/main/THIRD_PARTY_NOTICES.md) in the VTM Spark repo.
|
|
|
|
| 1 |
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: vtm-spark-mixed
|
| 4 |
+
license_link: https://github.com/sin-boo/VTM-Spark/blob/main/THIRD_PARTY_NOTICES.md
|
| 5 |
language:
|
| 6 |
- en
|
| 7 |
pipeline_tag: image-to-image
|
|
|
|
| 21 |
- cuda
|
| 22 |
---
|
| 23 |
|
| 24 |
+
# VTM-1.5.1
|
| 25 |
|
| 26 |
**Turn one anime picture into a live VTuber.** A keypoint-driven DiT draws your character in whatever pose and expression your face gives it, frame by frame.
|
| 27 |
|
|
|
|
| 42 |
|
| 43 |
<sub>Every frame on the right was drawn by `VTM-1.5.1.pt` from the one picture on the left. It was driven through VTM Spark's live tracking pipeline (head turn, nod and tilt, blinks, eye direction, mouth shapes), fed with scripted iPhone-style motion instead of a real face.</sub>
|
| 44 |
|
| 45 |
+
## How it works
|
| 46 |
+
|
| 47 |
+
<img src="https://huggingface.co/sinBoo1/VTM-Spark/resolve/main/media/pipeline.png" alt="Webcam or iPhone, then Track Lab, then 37 keypoints, then the VTM-1.5.1 DiT (fed with your character picture), then the SD VAE, then the VTM Spark camera" width="100%">
|
| 48 |
+
|
| 49 |
---
|
| 50 |
|
| 51 |
## Use it
|
|
|
|
| 83 |
| `openseeface/*` | 21 MB | OpenSeeFace webcam face tracking models |
|
| 84 |
| `media/*` | | Demo images for this card |
|
| 85 |
|
| 86 |
+
VTM Spark places them under `models/dit/`, `models/trackers/` and `vendor/tools/openseeface/models/`. On first setup it also fetches a few models straight from their publishers (SD VAE, a tiny VAE and two anime-face detectors), about 1.24 GB in all:
|
| 87 |
+
|
| 88 |
+
<img src="https://huggingface.co/sinBoo1/VTM-Spark/resolve/main/media/model-downloads.png" alt="Model downloads by size: hair segmentation 432 MB, generator 360 MB, SD VAE 335 MB, anime face landmarks 39 MB, body keypoints 23 MB, OpenSeeFace 21 MB, tiny VAE 9.8 MB, iris 6.4 MB, anime face box 6.2 MB, body tracking 5.8 MB" width="100%">
|
| 89 |
|
| 90 |
---
|
| 91 |
|
|
|
|
| 152 |
|
| 153 |
## Licences
|
| 154 |
|
| 155 |
+
No single licence covers every file here, so this repo is tagged `license: other`. Apache License 2.0 covers **our** training work: `VTM-1.5.1.pt` and our tracker fine-tunes. It does **not** re-license anyone else's weights, and some files here start from third-party weights with their own terms:
|
| 156 |
|
| 157 |
+
| File | Licence | Commercial use |
|
| 158 |
+
|---|---|---|
|
| 159 |
+
| `VTM-1.5.1.pt` | Apache-2.0 (ours) | Yes |
|
| 160 |
+
| `trackers/animeseg_hair3.pt` | Our fine-tune of Mask2Former ADE20k weights, which Meta licenses **CC BY-NC 4.0** | **No** |
|
| 161 |
+
| `trackers/iris_pose.pt`, `trackers/dwpose_v2.pt` | Our fine-tunes of Ultralytics YOLO-pose pretrained weights, which Ultralytics licenses **AGPL-3.0** | Under AGPL terms |
|
| 162 |
+
| `trackers/pose_landmarker_lite.task` | Apache-2.0 (Google / MediaPipe) | Yes |
|
| 163 |
+
| `openseeface/*` | BSD 2-Clause ([emilianavt/OpenSeeFace](https://github.com/emilianavt/OpenSeeFace)) | Yes |
|
| 164 |
|
| 165 |
Full inventory and sources: [THIRD_PARTY_NOTICES.md](https://github.com/sin-boo/VTM-Spark/blob/main/THIRD_PARTY_NOTICES.md) in the VTM Spark repo.
|
VTM-1.5.1.pt
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:11e9950405abe5af8ad305f6b38ba053a5794af316fbb004e0c3f5241c895836
|
| 3 |
+
size 359464372
|
media/model-downloads.png
ADDED
|
media/pipeline.png
ADDED
|