sinBoo1 commited on
Commit
2fc3ec1
·
verified ·
1 Parent(s): 1e56214

VTM-1.5.1.pt = the release checkpoint (10-step teacher, 1 or 2 steps); card: how it works, download chart, commercial-use column, license: other

Browse files
Files changed (4) hide show
  1. README.md +19 -11
  2. VTM-1.5.1.pt +2 -2
  3. media/model-downloads.png +0 -0
  4. media/pipeline.png +0 -0
README.md CHANGED
@@ -1,5 +1,7 @@
1
  ---
2
- license: apache-2.0
 
 
3
  language:
4
  - en
5
  pipeline_tag: image-to-image
@@ -19,7 +21,7 @@ tags:
19
  - cuda
20
  ---
21
 
22
- # VTM-Elf 0.01
23
 
24
  **Turn one anime picture into a live VTuber.** A keypoint-driven DiT draws your character in whatever pose and expression your face gives it, frame by frame.
25
 
@@ -40,6 +42,10 @@ These are the weights behind **[VTM Spark](https://github.com/sin-boo/VTM-Spark)
40
 
41
  <sub>Every frame on the right was drawn by `VTM-1.5.1.pt` from the one picture on the left. It was driven through VTM Spark's live tracking pipeline (head turn, nod and tilt, blinks, eye direction, mouth shapes), fed with scripted iPhone-style motion instead of a real face.</sub>
42
 
 
 
 
 
43
  ---
44
 
45
  ## Use it
@@ -77,7 +83,9 @@ VTM Spark ships this blueprint and a ready-to-paste image-AI prompt in `characte
77
  | `openseeface/*` | 21 MB | OpenSeeFace webcam face tracking models |
78
  | `media/*` | | Demo images for this card |
79
 
80
- VTM Spark places them under `models/dit/`, `models/trackers/` and `vendor/tools/openseeface/models/`.
 
 
81
 
82
  ---
83
 
@@ -144,14 +152,14 @@ Live VTubing and research on pose → image pipelines (live drive, pose retarget
144
 
145
  ## Licences
146
 
147
- Apache License 2.0 covers **our** training work: `VTM-1.5.1.pt` and our tracker fine-tunes. It does **not** re-license anyone else's weights, and some files here start from third-party weights with their own terms:
148
 
149
- | File | Licence |
150
- |---|---|
151
- | `VTM-1.5.1.pt` | Apache-2.0 (ours) |
152
- | `trackers/animeseg_hair3.pt` | Our fine-tune of Mask2Former ADE20k weights, which Meta licenses **CC BY-NC 4.0: non-commercial use only** |
153
- | `trackers/iris_pose.pt`, `trackers/dwpose_v2.pt` | Our fine-tunes of Ultralytics YOLO-pose pretrained weights, which Ultralytics licenses **AGPL-3.0** |
154
- | `trackers/pose_landmarker_lite.task` | Apache-2.0 (Google / MediaPipe) |
155
- | `openseeface/*` | BSD 2-Clause ([emilianavt/OpenSeeFace](https://github.com/emilianavt/OpenSeeFace)) |
156
 
157
  Full inventory and sources: [THIRD_PARTY_NOTICES.md](https://github.com/sin-boo/VTM-Spark/blob/main/THIRD_PARTY_NOTICES.md) in the VTM Spark repo.
 
1
  ---
2
+ license: other
3
+ license_name: vtm-spark-mixed
4
+ license_link: https://github.com/sin-boo/VTM-Spark/blob/main/THIRD_PARTY_NOTICES.md
5
  language:
6
  - en
7
  pipeline_tag: image-to-image
 
21
  - cuda
22
  ---
23
 
24
+ # VTM-1.5.1
25
 
26
  **Turn one anime picture into a live VTuber.** A keypoint-driven DiT draws your character in whatever pose and expression your face gives it, frame by frame.
27
 
 
42
 
43
  <sub>Every frame on the right was drawn by `VTM-1.5.1.pt` from the one picture on the left. It was driven through VTM Spark's live tracking pipeline (head turn, nod and tilt, blinks, eye direction, mouth shapes), fed with scripted iPhone-style motion instead of a real face.</sub>
44
 
45
+ ## How it works
46
+
47
+ <img src="https://huggingface.co/sinBoo1/VTM-Spark/resolve/main/media/pipeline.png" alt="Webcam or iPhone, then Track Lab, then 37 keypoints, then the VTM-1.5.1 DiT (fed with your character picture), then the SD VAE, then the VTM Spark camera" width="100%">
48
+
49
  ---
50
 
51
  ## Use it
 
83
  | `openseeface/*` | 21 MB | OpenSeeFace webcam face tracking models |
84
  | `media/*` | | Demo images for this card |
85
 
86
+ VTM Spark places them under `models/dit/`, `models/trackers/` and `vendor/tools/openseeface/models/`. On first setup it also fetches a few models straight from their publishers (SD VAE, a tiny VAE and two anime-face detectors), about 1.24 GB in all:
87
+
88
+ <img src="https://huggingface.co/sinBoo1/VTM-Spark/resolve/main/media/model-downloads.png" alt="Model downloads by size: hair segmentation 432 MB, generator 360 MB, SD VAE 335 MB, anime face landmarks 39 MB, body keypoints 23 MB, OpenSeeFace 21 MB, tiny VAE 9.8 MB, iris 6.4 MB, anime face box 6.2 MB, body tracking 5.8 MB" width="100%">
89
 
90
  ---
91
 
 
152
 
153
  ## Licences
154
 
155
+ No single licence covers every file here, so this repo is tagged `license: other`. Apache License 2.0 covers **our** training work: `VTM-1.5.1.pt` and our tracker fine-tunes. It does **not** re-license anyone else's weights, and some files here start from third-party weights with their own terms:
156
 
157
+ | File | Licence | Commercial use |
158
+ |---|---|---|
159
+ | `VTM-1.5.1.pt` | Apache-2.0 (ours) | Yes |
160
+ | `trackers/animeseg_hair3.pt` | Our fine-tune of Mask2Former ADE20k weights, which Meta licenses **CC BY-NC 4.0** | **No** |
161
+ | `trackers/iris_pose.pt`, `trackers/dwpose_v2.pt` | Our fine-tunes of Ultralytics YOLO-pose pretrained weights, which Ultralytics licenses **AGPL-3.0** | Under AGPL terms |
162
+ | `trackers/pose_landmarker_lite.task` | Apache-2.0 (Google / MediaPipe) | Yes |
163
+ | `openseeface/*` | BSD 2-Clause ([emilianavt/OpenSeeFace](https://github.com/emilianavt/OpenSeeFace)) | Yes |
164
 
165
  Full inventory and sources: [THIRD_PARTY_NOTICES.md](https://github.com/sin-boo/VTM-Spark/blob/main/THIRD_PARTY_NOTICES.md) in the VTM Spark repo.
VTM-1.5.1.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:25577fa38195bb3ee5eafba669ee2d9a2241aedd598d2b5be431bdaaca1ffe97
3
- size 359465662
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:11e9950405abe5af8ad305f6b38ba053a5794af316fbb004e0c3f5241c895836
3
+ size 359464372
media/model-downloads.png ADDED
media/pipeline.png ADDED