File size: 2,705 Bytes
458c218 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 | ---
license: apache-2.0
tags:
- motion-generation
- human-pose
- music-video
- transformer
- boostify
- motiondna
library_name: pytorch
pipeline_tag: other
---
# Boostify MotionDNA v0
**MotionDNA** is Boostify's human-motion model: a compact autoregressive Transformer
that learns the *DNA* of how performers move in real music videos and can generate
new, plausible motion sequences (2D skeletal pose) from a short seed.
This is the **v0** checkpoint, trained end-to-end on motion extracted from real
music-video clips. It is intentionally lightweight so it runs (and trains) on a
single NVIDIA L4 (23 GB) — in fact training uses under 1 GB of VRAM.
## What it does
- Input: a sequence of 2D human poses (17 COCO keypoints, x/y).
- Output: the predicted next pose, applied autoregressively to roll out new motion.
- Use cases: motion previews, choreography ideation, b-roll motion guides,
music-video editing assistance inside Boostify Dataset Studio.
## Architecture
- Causal Transformer encoder (next-pose prediction).
- `d_model=256`, `n_head=8`, `n_layer=6`, window `WIN=64`, input dim `34` (17×2).
- ~4.77M parameters. Smooth-L1 loss. AdamW + cosine schedule.
## Training data
Pose sequences extracted with YOLOv8-pose from real music-video clips
(46 clips, 11,038 frames). Poses are hip-centered and torso-scaled, then
standardized (mean/std stored inside the checkpoint).
> Source footage is used internally for research/demo. The released artifact
> contains only the trained weights and normalization stats — no source video.
## Files
| File | Description |
|------|-------------|
| `motiondna_v0_best.pt` | Best checkpoint (lowest train loss). |
| `motiondna_v0_final.pt` | Final-epoch checkpoint. |
| `motiondna_sample.npy` | Example generated motion `[136, 17, 2]`. |
| `loss_history.json` | Per-epoch training loss. |
| `extract_poses.py` | Pose-extraction pipeline (YOLOv8-pose → .npy). |
| `train_motiondna.py` | Training script. |
## Inference
```python
import torch, numpy as np
ckpt = torch.load("motiondna_v0_best.pt", map_location="cpu")
cfg = ckpt["cfg"]; mean = ckpt["mean"]; std = ckpt["std"]
# rebuild the MotionDNA module from train_motiondna.py, load_state_dict(ckpt["model"]),
# seed with N normalized poses and roll out model(x)[:, -1:] autoregressively,
# then de-normalize with: pose = gen * std + mean
```
## Limitations
- v0 trained on a small clip set → motion is coherent but limited in style range.
- 2D pose only (no 3D, no global translation, no audio conditioning yet).
- Roadmap: more clips, music/beat conditioning, 3D lift, style control.
## Citation
Boostify MotionDNA (2026). Trained with the Boostify Dataset Studio pipeline.
|