File size: 5,247 Bytes
d73d1c1 5568c23 6f35e65 d73d1c1 6f35e65 d73d1c1 6f35e65 d73d1c1 6f35e65 d73d1c1 5568c23 6f35e65 5568c23 6f35e65 d73d1c1 6f35e65 5568c23 6f35e65 5568c23 6f35e65 5568c23 6f35e65 d73d1c1 5568c23 d73d1c1 6f35e65 5568c23 6f35e65 d73d1c1 6f35e65 5568c23 6f35e65 5568c23 6f35e65 5568c23 6f35e65 5568c23 6f35e65 5568c23 6f35e65 5568c23 6f35e65 5568c23 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 | ---
library_name: motius
pipeline_tag: other
tags:
- motion-generation
- text-to-motion
- motionstreamer
- humanml3d-272
license: other
---
<!-- This model card is synchronized from docs/model_zoo/motionstreamer.md by tools/sync_model_zoo_cards.py. -->
# MotionStreamer
Streaming/autoregressive text-to-motion baseline integrated into the motius
Model Zoo. Our reproduction is **fully self-contained and independent of
`ref_repo`**: the causal TAE, the LLaMA autoregressive transformer, the
per-token diffusion head and the OpenAI-style Gaussian-diffusion sampler are all
vendored into `motius.models.motion.motionstreamer._ms`. The `save_pretrained` /
`from_pretrained` round-trip is **bit-identical** (`max-abs-diff = 0.0` for both
the TAE and the AR weights).
| | |
|---|---|
| **Task** | Text-to-Motion (T2M) |
| **Bundle / Pipeline** | `MotionStreamerBundle` / `MotionStreamerPipeline` |
| **Processed HF artifact** | [`ZeyuLing/Motius-MotionStreamer-HumanML272`](https://huggingface.co/ZeyuLing/Motius-MotionStreamer-HumanML272) |
| **Motion representation** | **MotionStreamer-272** (272-dim, 30 fps) |
| **Text encoder** | SentenceT5-XXL (`sentence-transformers/sentence-t5-xxl`, frozen) |
| **Paper** | *MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model*, 2025 — [arXiv:2503.15451](https://arxiv.org/abs/2503.15451) |
| **Original code** | https://github.com/zju3dv/MotionStreamer |
---
## Weights
Current motius artifact (diffusers-style `from_pretrained`):
| Artifact | Location | Contents | Status |
|---|---|---|---|
| MotionStreamer HumanML3D-272 | [`ZeyuLing/Motius-MotionStreamer-HumanML272`](https://huggingface.co/ZeyuLing/Motius-MotionStreamer-HumanML272) | `tae.safetensors` + `ar.safetensors` + `ms_config.json` + `Mean.npy` / `Std.npy` | public Hub artifact; complete SentenceT5 packaging pending |
| local mirror | `checkpoints/motionstreamer/t2m_humanml272` | same layout | optional local cache |
**Use directly from the Hub:**
```python
from motius.pipelines.motionstreamer import MotionStreamerPipeline
pipe = MotionStreamerPipeline.from_pretrained(
"ZeyuLing/Motius-MotionStreamer-HumanML272",
device="cuda",
)
motions = pipe.infer_t2m(["a person walks forward then turns around"], [120]) # list of (T, 272)
```
Complete text-encoder packaging is still pending for the current public
MotionStreamer artifact: the TAE/AR weights reload through
`MotionStreamerPipeline.from_pretrained`, but SentenceT5-XXL is currently
resolved by name rather than stored inside the repo.
**Or download to disk first:**
```bash
huggingface-cli download ZeyuLing/Motius-MotionStreamer-HumanML272 \
--local-dir checkpoints/motionstreamer/t2m_humanml272
```
---
## Motion representation
**MotionStreamer-272**, a 272-dim global motion representation at 30 fps
(see the [272-dim representation repo](https://github.com/Li-xingXiao/272-dim-Motion-Representation)).
Generation path:
```
text -> SentenceT5-XXL -> LLaMA AR (CFG, per-token diffusion sampling)
-> latent tokens (dim 16) -> causal TAE decoder (×4 upsample) -> 272-dim motion
```
Convert to/from HumanML3D-263 with `motius.motion.representation.convert`
(`hml263_to_motion272`, etc.).
---
## Evaluation
Generation pairs mirror `MotionStreamer272Evaluator.load_test_pairs()` (per
`(name, caption)` on the released `humanml3d_272` test split); each prediction is
scored against its GT/caption with the persisted MS-272 evaluator. Reproduce
with:
```bash
# 1) generate (8-GPU sharded)
bash scripts/eval/_run_ms_h3d272_shards.sh
# 2) score
python3 scripts/eval/eval_ms_h3d272.py --pred_dir outputs/evaluation/ms_h3d272/ms_272
```
### MotionStreamer-272 evaluator (native space)
The motius `MotionStreamer272Evaluator` is the same TMR-style evaluator used
in the paper (matching feature scale: MM-Dist ≈ 15, Diversity ≈ 27). Paper
numbers below are from the ICCV 2025 HumanML3D test-set table.
> _Full-set generation (7412 pairs, 8 GPUs) is in progress; the `motius`
> column is filled in once scoring completes._
| Metric | motius | MotionStreamer paper (ICCV'25) |
|---|---|---|
| FID ↓ | _pending_ | 11.790 |
| R-Precision Top-1 / 2 / 3 ↑ | _pending_ | 0.631 / 0.802 / 0.859 |
| MM-Dist ↓ | _pending_ | 16.081 |
| Diversity → | _pending_ | 27.284 |
| **GT(real)** FID / R@1 / R@3 / MM / Div | 0.0 / 0.706 / 0.911 / 15.01 / 27.36 | 0.002 / 0.702 / 0.914 / 15.151 / 27.492 |
The GT(real) row already reproduces the MotionStreamer paper *Real motion* row,
confirming the evaluator; the model row follows once generation finishes.
---
## Implementation notes
- **Vendored, ref_repo-independent**: `motius/models/motionstreamer/_ms/` holds
`tae.py` / `causal_cnn.py` / `resnet.py` (causal TAE), `llama_model.py` (LLaMA
AR), `diffloss.py` + `diffusion/` (per-token diffusion head). Only relative
imports were changed from the upstream source.
- **Text encoder reloaded by name**: SentenceT5-XXL is frozen and not duplicated
into the artifact (like CLIP for MDM).
- **Guidance**: classifier-free, default scale `4.0`, token unit length `4`.
## Direct Loading
```python
from motius import Pipeline
pipeline = Pipeline.from_pretrained("ZeyuLing/Motius-MotionStreamer-HumanML272")
```
|