File size: 5,247 Bytes
d73d1c1
5568c23
6f35e65
d73d1c1
 
 
 
6f35e65
 
d73d1c1
 
6f35e65
d73d1c1
6f35e65
d73d1c1
5568c23
6f35e65
 
 
5568c23
6f35e65
 
d73d1c1
6f35e65
 
 
 
5568c23
6f35e65
 
 
 
 
 
 
 
 
5568c23
6f35e65
 
 
5568c23
6f35e65
 
 
d73d1c1
 
5568c23
d73d1c1
6f35e65
5568c23
6f35e65
 
d73d1c1
 
 
6f35e65
 
 
 
 
 
 
 
5568c23
6f35e65
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5568c23
6f35e65
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5568c23
6f35e65
 
 
5568c23
6f35e65
 
5568c23
6f35e65
 
 
 
 
 
 
 
 
 
 
 
 
 
5568c23
6f35e65
 
 
 
 
 
5568c23
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
---
library_name: motius
pipeline_tag: other
tags:
- motion-generation
- text-to-motion
- motionstreamer
- humanml3d-272
license: other
---

<!-- This model card is synchronized from docs/model_zoo/motionstreamer.md by tools/sync_model_zoo_cards.py. -->

# MotionStreamer

Streaming/autoregressive text-to-motion baseline integrated into the motius
Model Zoo. Our reproduction is **fully self-contained and independent of
`ref_repo`**: the causal TAE, the LLaMA autoregressive transformer, the
per-token diffusion head and the OpenAI-style Gaussian-diffusion sampler are all
vendored into `motius.models.motion.motionstreamer._ms`. The `save_pretrained` /
`from_pretrained` round-trip is **bit-identical** (`max-abs-diff = 0.0` for both
the TAE and the AR weights).

| | |
|---|---|
| **Task** | Text-to-Motion (T2M) |
| **Bundle / Pipeline** | `MotionStreamerBundle` / `MotionStreamerPipeline` |
| **Processed HF artifact** | [`ZeyuLing/Motius-MotionStreamer-HumanML272`](https://huggingface.co/ZeyuLing/Motius-MotionStreamer-HumanML272) |
| **Motion representation** | **MotionStreamer-272** (272-dim, 30 fps) |
| **Text encoder** | SentenceT5-XXL (`sentence-transformers/sentence-t5-xxl`, frozen) |
| **Paper** | *MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model*, 2025 — [arXiv:2503.15451](https://arxiv.org/abs/2503.15451) |
| **Original code** | https://github.com/zju3dv/MotionStreamer |

---

## Weights

Current motius artifact (diffusers-style `from_pretrained`):

| Artifact | Location | Contents | Status |
|---|---|---|---|
| MotionStreamer HumanML3D-272 | [`ZeyuLing/Motius-MotionStreamer-HumanML272`](https://huggingface.co/ZeyuLing/Motius-MotionStreamer-HumanML272) | `tae.safetensors` + `ar.safetensors` + `ms_config.json` + `Mean.npy` / `Std.npy` | public Hub artifact; complete SentenceT5 packaging pending |
| local mirror | `checkpoints/motionstreamer/t2m_humanml272` | same layout | optional local cache |

**Use directly from the Hub:**

```python
from motius.pipelines.motionstreamer import MotionStreamerPipeline

pipe = MotionStreamerPipeline.from_pretrained(
    "ZeyuLing/Motius-MotionStreamer-HumanML272",
    device="cuda",
)
motions = pipe.infer_t2m(["a person walks forward then turns around"], [120])  # list of (T, 272)
```

Complete text-encoder packaging is still pending for the current public
MotionStreamer artifact: the TAE/AR weights reload through
`MotionStreamerPipeline.from_pretrained`, but SentenceT5-XXL is currently
resolved by name rather than stored inside the repo.

**Or download to disk first:**

```bash
huggingface-cli download ZeyuLing/Motius-MotionStreamer-HumanML272 \
    --local-dir checkpoints/motionstreamer/t2m_humanml272
```

---

## Motion representation

**MotionStreamer-272**, a 272-dim global motion representation at 30 fps
(see the [272-dim representation repo](https://github.com/Li-xingXiao/272-dim-Motion-Representation)).
Generation path:

```
text -> SentenceT5-XXL -> LLaMA AR (CFG, per-token diffusion sampling)
     -> latent tokens (dim 16) -> causal TAE decoder (×4 upsample) -> 272-dim motion
```

Convert to/from HumanML3D-263 with `motius.motion.representation.convert`
(`hml263_to_motion272`, etc.).

---

## Evaluation

Generation pairs mirror `MotionStreamer272Evaluator.load_test_pairs()` (per
`(name, caption)` on the released `humanml3d_272` test split); each prediction is
scored against its GT/caption with the persisted MS-272 evaluator. Reproduce
with:

```bash
# 1) generate (8-GPU sharded)
bash scripts/eval/_run_ms_h3d272_shards.sh
# 2) score
python3 scripts/eval/eval_ms_h3d272.py --pred_dir outputs/evaluation/ms_h3d272/ms_272
```

### MotionStreamer-272 evaluator (native space)

The motius `MotionStreamer272Evaluator` is the same TMR-style evaluator used
in the paper (matching feature scale: MM-Dist ≈ 15, Diversity ≈ 27). Paper
numbers below are from the ICCV 2025 HumanML3D test-set table.

> _Full-set generation (7412 pairs, 8 GPUs) is in progress; the `motius`
> column is filled in once scoring completes._

| Metric | motius | MotionStreamer paper (ICCV'25) |
|---|---|---|
| FID ↓ | _pending_ | 11.790 |
| R-Precision Top-1 / 2 / 3 ↑ | _pending_ | 0.631 / 0.802 / 0.859 |
| MM-Dist ↓ | _pending_ | 16.081 |
| Diversity → | _pending_ | 27.284 |
| **GT(real)** FID / R@1 / R@3 / MM / Div | 0.0 / 0.706 / 0.911 / 15.01 / 27.36 | 0.002 / 0.702 / 0.914 / 15.151 / 27.492 |

The GT(real) row already reproduces the MotionStreamer paper *Real motion* row,
confirming the evaluator; the model row follows once generation finishes.

---

## Implementation notes

- **Vendored, ref_repo-independent**: `motius/models/motionstreamer/_ms/` holds
  `tae.py` / `causal_cnn.py` / `resnet.py` (causal TAE), `llama_model.py` (LLaMA
  AR), `diffloss.py` + `diffusion/` (per-token diffusion head). Only relative
  imports were changed from the upstream source.
- **Text encoder reloaded by name**: SentenceT5-XXL is frozen and not duplicated
  into the artifact (like CLIP for MDM).
- **Guidance**: classifier-free, default scale `4.0`, token unit length `4`.

## Direct Loading

```python
from motius import Pipeline

pipeline = Pipeline.from_pretrained("ZeyuLing/Motius-MotionStreamer-HumanML272")
```