File size: 14,226 Bytes
da12144
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c6c36dc
da12144
fd9f18a
 
da12144
 
 
 
 
 
 
 
 
 
4abd650
da12144
 
fd9f18a
da12144
4abd650
 
 
 
 
5b6e273
 
 
da12144
 
 
 
 
c6c36dc
 
da12144
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c6c36dc
 
 
 
 
 
 
 
 
 
fd9f18a
c6c36dc
 
 
 
 
 
 
 
fd9f18a
 
 
 
 
 
 
 
 
c6c36dc
 
 
 
 
 
da12144
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c6c36dc
 
5b6e273
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fd9f18a
da12144
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5b6e273
 
c6c36dc
da12144
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8584083
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
---
license: other
library_name: motius
tags:
  - multimodal-generation
  - motion-generation
  - music-generation
  - motion-captioning
  - music-captioning
  - humanml3d
  - aistplusplus
  - unimumo
---

<h1 align="center">UniMuMo Model Card</h1>

<p align="center">
  <strong>One checkpoint for generation and translation across text, music, and motion.</strong>
</p>

<p align="center">
  <a href="https://arxiv.org/abs/2410.04534">Paper</a> |
  <a href="https://hanyangclarence.github.io/unimumo_demo/">Project Page</a> |
  <a href="https://github.com/hanyangclarence/UniMuMo">Original GitHub</a> |
  <a href="https://huggingface.co/ClarenceY/unimumo">Original Checkpoint</a> |
  <a href="https://huggingface.co/ZeyuLing/Motius-UniMuMo">Motius Checkpoint</a> |
  <a href="https://huggingface.co/spaces/ZeyuLing/t2m-humanml3d-leaderboard">T2M Leaderboard</a> |
  <a href="https://huggingface.co/spaces/ZeyuLing/m2t-humanml3d-leaderboard">M2T Leaderboard</a> |
  <a href="https://huggingface.co/spaces/ZeyuLing/music-to-dance-aistpp-leaderboard">M2D Leaderboard</a> |
  <a href="https://huggingface.co/spaces/ZeyuLing/dance-to-music-aistpp-leaderboard">D2M Leaderboard</a>
</p>

UniMuMo is the unified text, music, and motion model introduced in *UniMuMo:
Unified Text, Music and Motion Generation*. Motius independently implements its
inference architecture and converts the authors' published weights into one
self-contained safe artifact. Runtime inference does not import an upstream
checkout or download a second text, audio, motion, or caption model.

## Preview

- [HumanML3D all-case text-to-motion comparison](https://zeyuling-t2m-humanml3d-leaderboard.static.hf.space/cases/index.html)
- [HumanML3D all-case motion caption comparison](https://zeyuling-m2t-humanml3d-leaderboard.static.hf.space/cases/index.html)
- [Audio-synchronized AIST++ all-case dance comparison](https://zeyuling-music-to-dance-aistpp-leaderboard.static.hf.space/cases/index.html)
- [Motion-synchronized AIST++ all-case music comparison](https://zeyuling-dance-to-music-aistpp-leaderboard.static.hf.space/cases/index.html)

The T2M page compares UniMuMo's generated SMPL Mesh with every released
baseline over all 4,042 selected-caption cases. The M2T page shows the same
animated input SMPL Mesh beside every baseline caption for all 4,400 protocol
samples. The M2D page contains all 40 AIST++ cases, synchronized audio, native
SMPL-24 joints, the fitted SMPL Mesh, orbit, zoom, timeline seeking, and
downloadable motion assets. The D2M page contains the official 86 two-second
D2M-GAN clips and lets the viewer switch between reference and generated audio
while the same input mesh plays.

## Release Snapshot

| Item | Value |
| ---- | ----- |
| Tasks | T2M, M2T, Music-to-Dance, Dance-to-Music |
| Additional pipeline routes | Text-to-Music and joint Text-to-Music-Motion |
| Motion representation | HumanML3D-263 at 60 fps |
| Audio representation | Encodec, 32 kHz, four 2,048-entry RVQ codebooks |
| Shared code rate | 50 Hz |
| Generator | 24-layer, 1,024D dual-stream autoregressive Transformer |
| Text conditioning and captioning | T5-base encoder and T5-base captioner |
| Maximum duration | 10 seconds per call |
| Checkpoint | [`ZeyuLing/Motius-UniMuMo`](https://huggingface.co/ZeyuLing/Motius-UniMuMo) |
| Pipeline | `motius.pipelines.unimumo.UniMuMoPipeline` |
| Upstream revision | `hanyangclarence/UniMuMo@a75ddac791ff6806b5bd511d1ce887a1980e20d5` |

The artifact includes both core shards, Encodec, T5 encoder, T5 captioner,
SentencePiece tokenizer, HumanML3D normalization statistics, configuration,
and provenance. `UniMuMoPipeline.from_pretrained` is the only loader needed.

## Usage

Install the UniMuMo dependencies and load the complete artifact:

```bash
python -m pip install -e '.[unimumo]'
```

```python
from motius.pipelines.unimumo import UniMuMoPipeline

pipe = UniMuMoPipeline.from_pretrained(
    "ZeyuLing/Motius-UniMuMo",
    device="cuda",
)
```

Generate synchronized music and motion from two optional text descriptions:

```python
result = pipe.infer_text_to_music_motion(
    music_prompt="an upbeat electronic dance track",
    motion_prompt="a person dances energetically",
    duration_seconds=8.0,
    guidance_scale=4.0,
    seed=7,
)

print(result.waveform.shape, result.sample_rate)  # (256000,), 32000
print(result.motion.shape, result.motion_fps)     # (480, 263), 60.0
print(result.joints.shape)                       # (480, 22, 3)
```

Use the task-specific routes with the same loaded pipeline:

```python
t2m = pipe.infer_text_to_motion(
    "a person walks in a circle",
    duration_seconds=6.0,
    seed=7,
)
t2music = pipe.infer_text_to_music(
    "a quiet piano melody",
    duration_seconds=6.0,
    seed=7,
)
m2d = pipe.infer_music_to_motion(
    "music.wav",
    motion_prompt="a person performs a street dance",
    guidance_scale=3.0,
    seed=7,
)
motion_music = pipe.infer_motion_to_music(
    t2m.motion,
    input_fps=60.0,
    music_prompt="upbeat percussion",
    seed=7,
)
motion_caption = pipe.infer_motion_to_text(t2m.motion, input_fps=60.0)
music_caption = pipe.infer_music_to_text("music.wav")
```

Array audio inputs also require `sample_rate=...`. Motion inputs must be
HumanML3D-263; other Motius representations should first be converted through
the motion representation API.

## Evaluation Results

### HumanML3D Text-to-Motion

UniMuMo exposes text-to-motion as a zero-shot route of its joint music-motion
generator. Motius evaluates all 4,042 selected-caption protocol cases with one
deterministic pass and retrieval groups of 32. HumanML3D Official uses the
4,012 cases for which the released `new_joint_vecs` reference exists; its
retrieval computation uses the largest complete 32-case groups (`n=4,000`).

| Evaluator | n | R@1 | R@2 | R@3 | FID | MM-Dist | Diversity |
| --------- | -: | --: | --: | --: | --: | ------: | --------: |
| HumanML3D Official | 4,000 | 0.1000 | 0.1775 | 0.2468 | 1.4849 | 6.6372 | 9.0766 |
| MotionStreamer Evaluator | 4,032 | 0.0655 | 0.1138 | 0.1617 | 373.2192 | 25.6637 | 18.8368 |
| Motius Joint-Position Evaluator | 4,032 | 0.0704 | 0.1471 | 0.2093 | 0.6788 | 54.0101 | 46.7609 |

HumanML3D and MotionStreamer FID use their native evaluator spaces; uTMR FID
uses per-sample L2-normalized embeddings. The weak retrieval scores are
reported as measured: this checkpoint supports T2M, but it was not optimized
as a dedicated HumanML3D text-to-motion model.

The HumanML3D result above is computed from the codec's native 60 fps output
with the phase-aligned `[1::3]` inverse used by the official UniMuMo data
pipeline. A parity audit found and fixed a top-k sampling-order discrepancy;
under the upstream dependency versions, motion codes, sampled tokens, and
decoded features now match the released implementation exactly. The UniMuMo
paper does not report a standalone HumanML3D T2M leaderboard result, so the
remaining low retrieval score is recorded as the zero-shot operating point of
the released joint model rather than treated as a paper-parity target.

Physical diagnostics over all 4,042 generated SMPL-22 joint sequences:

| Slide | Float | Jitter | Dynamic | Penetration |
| ----: | ----: | -----: | ------: | ----------: |
| 21.7377 | 58.7669 | 11.8780 | 54.3552 | 0.0000 |

### HumanML3D Motion-to-Text

The shared M2T protocol contains 4,400 official test motions and temporal
subclips, three references per sample, semantic retrieval groups of 32, and
one deterministic evaluation pass. UniMuMo follows the authors' 10-second
captioning protocol: 20 fps HumanML3D input is padded to 200 frames, linearly
resampled to 60 fps, encoded, and captioned.

| Method | BLEU-1 | BLEU-4 | ROUGE-L | CIDEr | BERT raw | R@1 | R@2 | R@3 | MM-Dist |
| ------ | -----: | -----: | ------: | ----: | -------: | --: | --: | --: | ------: |
| UniMuMo | 0.3534 | 0.0457 | 0.2822 | 0.0635 | 0.9006 | 0.5162 | 0.7032 | 0.7984 | 2.9658 |

The lexical metrics use the TM2T token/lemma references. BERT raw is the
unrescaled RoBERTa-large layer-17 cosine score; the corresponding
baseline-rescaled BERTScore is `0.4109`. The paper reports `R@1=0.520`,
`R@3=0.806`, and `MM-Dist=2.958`, which closely matches this independent run.
The leaderboard's raw-reference diagnostic reports BLEU-4 `0.1271`, ROUGE-L
`0.3560`, CIDEr `0.3114`, and raw BERTScore `0.9075`.

### AIST++ Music-to-Dance

The common leaderboard evaluates all 40 public cross-modal cases against the
complete 1,320-motion AIST++ reference pool. FID_k/FID_g and diversity use the
released Bailando 60 fps protocol. uTMR FID uses canonical SMPL-22 joints at
30 fps with per-sample L2-normalized embeddings.

| Result | FID_k | FID_g | uTMR FID | Div_k | Div_g | BeatAlign |
| ------ | ----: | ----: | --------: | ----: | ----: | --------: |
| UniMuMo | 17.7250 | 38.6446 | 0.2823 | 8.8767 | 8.4657 | 0.2430 |
| Motius GT | 17.1589 | 10.6618 | 0.1829 | 8.1666 | 7.4893 | 0.2247 |

Physical diagnostics on the same generated clips:

| Jitter | Dynamic | Penetration | Float | Slide |
| -----: | ------: | ----------: | ----: | ----: |
| 0.00982 | 0.02523 | 0.00000 | 0.19176 | 0.00523 |

For paper parity, a separate first-five-second evaluation gives
`FID_k=10.7721`, `FID_g=27.3115`, and `BeatAlign=0.2430`; the paper reports
BeatAlign `0.24`. The leaderboard uses full generated timelines for every
method and does not mix the shorter parity result into rankings.

### AIST++ Dance-to-Music

Motion-to-music is a zero-shot route of the same released checkpoint. Motius
uses the D2M-GAN list of 86 two-second AIST++ clips, no text prompt, `CFG=3`,
temperature `1.0`, top-k `250`, and the released onset detector and one-second
beat-bin aggregation.

| Protocol | Samples | Beat Count Ratio (target 100%) | Beat Hit |
| -------- | ------: | -----------------------------: | -------: |
| UniMuMo paper, D2M-GAN protocol | 86 | 93.0% | 88.4% |
| Motius reproduction, official checkpoint | 86 | 84.30% | 80.81% |

The paper calls the first metric *Beats Coverage*, but its released evaluator
defines it as generated beat bins divided by reference beat bins. It is
therefore unbounded: values above 100% mean excess detected beats, not a score
better than 100%. Motius displays the less ambiguous name **Beat Count Ratio**
and does not rank it with a higher-is-better arrow.

An earlier 40-case result was invalid and has been removed. The published
audio codec stores modern parametrized weight-normalization keys, while the
official runtime expects `weight_g` and `weight_v`; the permissive loader had
silently left convolution weights randomly initialized. The Motius loader now
maps the layouts explicitly and rejects every missing, unexpected, or
shape-mismatched tensor. After the fix, motion codes, music codes, conditioned
motion codes, and decoded waveforms match the official implementation exactly
(waveform RMSE and maximum absolute error are both zero).

The public UniMuMo motion archive contains exact HumanML3D-263 features for 13
of the 20 AIST++ source dances in this protocol. The seven filtered dances are
reconstructed from the official AIST++ SMPL release using the same public SMPL
web rig and HumanML3D preprocessing. The six 119-frame tail clips retain their
native length instead of being padded to 120 frames.

Every generated WAV, codec output, input-motion provenance record, seed, and
per-case metric is available in the
[`dance_to_music_d2mgan_official86` benchmark folder](https://huggingface.co/ZeyuLing/Motius-UniMuMo/tree/main/benchmarks/dance_to_music_d2mgan_official86).

[Open the Dance-to-Music leaderboard and synchronized 86-case SMPL/audio viewer](https://huggingface.co/spaces/ZeyuLing/dance-to-music-aistpp-leaderboard).

## Motion Representation

The native motion stream is HumanML3D-263. Its root velocities, root height,
root-invariant positions, continuous 6D local rotations, local velocities, and
foot contacts are normalized with the authors' published statistics and
encoded jointly with zero-audio embeddings. The motion codec maps 60 fps motion
to the same 50 Hz, four-codebook token clock used by music.

For the AIST++ viewer and evaluator, Motius decodes HumanML3D joints, converts
the common SMPL-22 body directly, and extrapolates AIST++ hand joints 22 and 23
from each elbow-to-wrist direction by `0.35x`. The official feature evaluator
uses 24 joints; uTMR uses only the common SMPL-22 body. SMPL Mesh preview uses
position IK because generated joint positions do not uniquely determine axial
twist. Across all 40 clips, the fitted mesh has `28.16 mm` mean joint MPJPE.

## Reproduction Audit

| Check | Result |
| ----- | ------ |
| Motion-code encoding vs official implementation | Exact equality, zero differing codes |
| Seeded music generation vs official implementation | Exact equality, zero differing codes |
| Seeded motion generation vs official implementation | Exact equality, zero differing codes |
| Motion caption vs official implementation | Exact string equality |
| Official tensor load | 880 core tensors loaded across two safe shards |
| HumanML3D caption generation | 4,400/4,400 protocol samples |
| AIST++ dance generation | 40/40 cases, all finite |
| AIST++ music generation from dance | 86/86 official D2M-GAN clips, all WAV and codec metadata public |
| Encodec compatibility parity | Exact codes and decoded waveform; RMSE 0, maximum absolute error 0 |
| HumanML3D zero-shot motion generation | 4,042/4,042 selected-caption cases, all four output representations finite |
| Runtime boundary | No import from `ref_repo` or an upstream checkout |

The audited upstream source and checkpoint declare no license. Motius records
that fact rather than inferring redistribution terms. Users remain responsible
for obtaining permission appropriate to their use case.

## Citation

```bibtex
@article{yang2024unimumo,
  title={UniMuMo: Unified Text, Music and Motion Generation},
  author={Yang, Han and Su, Kun and Zhang, Yutong and Chen, Jiaben and Qian, Kaizhi and Liu, Gaowen and Gan, Chuang},
  journal={arXiv preprint arXiv:2410.04534},
  year={2024}
}
```

## Direct Loading

```python
from motius import Pipeline

pipeline = Pipeline.from_pretrained("ZeyuLing/Motius-UniMuMo")
```