Text Generation
Transformers
Safetensors
English
mugen
motion
human-motion
text-to-motion
motion-to-text
motion-captioning
motion-generation
autoencoder
gpt2
custom_code
Instructions to use zy22b/MUGEN with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use zy22b/MUGEN with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="zy22b/MUGEN", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("zy22b/MUGEN", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use zy22b/MUGEN with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "zy22b/MUGEN" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zy22b/MUGEN", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/zy22b/MUGEN
- SGLang
How to use zy22b/MUGEN with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "zy22b/MUGEN" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zy22b/MUGEN", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "zy22b/MUGEN" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zy22b/MUGEN", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use zy22b/MUGEN with Docker Model Runner:
docker model run hf.co/zy22b/MUGEN
| license: mit | |
| language: | |
| - en | |
| tags: | |
| - motion | |
| - human-motion | |
| - text-to-motion | |
| - motion-to-text | |
| - motion-captioning | |
| - motion-generation | |
| - autoencoder | |
| - gpt2 | |
| datasets: | |
| - humanml3d | |
| pipeline_tag: text-generation | |
| library_name: transformers | |
| base_model: | |
| - openai-community/gpt2 | |
| # MUGEN | |
| **A Unified Framework for Efficient Motion Understanding and Generation** | |
| MUGEN is a unified motion-language model. One model turns a description into | |
| human motion, and describes an observed motion in words. Both directions share a | |
| single motion representation, and that representation is continuous: **no | |
| codebook, one draw.** | |
| - 📦 Code: [github.com/JYe16/MUGEN](https://github.com/JYe16/MUGEN) | |
| - 🏋️ This checkpoint: HumanML3D, **K = 2** latent slots, 183.7M parameters | |
| --- | |
| ## What makes it different | |
| Unified motion-language systems have traditionally coupled generation and | |
| understanding through a shared *discrete* motion codebook, but quantization | |
| limits generation quality. The strongest generators buy that quality back at | |
| growing cost: stacked residual codebooks enlarge the representation, and masked | |
| decoding stages, long autoregressive rollouts and denoising chains of tens to | |
| hundreds of steps stretch inference. None of that decoding machinery serves | |
| understanding. | |
| MUGEN pays neither cost. Generating a motion costs **K language-model steps, one | |
| draw, and one decoder pass** — with K = 2 here, that is two rollout steps. | |
| ## Architecture | |
| ``` | |
| MugenForConditionalGeneration 183,690,603 params | |
| ├── motion_autoencoder (AdaptiveLengthAutoEncoder) 48.9M | |
| │ ├── encoder_backbone dilated Conv1d ResNet, per-frame features | |
| │ ├── latent_queries K learnable queries -> the K slots | |
| │ ├── encoder_blocks 4 x cross-attention (frames -> slots) | |
| │ ├── decoder_blocks 4 x cross-attention (slots -> frames) | |
| │ └── decoder_refiner dilated Conv1d ResNet -> 263-d motion | |
| ├── language_model (GPT2LMHeadModel) 124.4M | |
| │ └── 12 layers, 768 hidden, vocab 50261 (GPT-2 + <MOT> + 3 reserved) | |
| ├── layer router 9.1M | |
| │ ├── text_mem_proj prompt tokens -> routing memory | |
| │ ├── cross_attn_blocks 2 x cross-attention, per-slot routing queries | |
| │ ├── static_router_logits (K, 12) text-independent component | |
| │ └── layer_router MLP -> (K, 12) text-conditional component | |
| ├── calibrated latent head 0.9M | |
| │ ├── projector hidden -> (mu, logvar) per slot | |
| │ ├── latent_factors (K*512, 64) global covariance basis U | |
| │ └── factor_scale_head hidden -> per-factor log amplitudes | |
| └── motion_in_projector latent slots -> GPT-2 embeddings (captioning) 0.4M | |
| ``` | |
| **Adaptive-length autoencoder.** Cross-attention compresses a clip of *any* | |
| length into K = 2 continuous vectors of width 512, and a second cross-attention | |
| stack expands them back to *any* requested frame count. Decoder queries encode a | |
| frame's relative phase within the clip rather than an absolute frame index, which | |
| is why one decoder serves every length, and why you can decode the same latents | |
| at 60 frames and at 200 frames. | |
| **Depth-routed hidden states.** A text-conditioned router gives each latent slot | |
| its own soft mixture over all twelve transformer layers, so a slot reads from the | |
| depth it needs instead of squeezing every piece of motion evidence through the | |
| final layer. Both halves of the routing logit are tanh-bounded, so no logit | |
| margin can saturate the routing softmax. | |
| **Calibrated latent head.** The head predicts | |
| `N(mu, U diag(a)^2 U^T + diag(sigma^2))` over the whole flattened `K x 512` | |
| latent, with rank 64, trained by exact maximum likelihood. One draw therefore | |
| carries text-conditional variance that is *correlated across slots*, which is | |
| what a single-step sampler has to supply all at once. | |
| | | | | |
| |---|---| | |
| | Motion features | HumanML3D, 263-d per frame, 20 fps, 22 joints | | |
| | Latent slots | K = 2, width 512 | | |
| | Language model | GPT-2 (124M), fine-tuned | | |
| | Covariance rank | 64 | | |
| | Training | Joint generation + understanding, 240 epochs (Stage 2) | | |
| | Parameters | 183,690,603 | | |
| | Precision | float32 | | |
| --- | |
| ## Installation | |
| ```bash | |
| pip install torch transformers safetensors numpy | |
| ``` | |
| `trust_remote_code=True` is required: the autoencoder, the router and the | |
| calibrated head are custom modules that ship with this repository. | |
| **Nothing else is needed.** The model carries its own tokenizer and its own | |
| HumanML3D feature statistics, so it returns motion in real units without you | |
| downloading the dataset. | |
| --- | |
| ## Quick start | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model = AutoModelForCausalLM.from_pretrained( | |
| "zy22b/MUGEN", trust_remote_code=True | |
| ).eval() | |
| tokenizer = AutoTokenizer.from_pretrained("zy22b/MUGEN") | |
| features = model.generate_motion( | |
| ["a person walks forward and then waves with the right hand."], | |
| lengths=[120], # frames at 20 fps, so 6 seconds | |
| tokenizer=tokenizer, | |
| ) | |
| print(features.shape) # torch.Size([1, 120, 263]) | |
| joints = model.features_to_joints(features) | |
| print(joints.shape) # torch.Size([1, 120, 22, 3]) metres | |
| print(model.generate_caption(features, tokenizer=tokenizer)) | |
| # ['a person walks forward, turns to the left, and walks back.'] | |
| ``` | |
| That last line captions the motion the model just generated, so it changes from | |
| call to call: the motion is *sampled*, not a fixed function of the prompt. | |
| --- | |
| ## Inference tutorial | |
| ### 1. Text to motion | |
| ```python | |
| prompts = [ | |
| "a person walks forward and then waves with the right hand.", | |
| "the person jumps up and lands with both feet together.", | |
| "a man slowly sits down on a chair.", | |
| "a person runs in a circle to the left.", | |
| ] | |
| features = model.generate_motion(prompts, lengths=120, tokenizer=tokenizer) | |
| ``` | |
| `generate_motion` returns `(batch, frames, 263)` denormalised HumanML3D features. | |
| | Argument | Meaning | | |
| |---|---| | |
| | `texts` | one description, or a list of them | | |
| | `lengths` | frames per description at 20 fps; an `int` applies to the whole batch | | |
| | `temperature` | multiplier on the sampled perturbation; defaults to the calibrated `1.0` | | |
| | `generator` | a `torch.Generator` for a reproducible draw | | |
| | `denormalize` | `True` (default) returns raw HumanML3D units | | |
| **Choosing a length.** HumanML3D clips are multiples of 4 frames and at most 196 | |
| (9.8 s), so stay in `[20, 196]` and on the multiple-of-4 grid to stay in | |
| distribution. The decoder will happily produce other lengths, but you are then | |
| outside what it was trained on. | |
| **Per-row lengths.** Because decoder queries are relative-phase, a row must be | |
| decoded at its own length rather than cropped from a longer one. Pass a list and | |
| the model handles it; the returned tensor is zero-padded to the longest row. | |
| ```python | |
| features = model.generate_motion(prompts[:2], lengths=[60, 196], tokenizer=tokenizer) | |
| # (2, 196, 263); row 0 is valid up to frame 60 | |
| ``` | |
| ### 2. Temperature, and why the default is 1.0 | |
| The head predicts a distribution, and `temperature` scales the whole zero-mean | |
| perturbation around its mean. | |
| ```python | |
| # The calibrated conditional distribution. This is the protocol every reported | |
| # number uses, and different calls give different motions for the same prompt. | |
| features = model.generate_motion(prompts, lengths=120, tokenizer=tokenizer) | |
| # The distribution mean. Deterministic, and a *regression* protocol: numbers | |
| # obtained this way are not comparable to generative ones. | |
| mean_motion = model.generate_motion(prompts, lengths=120, tokenizer=tokenizer, | |
| temperature=0.0) | |
| # Reproducible sampling. | |
| g = torch.Generator().manual_seed(1234) | |
| features = model.generate_motion(prompts, lengths=120, tokenizer=tokenizer, generator=g) | |
| ``` | |
| Sampling at temperature 1.0 is what makes the same prompt yield genuinely | |
| different, plausible motions. Captioning five seeded draws of the same prompt | |
| shows how far apart they land: | |
| ``` | |
| prompt: "a person walks forward and then waves with the right hand." | |
| seed 0 a person walks forward and then turns around and walks back. | |
| seed 1 a person walks forward and reaches out with their right hand. | |
| seed 2 a person walks forward and then waves with their left hand. | |
| seed 3 a person walks forward and puts their right hand on their head. | |
| seed 4 a person walks forward and then puts something on a counter. | |
| ``` | |
| Lowering the temperature trades that diversity for proximity to the conditional | |
| mean. | |
| ### 3. Motion to text | |
| ```python | |
| import numpy as np | |
| motion = np.load("000004.npy") # (T, 263) HumanML3D features | |
| clip = torch.from_numpy(motion).float().unsqueeze(0) | |
| print(model.generate_caption(clip, tokenizer=tokenizer)) | |
| # ['a person bends over to touch the ground with their hands, then stands up straight.'] | |
| ``` | |
| Pass `normalized=True` if your features are already standardised by the dataset | |
| statistics. Captioning is greedy by default; `max_new_tokens` and `num_beams` | |
| override the config values. | |
| ### 4. The shared latent interface | |
| The K slots are the only motion representation in the system: the language model | |
| generates them for text-to-motion and reads these same vectors back for | |
| captioning. You can work with them directly. | |
| ```python | |
| latents = model.encode_motion(clip) # (1, 2, 512): a whole clip, two vectors | |
| print(latents.shape) | |
| # Decode at the original length, and at a different one. | |
| same = model.decode_motion(latents, clip.shape[1]) | |
| longer = model.decode_motion(latents, 200) | |
| # Caption straight from latents, skipping the encoder. | |
| print(model.generate_caption(latents=latents, tokenizer=tokenizer)) | |
| # Text -> latents, without decoding. | |
| z = model.text_to_latents(["a person waves."], tokenizer=tokenizer) | |
| ``` | |
| Interpolating between two clips is then a two-line operation: | |
| ```python | |
| a = model.encode_motion(clip_a) | |
| b = model.encode_motion(clip_b) | |
| blend = model.decode_motion(0.5 * a + 0.5 * b, 120) | |
| ``` | |
| ### 5. Running on GPU, and in batches | |
| ```python | |
| model = model.to("cuda") | |
| features = model.generate_motion(prompts, lengths=120, tokenizer=tokenizer) # follows the model | |
| ``` | |
| Batched prompts are left-padded, which reproduces the training and evaluation | |
| behaviour exactly. Single-prompt calls involve no padding at all. | |
| ### 6. Saving and rendering | |
| `features_to_joints` gives `(batch, frames, 22, 3)` positions in metres, on the | |
| HumanML3D skeleton, with the y axis vertical. | |
| ```python | |
| joints = model.features_to_joints(features) | |
| np.save("sample.npy", joints[0].cpu().numpy()) | |
| ``` | |
| For animation, feed those positions to any HumanML3D-compatible renderer. The | |
| [code repository](https://github.com/JYe16/MUGEN) includes | |
| `motGPT/utils/render_utils.py`, which writes an mp4 directly from a `[T, J, 3]` | |
| array. | |
| ### 7. Sanity check | |
| If you swap in different feature statistics, the failure is silent: shapes stay | |
| right and the numbers look plausible while the whole skeleton is mis-scaled. One | |
| assertion catches it. | |
| ```python | |
| joints = model.features_to_joints(features) | |
| high, low = float(joints[..., 1].max()), float(joints[..., 1].min()) | |
| assert 1.0 < high < 2.5 and -0.5 < low < 0.5, "implausible body height" | |
| ``` | |
| --- | |
| ## Reproducibility | |
| The inference path in `modeling_mugen.py` was verified against the original | |
| training code on the same checkpoint. Routed hidden states, `mu`, `logvar`, the | |
| factor log-amplitudes, the autoencoder encode and decode outputs, and a | |
| seeded draw all agree to **`max|diff| = 0.000e+00`**, and generated captions match | |
| token for token. | |
| Two consequences worth knowing: | |
| - A given seed reproduces the same draw here as it does in the research code, | |
| because the factor noise is drawn before the diagonal noise in both. | |
| - `router_eval_tau` (1.5) is the converged end of the training temperature | |
| anneal. Changing it changes which transformer depth each slot reads from, so | |
| leave it alone unless you are deliberately studying the router. | |
| Tested with `transformers` 4.47 and PyTorch 2.x. The rollout passes explicit | |
| `position_ids`, so behaviour does not drift with the transformers version. | |
| ## Limitations | |
| - Trained on **HumanML3D** only: everyday single-person motion, English | |
| descriptions, 20 fps, 22 joints. Multi-person interaction, object manipulation, | |
| and highly stylised motion are out of distribution. | |
| - Length is an **input**, not something the model infers from the text. Ask for | |
| 120 frames and you get 120 frames whether or not the description warrants it. | |
| - K = 2 slots is a deliberately tight budget, chosen for this dataset. It is not | |
| a universal setting: the SnapMoGen model in the code repository uses K = 4, and | |
| the best budget is dataset-dependent. | |
| - Captions come from a fine-tuned GPT-2 (124M) and are short and generic | |
| relative to human references. They describe the dominant action reliably and | |
| fine detail unreliably. On HumanML3D clip `000004`, whose reference reads *"a | |
| man is pretending to be a chicken, constantly pecking at the ground and waving | |
| his arms like a chicken"*, the model produces *"a person bends over to touch | |
| the ground with their hands, then stands up straight"*: the gross movement is | |
| right, the intent and the arm motion are gone. | |
| - Not a safety-filtered model. It generates skeletal motion and short English | |
| captions, and it has no content filtering of any kind. | |
| <!-- ## Citation | |
| The paper is under review. A citation entry will be added here once it appears; | |
| in the meantime please cite the repository and this model card. | |
| ```bibtex | |
| @misc{mugen2026, | |
| title = {MUGEN: A Unified Framework for Efficient Motion Understanding and Generation}, | |
| note = {Under review. Code: https://github.com/JYe16/MUGEN, | |
| model: https://huggingface.co/zy22b/MUGEN}, | |
| year = {2026} | |
| } | |
| ``` --> | |
| ## Acknowledgements | |
| Built on [HumanML3D](https://github.com/EricGuo5513/HumanML3D), | |
| [MotionGPT](https://github.com/OpenMotionLab/MotionGPT), | |
| [MLD](https://github.com/ChenFengYe/motion-latent-diffusion), and our earlier | |
| discrete-codebook system | |
| [GeoMotionGPT](https://huggingface.co/zy22b/GeoMotionGPT). | |
| ## License | |
| MIT | |