EchoMask official checkpoints
Official checkpoints for EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation (ACM Multimedia 2025), a speech-conditioned approach to holistic co-speech gesture generation.
Files
| Archive | Protocol | Contents |
|---|---|---|
weights_echomask_no_smplx.zip |
Speaker 2, paper protocol | Original pretrained representation models and EchoMask checkpoint; SMPL-X files excluded |
EchoMask_all_speakers_weights_25spk.zip |
25 English BEAT2 speakers | Reusable representation models and selected EchoMask checkpoint |
The two archives implement different training protocols and should not be treated as a controlled single-speaker versus multi-speaker ablation. See the code repository for installation paths and configuration details.
Verify both downloads with sha256sum -c SHA256SUMS.
Reported results
| Protocol | Speakers | FGD ↓ | BC ↑ | DIV ↑ | MSE ↓ | LVD ↓ |
|---|---|---|---|---|---|---|
| Speaker 2 (paper) | 1 | 0.4623 | 0.7738 | 13.370 | 6.761e-8 | 7.290e-5 |
| Released all-speaker checkpoint | 25 | 0.5656 | 0.4951 | 9.299 | 4.700e-8 | 6.090e-5 |
The first row is from the paper. The second row is from the released all-speaker checkpoint. Their protocols differ.
Terms
This repository mirrors research artifacts released by the authors. No new license is granted by this model card. The Hugging Face archives intentionally exclude SMPLX_NEUTRAL_2020.npz; obtain SMPL-X directly from the official source under its terms. Code, BEAT2, pretrained encoders, and other third-party assets remain subject to their respective terms.
Citation
@inproceedings{zhang2025echomask,
title={EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation},
author={Zhang, Xiangyue and Li, Jianfang and Zhang, Jiaxu and Ren, Jianqiang and Bo, Liefeng and Tu, Zhigang},
booktitle={Proceedings of the 33rd ACM International Conference on Multimedia},
pages={10827--10836},
year={2025},
doi={10.1145/3746027.3754847}
}