--- library_name: pytorch license: mit pipeline_tag: text-to-speech tags: - accent-tts - mandarin - joycent - grad-tts - whisaid --- # Joycent trained with WhisAID Medium GRL accent embeddings This is a Joycent Mandarin accent TTS acoustic model trained using accent embeddings extracted by [`walston/whisaid-medium-grl`](https://huggingface.co/walston/whisaid-medium-grl). The released checkpoint is epoch 100. - **Paper:** [Joycent: Diffusion-based Accent TTS without Accented Phone Prediction](https://huggingface.co/papers/2606.16417) - **Code:** [oshindow/Joycent-code](https://github.com/oshindow/Joycent-code) - **Accent encoder:** [walston/whisaid-medium-grl](https://huggingface.co/walston/whisaid-medium-grl) - **Vocoder:** [walston/joycent-vocoder](https://huggingface.co/walston/joycent-vocoder) ## Download ```python from huggingface_hub import hf_hub_download checkpoint_path = hf_hub_download( repo_id="walston/joycent-medium-grl", filename="grad_100.pt", ) ``` Pass the downloaded checkpoint to `joycent/inference_joycent.py` with the `--acoustic-checkpoint` argument. Full synthesis also requires the Joycent vocoder and reference-audio feature extraction dependencies described in the Joycent repository. ## Checkpoint - Epoch: 100 - Acoustic model: Joycent / Grad-TTS - Accent embedding model: WhisAID Whisper Medium GRL (lambda 0.05) - Accent embedding dimension: 256 ## Citation ```bibtex @misc{wang2026joycentdiffusionbasedaccenttts, title={Joycent: Diffusion-based Accent TTS without Accented Phone Prediction}, author={Xintong Wang and Ye Wang}, year={2026}, eprint={2606.16417}, archivePrefix={arXiv}, primaryClass={cs.SD}, } ```