Record or upload 15β30 seconds of one person speaking clearly, with no background noise or music.
Use Speaker 1: β¦ Speaker 4: tags, one turn per line. Untagged text becomes Speaker 1. Named characters (Host:, Mom:) are auto-mapped to speaker slots.
Speaker 1:
Speaker 4:
Host:
Mom:
VibeVoice generates expressive, long-form, multi-speaker conversational audio from text. It uses continuous speech tokenizers at an ultra-low 7.5 Hz frame rate and a next-token diffusion framework to synthesize up to 90 minutes of speech with up to 4 distinct speakers.