Chorus VOICE STUDIO

What's the scene?

Untitled conversation

No scene yet
Describe a scenario above and click Write with AI, import a script, or start typing your own line below.
Add dialogue to begin
NOW PLAYING Β·
0:00 / 0:00
Download WAV MP3

Choose your voices

Clone a voice

Record or upload 15–30 seconds of one person speaking clearly, with no background noise or music.

Import a script

Use Speaker 1: … Speaker 4: tags, one turn per line. Untagged text becomes Speaker 1. Named characters (Host:, Mom:) are auto-mapped to speaker slots.

No file chosen

How Chorus works

VibeVoice generates expressive, long-form, multi-speaker conversational audio from text. It uses continuous speech tokenizers at an ultra-low 7.5 Hz frame rate and a next-token diffusion framework to synthesize up to 90 minutes of speech with up to 4 distinct speakers.

VibeVoice architecture diagram