VoxAudio / README.md
nielsr's picture
nielsr HF Staff
Add pipeline tag and links to model card
b727737 verified
|
Raw
History Blame
711 Bytes
metadata
license: cc-by-sa-4.0
pipeline_tag: text-to-audio

VoxAudio

VoxAudio is a streaming chunk-autoregressive flow matching model for vocalized audio generation: given a text caption that may quote explicit speech, it generates 24 kHz audio containing articulate speech together with the surrounding soundscape.

For setup and inference instructions, please refer to the GitHub repository.