Real-time speech and talking-head models Collection Open models for real-time voice pipelines: VAD, diarization, streaming ASR, TTS, full-duplex speech, and talking-head animation. • 44 items • Updated 7 days ago
facebook/seamless-m4t-v2-large Automatic Speech Recognition • 2B • Updated Jan 4, 2024 • 333k • 1.01k