FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Paper • 2606.31247 • Published Jun 30 • 5
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Paper • 2606.31247 • Published Jun 30 • 5
SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion Paper • 2402.12660 • Published Feb 20, 2024
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation Paper • 2505.13000 • Published May 19, 2025 • 1
Vevo2: Bridging Controllable Speech and Singing Voice Generation via Unified Prosody Learning Paper • 2508.16332 • Published Aug 22, 2025
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment Paper • 2505.04113 • Published May 7, 2025
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness Paper • 2511.07931 • Published Nov 11, 2025
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation Paper • 2407.05361 • Published Jul 7, 2024 • 2
Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation Paper • 2501.15907 • Published Jan 27, 2025 • 19
Closing the Modality Reasoning Gap for Speech Large Language Models Paper • 2601.05543 • Published Jan 9 • 1
BATON: Aligning Text-to-Audio Model with Human Preference Feedback Paper • 2402.00744 • Published Feb 1, 2024
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis Paper • 2409.08628 • Published Sep 13, 2024
AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward Paper • 2411.18654 • Published Nov 27, 2024
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Paper • 2502.03128 • Published Feb 5, 2025 • 2
NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations Paper • 2508.04195 • Published Aug 6, 2025 • 2