Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction Paper • 2609.21392 • Published 8 days ago
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Paper • 2609.08936 • Published 18 days ago • 165
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction Paper • 2608.26005 • Published Aug 26 • 163
GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark Paper • 2606.28884 • Published Jun 27 • 1
GigaSpeech Series Collection Evolving, Large-Scale, and Multi-domain ASR Corpus • 6 items • Updated Jun 26
UAT: Unified Audio-Text Diffusion for Audio Generation, Editing, and Captioning Paper • 2606.04939 • Published Jun 3
Evaluating the Expressive Appropriateness of Speech in Rich Contexts Paper • 2605.09413 • Published May 10 • 5
Evaluating the Expressive Appropriateness of Speech in Rich Contexts Paper • 2605.09413 • Published May 10 • 5