Running Agents 75 Qwen3.5 Omni Online Demo 📚 75 Chat with a multimodal AI using text, image, audio, or video
Running Agents 119 Qwen3 TTS Voice Design 📈 119 Generate custom speech from text and voice description
SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity Paper • 2510.23541 • Published Oct 27, 2025 • 17
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Paper • 2509.08519 • Published Sep 10, 2025 • 132
ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing Paper • 2506.21448 • Published Jun 26, 2025 • 9
Running on Zero Agents Featured 685 ACE Step 😻 685 A Step Towards Music Generation Foundation Model