Training-Free Speech-Centric Omni Understanding with Frozen VLMs Paper • 2609.04242 • Published Aug 7 • 4
jdopensource/JoyAI-VL-Interaction-Preview Video-Text-to-Text • 9B • Updated Jun 22 • 703 • 84
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence Paper • 2606.14777 • Published Jun 10 • 217
Running Agents 13 LLaVA OneVision 1.5 📉 13 Interact with a multimodal chatbot using text and images