meta-models/Muse-Glimmer-30B Image-Text-to-Text • 30B • Updated 2 days ago • 121k • • 1.45k
nvidia/nemotron-speech-streaming-en-0.6b Automatic Speech Recognition • 0.6B • Updated 8 days ago • 128k • 603
Running on Zero Agents Featured 175 AudioX 👀 175 Generate audio from text, video, or audio prompts
Running on Zero Agents Featured 5.08k MusicGen 🎵 5.08k Generate music from a text description and optional melody
view article Article ColFlor: Towards BERT-Size Vision-Language Document Retrieval Models ahmed-masry • Oct 18, 2024 • 22
Qwen/Qwen2-VL-2B-Instruct-GPTQ-Int4 Image-Text-to-Text • 2B • Updated Sep 21, 2024 • 820 • 28