soundsgoodai/Zipformer-cr-ctc-transducer-XL-290M Automatic Speech Recognition • Updated 29 days ago • 111 • 2
HuggingFaceTB/SmolVLM2-500M-Video-Instruct Image-Text-to-Text • 0.5B • Updated Apr 8, 2025 • 1.47M • 170
microsoft/VibeVoice-ASR-BitNet Automatic Speech Recognition • 0.3B • Updated 23 days ago • 16.3k • 181
[Data] VLM SFT Collection This collection gathers clean multimodal datasets for VLM supervised fine-tuning. Covering visual QA, image dialogue and document reasoning. • 2 items • Updated 5 days ago
google/siglip2-base-patch16-224 Zero-Shot Image Classification • 0.4B • Updated Feb 21, 2025 • 1.09M • 128