Collections
Discover the best community collections!
Collections trending this week
-
Salesforce/instructblip-vicuna-7b
Image-Text-to-Text • 8B • Updated • 10.7k • 102 -
Salesforce/instructblip-vicuna-13b
Image-Text-to-Text • 14B • Updated • 168 • 43 -
Salesforce/instructblip-flan-t5-xxl
Image-Text-to-Text • 12B • Updated • 272 • 21 -
Salesforce/instructblip-flan-t5-xl
Image-Text-to-Text • 4B • Updated • 16.1k • 30
-
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
Paper • 2602.11858 • Published • 61 -
inclusionAI/ZwZ-4B
Image-Text-to-Text • 5B • Updated • 91 • 31 -
inclusionAI/ZwZ-8B
Image-Text-to-Text • 9B • Updated • 219 • 46 -
inclusionAI/ZwZ-RL-VQA
Viewer • Updated • 111k • 1.33k • 15
-
Salesforce/blip2-opt-2.7b
Image-Text-to-Text • 4B • Updated • 522k • 446 -
Salesforce/blip2-flan-t5-xxl
Image-Text-to-Text • 12B • Updated • 1.14k • 94 -
Salesforce/blip2-opt-6.7b-coco
Image-Text-to-Text • 8B • Updated • 1.12k • 35 -
Salesforce/blip2-opt-6.7b
Image-Text-to-Text • 8B • Updated • 99.4k • 80
-
inclusionAI/Ming-flash-omni-2.0
Any-to-Any • 104B • Updated • 2.47k • 267 -
inclusionAI/Ming-omni-tts-16.8B-A3B
Text-to-Speech • 18B • Updated • 1.24k • 54 -
inclusionAI/Ming-omni-tts-0.5B
Text-to-Speech • 2B • Updated • 9.38k • 36 -
inclusionAI/Ming-omni-tts-tokenizer-12Hz
Audio-to-Audio • 0.8B • Updated • 21 • 10
-
Salesforce/instructblip-vicuna-7b
Image-Text-to-Text • 8B • Updated • 10.7k • 102 -
Salesforce/instructblip-vicuna-13b
Image-Text-to-Text • 14B • Updated • 168 • 43 -
Salesforce/instructblip-flan-t5-xxl
Image-Text-to-Text • 12B • Updated • 272 • 21 -
Salesforce/instructblip-flan-t5-xl
Image-Text-to-Text • 4B • Updated • 16.1k • 30
-
Salesforce/blip2-opt-2.7b
Image-Text-to-Text • 4B • Updated • 522k • 446 -
Salesforce/blip2-flan-t5-xxl
Image-Text-to-Text • 12B • Updated • 1.14k • 94 -
Salesforce/blip2-opt-6.7b-coco
Image-Text-to-Text • 8B • Updated • 1.12k • 35 -
Salesforce/blip2-opt-6.7b
Image-Text-to-Text • 8B • Updated • 99.4k • 80
-
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
Paper • 2602.11858 • Published • 61 -
inclusionAI/ZwZ-4B
Image-Text-to-Text • 5B • Updated • 91 • 31 -
inclusionAI/ZwZ-8B
Image-Text-to-Text • 9B • Updated • 219 • 46 -
inclusionAI/ZwZ-RL-VQA
Viewer • Updated • 111k • 1.33k • 15
-
inclusionAI/Ming-flash-omni-2.0
Any-to-Any • 104B • Updated • 2.47k • 267 -
inclusionAI/Ming-omni-tts-16.8B-A3B
Text-to-Speech • 18B • Updated • 1.24k • 54 -
inclusionAI/Ming-omni-tts-0.5B
Text-to-Speech • 2B • Updated • 9.38k • 36 -
inclusionAI/Ming-omni-tts-tokenizer-12Hz
Audio-to-Audio • 0.8B • Updated • 21 • 10