The massive multimodal embedding benchmark
Detect objects, keypoints, or text in your images
Segment objects in images and videos using text prompts
Chat with an AI assistant using text and images