MODUS any-to-any multimodal demo (16 aligned modalities)
Test-time search over ordered visual tokens.
Segment images with prompts or automatic masks
Segment objects in images using text prompts or scribbles