Parse document images into structured Markdown
Unified foundation model for promptable segmentation
Detect and label objects in images and videos