--- title: MOSS-VL-Instruct-0708 emoji: 🧠 colorFrom: gray colorTo: blue sdk: gradio sdk_version: 6.15.1 app_file: app.py short_description: Image and video understanding with MOSS-VL multimodal model python_version: "3.12" startup_duration_timeout: 1h --- # MOSS-VL-Instruct-0708 An 11B parameter vision-language model from OpenMOSS that supports both image and video understanding. ## Capabilities - **Image understanding**: OCR, document parsing, fine-grained visual recognition, multi-image comparison - **Video understanding**: Long-form video comprehension, temporal reasoning, action recognition - **256K context window** for processing long videos and complex instructions ## Usage Upload an image or video and enter a text prompt describing what you want the model to do. The model will generate a text response based on its understanding of the visual input.