henrywch2huggingface's picture
Single-stage layout: one image/video block, chat/raw response view, fixed equal heights.
65be679
|
Raw
History Blame Contribute Delete
876 Bytes
---
title: MOSS-VL-Instruct-0708
emoji: 🧠
colorFrom: gray
colorTo: blue
sdk: gradio
sdk_version: 6.15.1
app_file: app.py
short_description: Image and video understanding with MOSS-VL multimodal model
python_version: "3.12"
startup_duration_timeout: 1h
---
# MOSS-VL-Instruct-0708
An 11B parameter vision-language model from OpenMOSS that supports both image and video understanding.
## Capabilities
- **Image understanding**: OCR, document parsing, fine-grained visual recognition, multi-image comparison
- **Video understanding**: Long-form video comprehension, temporal reasoning, action recognition
- **256K context window** for processing long videos and complex instructions
## Usage
Upload an image or video and enter a text prompt describing what you want the model to do. The model will generate a text response based on its understanding of the visual input.