File size: 876 Bytes
670fb9c
c346860
 
 
 
670fb9c
65be679
670fb9c
c346860
 
 
670fb9c
 
c346860
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
---
title: MOSS-VL-Instruct-0708
emoji: 🧠
colorFrom: gray
colorTo: blue
sdk: gradio
sdk_version: 6.15.1
app_file: app.py
short_description: Image and video understanding with MOSS-VL multimodal model
python_version: "3.12"
startup_duration_timeout: 1h
---

# MOSS-VL-Instruct-0708

An 11B parameter vision-language model from OpenMOSS that supports both image and video understanding.

## Capabilities

- **Image understanding**: OCR, document parsing, fine-grained visual recognition, multi-image comparison
- **Video understanding**: Long-form video comprehension, temporal reasoning, action recognition
- **256K context window** for processing long videos and complex instructions

## Usage

Upload an image or video and enter a text prompt describing what you want the model to do. The model will generate a text response based on its understanding of the visual input.