henrywch2huggingface's picture
Single-stage layout: one image/video block, chat/raw response view, fixed equal heights.
65be679
|
Raw
History Blame Contribute Delete
876 Bytes

A newer version of the Gradio SDK is available: 6.22.0

Upgrade
metadata
title: MOSS-VL-Instruct-0708
emoji: 🧠
colorFrom: gray
colorTo: blue
sdk: gradio
sdk_version: 6.15.1
app_file: app.py
short_description: Image and video understanding with MOSS-VL multimodal model
python_version: '3.12'
startup_duration_timeout: 1h

MOSS-VL-Instruct-0708

An 11B parameter vision-language model from OpenMOSS that supports both image and video understanding.

Capabilities

  • Image understanding: OCR, document parsing, fine-grained visual recognition, multi-image comparison
  • Video understanding: Long-form video comprehension, temporal reasoning, action recognition
  • 256K context window for processing long videos and complex instructions

Usage

Upload an image or video and enter a text prompt describing what you want the model to do. The model will generate a text response based on its understanding of the visual input.