--- license: apache-2.0 language: - en - zh library_name: transformers pipeline_tag: video-text-to-text base_model: OpenMOSS-Team/MOSS-VL-Instruct-0708 tags: - MOSS-VL - image-understanding - video-understanding - FP8 - compressed-tensors - quantized - SGLang - custom_code ---

MOSS-VL

English | 中文

# MOSS-VL-Instruct-0708 FP8 Dynamic + Transformers KV8 This is the 24 GiB quantized release of MOSS-VL-Instruct-0708. The same checkpoint directory is supported by both the standard Transformers inference path and the native SGLang MOSS-VL backend. ## Quantization profile | Component | Format | | --- | --- | | Most language layers | FP8 weights + per-token dynamic FP8 input activations | | Cross-attention, vision modules and `lm_head` | BF16 | | Transformers KV cache | HQQ INT8 | | SGLang KV cache | Native SGLang BF16 cache | The FP8 weights are shared across both engines. Transformers reads the HQQ KV8 configuration from `generation_config.json`; SGLang loads the same weights with `model_impl=sglang` and uses its own cache implementation. ## Quantization benchmark Across the selected benchmarks, the quantized models remain close to their non-quantized BF16 counterparts, showing that overall model quality is largely preserved after quantization.

MOSS-VL quantization benchmark comparison

## Hardware requirements The model is designed to run with the Transformers inference path on a single NVIDIA GPU with 24 GB of VRAM. SGLang memory requirements depend on its server configuration and KV cache allocation. ## Transformers environment ### Installation Start from the standard MOSS-VL release environment, then install the FP8 checkpoint loader and HQQ cache backend: ```bash git clone https://github.com/OpenMOSS/MOSS-VL.git cd MOSS-VL conda create -n moss_vl_quant python=3.12 pip -y conda activate moss_vl_quant pip install -i https://pypi.org/simple --no-build-isolation -r requirements.txt pip install -i https://pypi.org/simple \ compressed-tensors==0.14.0 \ hqq==0.2.8.post1 python -m pip check ``` Core versions used by the standard Transformers path: | Package | Version | | --- | --- | | Python | 3.12 | | PyTorch | 2.8.0 + CUDA 12.8 | | Transformers | 4.57.1 | | Accelerate | 1.12.0 | | FlashAttention | 2.8.1 | | compressed-tensors checkpoint format | 0.14.0 | | HQQ | 0.2.8.post1 | Video decoding also requires FFmpeg to be available in `PATH`. ### Load the model ```python import torch from transformers import AutoModelForCausalLM, AutoProcessor checkpoint = "/path/to/MOSS-VL-0708-Instruct-FP8-Dynamic-KV8-HQQ" processor = AutoProcessor.from_pretrained( checkpoint, trust_remote_code=True, frame_extract_num_threads=1, ) model = AutoModelForCausalLM.from_pretrained( checkpoint, trust_remote_code=True, device_map="auto", torch_dtype=torch.bfloat16, attn_implementation="flash_attention_2", ) model.eval() ``` `generation_config.json` enables HQQ KV8 automatically. Keep that file beside the checkpoint and do not pass a conflicting cache configuration. ### Image inference ```python text = model.offline_image_generate( processor, prompt="Describe this image.", image="data/example_image.jpg", shortest_edge=4096, longest_edge=16777216, multi_image_max_pixels=201326592, patch_size=16, temporal_patch_size=1, merge_size=2, image_mean=[0.5, 0.5, 0.5], image_std=[0.5, 0.5, 0.5], max_new_tokens=256, temperature=1.0, top_k=50, top_p=1.0, repetition_penalty=1.0, do_sample=False, vision_chunked_length=64, ) print(text) ``` ### Video inference ```python text = model.offline_video_generate( processor, prompt="Describe this video.", video="data/example_video.mp4", shortest_edge=4096, longest_edge=16777216, video_max_pixels=201326592, patch_size=16, temporal_patch_size=1, merge_size=2, video_fps=1.0, min_frames=1, max_frames=256, num_extract_threads=4, image_mean=[0.5, 0.5, 0.5], image_std=[0.5, 0.5, 0.5], max_new_tokens=256, temperature=1.0, top_k=50, top_p=1.0, repetition_penalty=1.0, do_sample=False, vision_chunked_length=64, ) print(text) ``` ### Repository inference runner The standard repository runner accepts JSON/JSONL queries with image or video content and loads this directory directly: ```bash python inference/run_inference.py \ --checkpoint /path/to/MOSS-VL-0708-Instruct-FP8-Dynamic-KV8-HQQ \ --mode offline \ --input inference/video_queries.json \ --output video_results.json ``` ## SGLang environment Use a separate environment for SGLang. The native backend in the `mossvl_sglang` release directory pins: | Package | Version | | --- | --- | | Python | 3.12.8 | | SGLang | 0.5.11 | | sglang-kernel | 0.4.2 | | PyTorch | 2.11.0 + CUDA 13.0 | | Transformers | 5.6.0 | | compressed-tensors | 0.17.1 | Create the environment from its complete lock file: ```bash cd /path/to/mossvl_sglang env -u INDEX_URL -u PIP_INDEX_URL conda create -y -n mossvl-sglang-0511 \ --override-channels \ -c https://mirrors.tuna.tsinghua.edu.cn/anaconda/pkgs/main \ python=3.12.8 pip=24.2 conda activate mossvl-sglang-0511 env -u INDEX_URL -u PIP_INDEX_URL python -m pip install -r requirements.txt python -m pip check ``` Start the OpenAI-compatible service with the quantized directory as the model path: ```bash conda activate mossvl-sglang-0511 cd /path/to/mossvl_sglang GPU_IDS=0 TP_SIZE=1 MEM_FRACTION_STATIC=0.35 \ ./start_sglang_moss_vl.sh \ --model-path /path/to/MOSS-VL-0708-Instruct-FP8-Dynamic-KV8-HQQ ``` The service listens on `http://127.0.0.1:30000` by default. Keep `MODEL_IMPL=sglang`; do not force the Transformers remote implementation. SGLang intentionally uses its native BF16 KV cache rather than the HQQ cache in `generation_config.json`. ## Configuration files - `config.json`: model and FP8 weight/activation configuration. - `generation_config.json`: Transformers HQQ KV8 configuration. - `modeling_moss_vl.py`: checkpoint-local cross-attention/QuantizedCache code.