MOSS-VL

English | 中文

MOSS-VL-Realtime W4A16 NF4 + KV8 HQQ

This is the 24 GiB quantized release of MOSS-VL-Realtime. It keeps the original timestamp-aware streaming interface and can also use the offline image/video helpers from the standard checkpoint.

Quantization profile

Component Format
Most language layers bitsandbytes NF4 4-bit weights with double quantization
First/last language layers and multimodal modules BF16
Activations and compute BF16
Transformers KV cache HQQ INT8
Attention backend FlashAttention 2

The checkpoint carries its bitsandbytes configuration, HQQ cache configuration, and MOSS-VL remote modeling code. Load the directory directly; do not add a second runtime quantization configuration.

Quantization benchmark

Across the selected benchmarks, the quantized models remain close to their non-quantized BF16 counterparts, showing that overall model quality is largely preserved after quantization.

MOSS-VL quantization benchmark comparison

Hardware requirements

The model is designed to run on a single NVIDIA GPU with 24 GB of VRAM. Use FlashAttention 2 and frame_queue_size=1 for the 24 GB realtime profile.

Environment

Installation

Use the standard MOSS-VL repository requirements, then add the two quantization backends required by this checkpoint:

git clone https://github.com/OpenMOSS/MOSS-VL.git
cd MOSS-VL

conda create -n moss_vl_quant python=3.12 pip -y
conda activate moss_vl_quant
pip install -i https://pypi.org/simple --no-build-isolation -r requirements.txt
pip install -i https://pypi.org/simple \
  bitsandbytes==0.49.2 \
  hqq==0.2.8.post1
python -m pip check

The standard release environment uses the following core stack:

Package Version
Python 3.12
PyTorch 2.8.0 + CUDA 12.8
Transformers 4.57.1
Accelerate 1.12.0
FlashAttention 2.8.1
bitsandbytes 0.49.2
HQQ 0.2.8.post1

Video decoding also requires FFmpeg to be available in PATH.

Load the model

Keep attn_implementation set to flash_attention_2 for the 24 GB profile.

import torch
from transformers import AutoModelForCausalLM, AutoProcessor

checkpoint = "/path/to/mossvl_streaming_w4a16_nf4_keep_first4_last4_kv8_hqq"

processor = AutoProcessor.from_pretrained(
    checkpoint,
    trust_remote_code=True,
    frame_extract_num_threads=1,
)
model = AutoModelForCausalLM.from_pretrained(
    checkpoint,
    trust_remote_code=True,
    device_map="auto",
    torch_dtype=torch.bfloat16,
    attn_implementation="flash_attention_2",
)
model.eval()

generation_config.json automatically enables the HQQ INT8 KV cache. Do not override it with a BF16/dynamic cache when using the 24 GB profile.

Realtime inference

The application supplies PIL-compatible frames with non-decreasing timestamps. Use frame_queue_size=1 for the 24 GB realtime profile.

import time
from PIL import Image

session = model.create_realtime_session(
    processor,
    initial_prompt=(
        "Describe important changes in the video as they happen. "
        "Stay silent when there is no meaningful update."
    ),
    frame_queue_size=1,
    max_tokens_per_turn=12,
    max_new_tokens=4096,
    do_sample=False,
)

frame_paths = [
    "data/frame_0001.jpg",
    "data/frame_0002.jpg",
    "data/frame_0003.jpg",
]

try:
    session.start()
    for index, frame_path in enumerate(frame_paths):
        image = Image.open(frame_path).convert("RGB")
        session.push_frame(image, timestamp=float(index))

        while True:
            chunk = session.poll_output(timeout=0.0)
            if chunk is None:
                break
            print(chunk, end="", flush=True)

        time.sleep(1.0)

    session.push_prompt("What changed in the latest frames?")
    deadline = time.monotonic() + 5.0
    while time.monotonic() < deadline:
        chunk = session.poll_output(timeout=0.1)
        if chunk is not None:
            print(chunk, end="", flush=True)
finally:
    session.close()

One model instance supports one active realtime session. The model may emit control tokens such as <|silence|>, <|round_start|>, and <|round_end|>; applications should filter or render them according to their protocol.

Offline video inference

text = model.offline_video_generate(
    processor,
    prompt="Describe this video.",
    video="data/example_video.mp4",
    shortest_edge=4096,
    longest_edge=16777216,
    video_max_pixels=201326592,
    patch_size=16,
    temporal_patch_size=1,
    merge_size=2,
    video_fps=1.0,
    min_frames=1,
    max_frames=256,
    num_extract_threads=4,
    image_mean=[0.5, 0.5, 0.5],
    image_std=[0.5, 0.5, 0.5],
    max_new_tokens=256,
    temperature=1.0,
    top_k=50,
    top_p=1.0,
    repetition_penalty=1.0,
    do_sample=False,
    vision_chunked_length=64,
)
print(text)

Configuration files

  • config.json: model and NF4 weight configuration.
  • generation_config.json: HQQ KV8 configuration.
  • modeling_moss_vl.py: checkpoint-local MOSS-VL and QuantizedCache code.
Downloads last month
-
Safetensors
Model size
11B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OpenMOSS-Team/MOSS-VL-Realtime-NF4

Quantized
(2)
this model

Collection including OpenMOSS-Team/MOSS-VL-Realtime-NF4