--- license: apache-2.0 language: - en - zh library_name: transformers pipeline_tag: video-text-to-text base_model: OpenMOSS-Team/MOSS-VL-Realtime tags: - MOSS-VL - realtime - streaming - video-understanding - bitsandbytes - NF4 - quantized - custom_code ---

MOSS-VL

English | 中文

# MOSS-VL-Realtime W4A16 NF4 + KV8 HQQ This is the 24 GiB quantized release of MOSS-VL-Realtime. It keeps the original timestamp-aware streaming interface and can also use the offline image/video helpers from the standard checkpoint. ## Quantization profile | Component | Format | | --- | --- | | Most language layers | bitsandbytes NF4 4-bit weights with double quantization | | First/last language layers and multimodal modules | BF16 | | Activations and compute | BF16 | | Transformers KV cache | HQQ INT8 | | Attention backend | FlashAttention 2 | The checkpoint carries its bitsandbytes configuration, HQQ cache configuration, and MOSS-VL remote modeling code. Load the directory directly; do not add a second runtime quantization configuration. ## Quantization benchmark Across the selected benchmarks, the quantized models remain close to their non-quantized BF16 counterparts, showing that overall model quality is largely preserved after quantization.

MOSS-VL quantization benchmark comparison

## Hardware requirements The model is designed to run on a single NVIDIA GPU with 24 GB of VRAM. Use FlashAttention 2 and `frame_queue_size=1` for the 24 GB realtime profile. ## Environment ### Installation Use the standard MOSS-VL repository requirements, then add the two quantization backends required by this checkpoint: ```bash git clone https://github.com/OpenMOSS/MOSS-VL.git cd MOSS-VL conda create -n moss_vl_quant python=3.12 pip -y conda activate moss_vl_quant pip install -i https://pypi.org/simple --no-build-isolation -r requirements.txt pip install -i https://pypi.org/simple \ bitsandbytes==0.49.2 \ hqq==0.2.8.post1 python -m pip check ``` The standard release environment uses the following core stack: | Package | Version | | --- | --- | | Python | 3.12 | | PyTorch | 2.8.0 + CUDA 12.8 | | Transformers | 4.57.1 | | Accelerate | 1.12.0 | | FlashAttention | 2.8.1 | | bitsandbytes | 0.49.2 | | HQQ | 0.2.8.post1 | Video decoding also requires FFmpeg to be available in `PATH`. ## Load the model Keep `attn_implementation` set to `flash_attention_2` for the 24 GB profile. ```python import torch from transformers import AutoModelForCausalLM, AutoProcessor checkpoint = "/path/to/mossvl_streaming_w4a16_nf4_keep_first4_last4_kv8_hqq" processor = AutoProcessor.from_pretrained( checkpoint, trust_remote_code=True, frame_extract_num_threads=1, ) model = AutoModelForCausalLM.from_pretrained( checkpoint, trust_remote_code=True, device_map="auto", torch_dtype=torch.bfloat16, attn_implementation="flash_attention_2", ) model.eval() ``` `generation_config.json` automatically enables the HQQ INT8 KV cache. Do not override it with a BF16/dynamic cache when using the 24 GB profile. ## Realtime inference The application supplies PIL-compatible frames with non-decreasing timestamps. Use `frame_queue_size=1` for the 24 GB realtime profile. ```python import time from PIL import Image session = model.create_realtime_session( processor, initial_prompt=( "Describe important changes in the video as they happen. " "Stay silent when there is no meaningful update." ), frame_queue_size=1, max_tokens_per_turn=12, max_new_tokens=4096, do_sample=False, ) frame_paths = [ "data/frame_0001.jpg", "data/frame_0002.jpg", "data/frame_0003.jpg", ] try: session.start() for index, frame_path in enumerate(frame_paths): image = Image.open(frame_path).convert("RGB") session.push_frame(image, timestamp=float(index)) while True: chunk = session.poll_output(timeout=0.0) if chunk is None: break print(chunk, end="", flush=True) time.sleep(1.0) session.push_prompt("What changed in the latest frames?") deadline = time.monotonic() + 5.0 while time.monotonic() < deadline: chunk = session.poll_output(timeout=0.1) if chunk is not None: print(chunk, end="", flush=True) finally: session.close() ``` One model instance supports one active realtime session. The model may emit control tokens such as `<|silence|>`, `<|round_start|>`, and `<|round_end|>`; applications should filter or render them according to their protocol. ## Offline video inference ```python text = model.offline_video_generate( processor, prompt="Describe this video.", video="data/example_video.mp4", shortest_edge=4096, longest_edge=16777216, video_max_pixels=201326592, patch_size=16, temporal_patch_size=1, merge_size=2, video_fps=1.0, min_frames=1, max_frames=256, num_extract_threads=4, image_mean=[0.5, 0.5, 0.5], image_std=[0.5, 0.5, 0.5], max_new_tokens=256, temperature=1.0, top_k=50, top_p=1.0, repetition_penalty=1.0, do_sample=False, vision_chunked_length=64, ) print(text) ``` ## Configuration files - `config.json`: model and NF4 weight configuration. - `generation_config.json`: HQQ KV8 configuration. - `modeling_moss_vl.py`: checkpoint-local MOSS-VL and QuantizedCache code.