Video-Text-to-Text
Transformers
Safetensors
English
Chinese
moss_vl
feature-extraction
MOSS-VL
realtime
streaming
video-understanding
bitsandbytes
NF4
quantized
custom_code
4-bit precision
Instructions to use OpenMOSS-Team/MOSS-VL-Realtime-NF4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OpenMOSS-Team/MOSS-VL-Realtime-NF4 with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("OpenMOSS-Team/MOSS-VL-Realtime-NF4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 5,710 Bytes
404f9d8 2e600c2 404f9d8 2e600c2 46a87a2 be02059 78cdbd3 be02059 2e600c2 2805586 be02059 2805586 2e600c2 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 | ---
license: apache-2.0
language:
- en
- zh
library_name: transformers
pipeline_tag: video-text-to-text
base_model: OpenMOSS-Team/MOSS-VL-Realtime
tags:
- MOSS-VL
- realtime
- streaming
- video-understanding
- bitsandbytes
- NF4
- quantized
- custom_code
---
<p align="center">
<img src="assets/logo.png" width="300" alt="MOSS-VL"/>
</p>
<p align="center">
English | <a href="https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-NF4/blob/main/README_zh.md">中文</a>
</p>
# MOSS-VL-Realtime W4A16 NF4 + KV8 HQQ
This is the 24 GiB quantized release of MOSS-VL-Realtime. It keeps the original
timestamp-aware streaming interface and can also use the offline image/video
helpers from the standard checkpoint.
## Quantization profile
| Component | Format |
| --- | --- |
| Most language layers | bitsandbytes NF4 4-bit weights with double quantization |
| First/last language layers and multimodal modules | BF16 |
| Activations and compute | BF16 |
| Transformers KV cache | HQQ INT8 |
| Attention backend | FlashAttention 2 |
The checkpoint carries its bitsandbytes configuration, HQQ cache configuration,
and MOSS-VL remote modeling code. Load the directory directly; do not add a
second runtime quantization configuration.
## Quantization benchmark
Across the selected benchmarks, the quantized models remain close to their
non-quantized BF16 counterparts, showing that overall model quality is largely
preserved after quantization.
<p align="center">
<img src="assets/mossvl_quantization_benchmark_comparison_en_4k.png" alt="MOSS-VL quantization benchmark comparison" width="100%"/>
</p>
## Hardware requirements
The model is designed to run on a single NVIDIA GPU with 24 GB of VRAM. Use
FlashAttention 2 and `frame_queue_size=1` for the 24 GB realtime profile.
## Environment
### Installation
Use the standard MOSS-VL repository requirements, then add the two quantization
backends required by this checkpoint:
```bash
git clone https://github.com/OpenMOSS/MOSS-VL.git
cd MOSS-VL
conda create -n moss_vl_quant python=3.12 pip -y
conda activate moss_vl_quant
pip install -i https://pypi.org/simple --no-build-isolation -r requirements.txt
pip install -i https://pypi.org/simple \
bitsandbytes==0.49.2 \
hqq==0.2.8.post1
python -m pip check
```
The standard release environment uses the following core stack:
| Package | Version |
| --- | --- |
| Python | 3.12 |
| PyTorch | 2.8.0 + CUDA 12.8 |
| Transformers | 4.57.1 |
| Accelerate | 1.12.0 |
| FlashAttention | 2.8.1 |
| bitsandbytes | 0.49.2 |
| HQQ | 0.2.8.post1 |
Video decoding also requires FFmpeg to be available in `PATH`.
## Load the model
Keep `attn_implementation` set to `flash_attention_2` for the 24 GB profile.
```python
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
checkpoint = "/path/to/mossvl_streaming_w4a16_nf4_keep_first4_last4_kv8_hqq"
processor = AutoProcessor.from_pretrained(
checkpoint,
trust_remote_code=True,
frame_extract_num_threads=1,
)
model = AutoModelForCausalLM.from_pretrained(
checkpoint,
trust_remote_code=True,
device_map="auto",
torch_dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
)
model.eval()
```
`generation_config.json` automatically enables the HQQ INT8 KV cache. Do not
override it with a BF16/dynamic cache when using the 24 GB profile.
## Realtime inference
The application supplies PIL-compatible frames with non-decreasing timestamps.
Use `frame_queue_size=1` for the 24 GB realtime profile.
```python
import time
from PIL import Image
session = model.create_realtime_session(
processor,
initial_prompt=(
"Describe important changes in the video as they happen. "
"Stay silent when there is no meaningful update."
),
frame_queue_size=1,
max_tokens_per_turn=12,
max_new_tokens=4096,
do_sample=False,
)
frame_paths = [
"data/frame_0001.jpg",
"data/frame_0002.jpg",
"data/frame_0003.jpg",
]
try:
session.start()
for index, frame_path in enumerate(frame_paths):
image = Image.open(frame_path).convert("RGB")
session.push_frame(image, timestamp=float(index))
while True:
chunk = session.poll_output(timeout=0.0)
if chunk is None:
break
print(chunk, end="", flush=True)
time.sleep(1.0)
session.push_prompt("What changed in the latest frames?")
deadline = time.monotonic() + 5.0
while time.monotonic() < deadline:
chunk = session.poll_output(timeout=0.1)
if chunk is not None:
print(chunk, end="", flush=True)
finally:
session.close()
```
One model instance supports one active realtime session. The model may emit
control tokens such as `<|silence|>`, `<|round_start|>`, and `<|round_end|>`;
applications should filter or render them according to their protocol.
## Offline video inference
```python
text = model.offline_video_generate(
processor,
prompt="Describe this video.",
video="data/example_video.mp4",
shortest_edge=4096,
longest_edge=16777216,
video_max_pixels=201326592,
patch_size=16,
temporal_patch_size=1,
merge_size=2,
video_fps=1.0,
min_frames=1,
max_frames=256,
num_extract_threads=4,
image_mean=[0.5, 0.5, 0.5],
image_std=[0.5, 0.5, 0.5],
max_new_tokens=256,
temperature=1.0,
top_k=50,
top_p=1.0,
repetition_penalty=1.0,
do_sample=False,
vision_chunked_length=64,
)
print(text)
```
## Configuration files
- `config.json`: model and NF4 weight configuration.
- `generation_config.json`: HQQ KV8 configuration.
- `modeling_moss_vl.py`: checkpoint-local MOSS-VL and QuantizedCache code.
|