File size: 5,710 Bytes
404f9d8
 
2e600c2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
404f9d8
2e600c2
46a87a2
 
 
 
be02059
78cdbd3
be02059
 
2e600c2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2805586
 
be02059
 
 
2805586
 
 
 
 
2e600c2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
---
license: apache-2.0
language:
- en
- zh
library_name: transformers
pipeline_tag: video-text-to-text
base_model: OpenMOSS-Team/MOSS-VL-Realtime
tags:
- MOSS-VL
- realtime
- streaming
- video-understanding
- bitsandbytes
- NF4
- quantized
- custom_code
---

<p align="center">
  <img src="assets/logo.png" width="300" alt="MOSS-VL"/>
</p>

<p align="center">
  English | <a href="https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-NF4/blob/main/README_zh.md">中文</a>
</p>

# MOSS-VL-Realtime W4A16 NF4 + KV8 HQQ

This is the 24 GiB quantized release of MOSS-VL-Realtime. It keeps the original
timestamp-aware streaming interface and can also use the offline image/video
helpers from the standard checkpoint.

## Quantization profile

| Component | Format |
| --- | --- |
| Most language layers | bitsandbytes NF4 4-bit weights with double quantization |
| First/last language layers and multimodal modules | BF16 |
| Activations and compute | BF16 |
| Transformers KV cache | HQQ INT8 |
| Attention backend | FlashAttention 2 |

The checkpoint carries its bitsandbytes configuration, HQQ cache configuration,
and MOSS-VL remote modeling code. Load the directory directly; do not add a
second runtime quantization configuration.

## Quantization benchmark

Across the selected benchmarks, the quantized models remain close to their
non-quantized BF16 counterparts, showing that overall model quality is largely
preserved after quantization.

<p align="center">
  <img src="assets/mossvl_quantization_benchmark_comparison_en_4k.png" alt="MOSS-VL quantization benchmark comparison" width="100%"/>
</p>

## Hardware requirements

The model is designed to run on a single NVIDIA GPU with 24 GB of VRAM. Use
FlashAttention 2 and `frame_queue_size=1` for the 24 GB realtime profile.

## Environment

### Installation

Use the standard MOSS-VL repository requirements, then add the two quantization
backends required by this checkpoint:

```bash
git clone https://github.com/OpenMOSS/MOSS-VL.git
cd MOSS-VL

conda create -n moss_vl_quant python=3.12 pip -y
conda activate moss_vl_quant
pip install -i https://pypi.org/simple --no-build-isolation -r requirements.txt
pip install -i https://pypi.org/simple \
  bitsandbytes==0.49.2 \
  hqq==0.2.8.post1
python -m pip check
```

The standard release environment uses the following core stack:

| Package | Version |
| --- | --- |
| Python | 3.12 |
| PyTorch | 2.8.0 + CUDA 12.8 |
| Transformers | 4.57.1 |
| Accelerate | 1.12.0 |
| FlashAttention | 2.8.1 |
| bitsandbytes | 0.49.2 |
| HQQ | 0.2.8.post1 |

Video decoding also requires FFmpeg to be available in `PATH`.

## Load the model

Keep `attn_implementation` set to `flash_attention_2` for the 24 GB profile.

```python
import torch
from transformers import AutoModelForCausalLM, AutoProcessor

checkpoint = "/path/to/mossvl_streaming_w4a16_nf4_keep_first4_last4_kv8_hqq"

processor = AutoProcessor.from_pretrained(
    checkpoint,
    trust_remote_code=True,
    frame_extract_num_threads=1,
)
model = AutoModelForCausalLM.from_pretrained(
    checkpoint,
    trust_remote_code=True,
    device_map="auto",
    torch_dtype=torch.bfloat16,
    attn_implementation="flash_attention_2",
)
model.eval()
```

`generation_config.json` automatically enables the HQQ INT8 KV cache. Do not
override it with a BF16/dynamic cache when using the 24 GB profile.

## Realtime inference

The application supplies PIL-compatible frames with non-decreasing timestamps.
Use `frame_queue_size=1` for the 24 GB realtime profile.

```python
import time
from PIL import Image

session = model.create_realtime_session(
    processor,
    initial_prompt=(
        "Describe important changes in the video as they happen. "
        "Stay silent when there is no meaningful update."
    ),
    frame_queue_size=1,
    max_tokens_per_turn=12,
    max_new_tokens=4096,
    do_sample=False,
)

frame_paths = [
    "data/frame_0001.jpg",
    "data/frame_0002.jpg",
    "data/frame_0003.jpg",
]

try:
    session.start()
    for index, frame_path in enumerate(frame_paths):
        image = Image.open(frame_path).convert("RGB")
        session.push_frame(image, timestamp=float(index))

        while True:
            chunk = session.poll_output(timeout=0.0)
            if chunk is None:
                break
            print(chunk, end="", flush=True)

        time.sleep(1.0)

    session.push_prompt("What changed in the latest frames?")
    deadline = time.monotonic() + 5.0
    while time.monotonic() < deadline:
        chunk = session.poll_output(timeout=0.1)
        if chunk is not None:
            print(chunk, end="", flush=True)
finally:
    session.close()
```

One model instance supports one active realtime session. The model may emit
control tokens such as `<|silence|>`, `<|round_start|>`, and `<|round_end|>`;
applications should filter or render them according to their protocol.

## Offline video inference

```python
text = model.offline_video_generate(
    processor,
    prompt="Describe this video.",
    video="data/example_video.mp4",
    shortest_edge=4096,
    longest_edge=16777216,
    video_max_pixels=201326592,
    patch_size=16,
    temporal_patch_size=1,
    merge_size=2,
    video_fps=1.0,
    min_frames=1,
    max_frames=256,
    num_extract_threads=4,
    image_mean=[0.5, 0.5, 0.5],
    image_std=[0.5, 0.5, 0.5],
    max_new_tokens=256,
    temperature=1.0,
    top_k=50,
    top_p=1.0,
    repetition_penalty=1.0,
    do_sample=False,
    vision_chunked_length=64,
)
print(text)
```

## Configuration files

- `config.json`: model and NF4 weight configuration.
- `generation_config.json`: HQQ KV8 configuration.
- `modeling_moss_vl.py`: checkpoint-local MOSS-VL and QuantizedCache code.