File size: 4,443 Bytes
be02059
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
78cdbd3
be02059
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
---
license: apache-2.0
language:
- en
- zh
library_name: transformers
pipeline_tag: video-text-to-text
base_model: OpenMOSS-Team/MOSS-VL-Realtime
tags:
- MOSS-VL
- realtime
- streaming
- video-understanding
- bitsandbytes
- NF4
- quantized
- custom_code
---

<p align="center">
  <img src="assets/logo.png" width="300" alt="MOSS-VL"/>
</p>

<p align="center">
  <a href="https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-NF4/blob/main/README.md">English</a> | 中文
</p>

# MOSS-VL-Realtime W4A16 NF4 + KV8 HQQ

这是 MOSS-VL-Realtime 的 24 GB 显存量化版本,保留了原模型基于时间戳
的实时流式推理能力,同时支持标准的离线图片和视频推理接口。

## 量化方法

| 模块 | 格式 |
| --- | --- |
| 大部分语言模型层 | bitsandbytes NF4 4-bit 权重,启用 double quantization |
| 首尾语言模型层和多模态模块 | BF16 |
| 激活与计算 | BF16 |
| Transformers KV Cache | HQQ INT8 |
| Attention 后端 | FlashAttention 2 |

模型保留了首尾语言层和多模态模块的 BF16 精度,对主要语言模型层使用
NF4 权重量化,并通过 HQQ 将 KV Cache 压缩为 INT8。量化配置已包含在
checkpoint 中,加载时无需再次传入量化参数。

## 量化前后性能

在所列 benchmark 上,量化模型与未量化 BF16 模型的整体表现接近,
说明量化后模型能力基本保持,没有受到明显影响。

<p align="center">
  <img src="assets/mossvl_quantization_benchmark_comparison_zh_4k.png" alt="MOSS-VL 量化前后 benchmark 对比" width="100%"/>
</p>

## 硬件要求

模型支持单张 24 GB 显存的 NVIDIA 消费级显卡。建议使用
FlashAttention 2,并将实时推理的 `frame_queue_size` 设为 1。

## 环境安装

```bash
git clone https://github.com/OpenMOSS/MOSS-VL.git
cd MOSS-VL

conda create -n moss_vl_quant python=3.12 pip -y
conda activate moss_vl_quant
pip install -i https://pypi.org/simple --no-build-isolation -r requirements.txt
pip install -i https://pypi.org/simple \
  bitsandbytes==0.49.2 \
  hqq==0.2.8.post1
python -m pip check
```

主要环境版本:

| 依赖 | 版本 |
| --- | --- |
| Python | 3.12 |
| PyTorch | 2.8.0 + CUDA 12.8 |
| Transformers | 4.57.1 |
| Accelerate | 1.12.0 |
| FlashAttention | 2.8.1 |
| bitsandbytes | 0.49.2 |
| HQQ | 0.2.8.post1 |

视频解码还需要确保 FFmpeg 已加入 `PATH`## 加载模型

```python
import torch
from transformers import AutoModelForCausalLM, AutoProcessor

checkpoint = "OpenMOSS-Team/MOSS-VL-Realtime-NF4"

processor = AutoProcessor.from_pretrained(
    checkpoint,
    trust_remote_code=True,
    frame_extract_num_threads=1,
)
model = AutoModelForCausalLM.from_pretrained(
    checkpoint,
    trust_remote_code=True,
    device_map="auto",
    torch_dtype=torch.bfloat16,
    attn_implementation="flash_attention_2",
)
model.eval()
```

`generation_config.json` 会自动启用 HQQ INT8 KV Cache,请不要使用
BF16 或 dynamic Cache 配置覆盖它。

## 实时流式推理

应用侧按时间顺序传入 PIL 图片和对应时间戳。24 GB 显存配置建议使用
`frame_queue_size=1````python
import time
from PIL import Image

session = model.create_realtime_session(
    processor,
    initial_prompt="持续描述视频中的重要变化,没有明显变化时保持安静。",
    frame_queue_size=1,
    max_tokens_per_turn=12,
    max_new_tokens=4096,
    do_sample=False,
)

frame_paths = [
    "data/frame_0001.jpg",
    "data/frame_0002.jpg",
    "data/frame_0003.jpg",
]

try:
    session.start()
    for index, frame_path in enumerate(frame_paths):
        image = Image.open(frame_path).convert("RGB")
        session.push_frame(image, timestamp=float(index))

        while True:
            chunk = session.poll_output(timeout=0.0)
            if chunk is None:
                break
            print(chunk, end="", flush=True)

        time.sleep(1.0)
finally:
    session.close()
```

一个模型实例同时支持一个实时会话。模型可能输出 `<|silence|>``<|round_start|>``<|round_end|>` 等控制 token,应用侧可以按需
过滤或渲染。

## 离线视频推理

```python
text = model.offline_video_generate(
    processor,
    prompt="请描述这段视频。",
    video="data/example_video.mp4",
    video_fps=1.0,
    min_frames=1,
    max_frames=256,
    max_new_tokens=256,
    do_sample=False,
    vision_chunked_length=64,
)
print(text)
```