Video-Text-to-Text
Transformers
Safetensors
English
Chinese
moss_vl
feature-extraction
MOSS-VL
realtime
streaming
video-understanding
bitsandbytes
NF4
quantized
custom_code
4-bit precision
Instructions to use OpenMOSS-Team/MOSS-VL-Realtime-NF4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OpenMOSS-Team/MOSS-VL-Realtime-NF4 with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("OpenMOSS-Team/MOSS-VL-Realtime-NF4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 4,443 Bytes
be02059 78cdbd3 be02059 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 | ---
license: apache-2.0
language:
- en
- zh
library_name: transformers
pipeline_tag: video-text-to-text
base_model: OpenMOSS-Team/MOSS-VL-Realtime
tags:
- MOSS-VL
- realtime
- streaming
- video-understanding
- bitsandbytes
- NF4
- quantized
- custom_code
---
<p align="center">
<img src="assets/logo.png" width="300" alt="MOSS-VL"/>
</p>
<p align="center">
<a href="https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-NF4/blob/main/README.md">English</a> | 中文
</p>
# MOSS-VL-Realtime W4A16 NF4 + KV8 HQQ
这是 MOSS-VL-Realtime 的 24 GB 显存量化版本,保留了原模型基于时间戳
的实时流式推理能力,同时支持标准的离线图片和视频推理接口。
## 量化方法
| 模块 | 格式 |
| --- | --- |
| 大部分语言模型层 | bitsandbytes NF4 4-bit 权重,启用 double quantization |
| 首尾语言模型层和多模态模块 | BF16 |
| 激活与计算 | BF16 |
| Transformers KV Cache | HQQ INT8 |
| Attention 后端 | FlashAttention 2 |
模型保留了首尾语言层和多模态模块的 BF16 精度,对主要语言模型层使用
NF4 权重量化,并通过 HQQ 将 KV Cache 压缩为 INT8。量化配置已包含在
checkpoint 中,加载时无需再次传入量化参数。
## 量化前后性能
在所列 benchmark 上,量化模型与未量化 BF16 模型的整体表现接近,
说明量化后模型能力基本保持,没有受到明显影响。
<p align="center">
<img src="assets/mossvl_quantization_benchmark_comparison_zh_4k.png" alt="MOSS-VL 量化前后 benchmark 对比" width="100%"/>
</p>
## 硬件要求
模型支持单张 24 GB 显存的 NVIDIA 消费级显卡。建议使用
FlashAttention 2,并将实时推理的 `frame_queue_size` 设为 1。
## 环境安装
```bash
git clone https://github.com/OpenMOSS/MOSS-VL.git
cd MOSS-VL
conda create -n moss_vl_quant python=3.12 pip -y
conda activate moss_vl_quant
pip install -i https://pypi.org/simple --no-build-isolation -r requirements.txt
pip install -i https://pypi.org/simple \
bitsandbytes==0.49.2 \
hqq==0.2.8.post1
python -m pip check
```
主要环境版本:
| 依赖 | 版本 |
| --- | --- |
| Python | 3.12 |
| PyTorch | 2.8.0 + CUDA 12.8 |
| Transformers | 4.57.1 |
| Accelerate | 1.12.0 |
| FlashAttention | 2.8.1 |
| bitsandbytes | 0.49.2 |
| HQQ | 0.2.8.post1 |
视频解码还需要确保 FFmpeg 已加入 `PATH`。
## 加载模型
```python
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
checkpoint = "OpenMOSS-Team/MOSS-VL-Realtime-NF4"
processor = AutoProcessor.from_pretrained(
checkpoint,
trust_remote_code=True,
frame_extract_num_threads=1,
)
model = AutoModelForCausalLM.from_pretrained(
checkpoint,
trust_remote_code=True,
device_map="auto",
torch_dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
)
model.eval()
```
`generation_config.json` 会自动启用 HQQ INT8 KV Cache,请不要使用
BF16 或 dynamic Cache 配置覆盖它。
## 实时流式推理
应用侧按时间顺序传入 PIL 图片和对应时间戳。24 GB 显存配置建议使用
`frame_queue_size=1`。
```python
import time
from PIL import Image
session = model.create_realtime_session(
processor,
initial_prompt="持续描述视频中的重要变化,没有明显变化时保持安静。",
frame_queue_size=1,
max_tokens_per_turn=12,
max_new_tokens=4096,
do_sample=False,
)
frame_paths = [
"data/frame_0001.jpg",
"data/frame_0002.jpg",
"data/frame_0003.jpg",
]
try:
session.start()
for index, frame_path in enumerate(frame_paths):
image = Image.open(frame_path).convert("RGB")
session.push_frame(image, timestamp=float(index))
while True:
chunk = session.poll_output(timeout=0.0)
if chunk is None:
break
print(chunk, end="", flush=True)
time.sleep(1.0)
finally:
session.close()
```
一个模型实例同时支持一个实时会话。模型可能输出 `<|silence|>`、
`<|round_start|>` 和 `<|round_end|>` 等控制 token,应用侧可以按需
过滤或渲染。
## 离线视频推理
```python
text = model.offline_video_generate(
processor,
prompt="请描述这段视频。",
video="data/example_video.mp4",
video_fps=1.0,
min_frames=1,
max_frames=256,
max_new_tokens=256,
do_sample=False,
vision_chunked_length=64,
)
print(text)
```
|