CCCCyx's picture
Fix language switch links
78cdbd3 verified
|
Raw
History Blame Contribute Delete
5.71 kB
---
license: apache-2.0
language:
- en
- zh
library_name: transformers
pipeline_tag: video-text-to-text
base_model: OpenMOSS-Team/MOSS-VL-Realtime
tags:
- MOSS-VL
- realtime
- streaming
- video-understanding
- bitsandbytes
- NF4
- quantized
- custom_code
---
<p align="center">
<img src="assets/logo.png" width="300" alt="MOSS-VL"/>
</p>
<p align="center">
English | <a href="https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-NF4/blob/main/README_zh.md">中文</a>
</p>
# MOSS-VL-Realtime W4A16 NF4 + KV8 HQQ
This is the 24 GiB quantized release of MOSS-VL-Realtime. It keeps the original
timestamp-aware streaming interface and can also use the offline image/video
helpers from the standard checkpoint.
## Quantization profile
| Component | Format |
| --- | --- |
| Most language layers | bitsandbytes NF4 4-bit weights with double quantization |
| First/last language layers and multimodal modules | BF16 |
| Activations and compute | BF16 |
| Transformers KV cache | HQQ INT8 |
| Attention backend | FlashAttention 2 |
The checkpoint carries its bitsandbytes configuration, HQQ cache configuration,
and MOSS-VL remote modeling code. Load the directory directly; do not add a
second runtime quantization configuration.
## Quantization benchmark
Across the selected benchmarks, the quantized models remain close to their
non-quantized BF16 counterparts, showing that overall model quality is largely
preserved after quantization.
<p align="center">
<img src="assets/mossvl_quantization_benchmark_comparison_en_4k.png" alt="MOSS-VL quantization benchmark comparison" width="100%"/>
</p>
## Hardware requirements
The model is designed to run on a single NVIDIA GPU with 24 GB of VRAM. Use
FlashAttention 2 and `frame_queue_size=1` for the 24 GB realtime profile.
## Environment
### Installation
Use the standard MOSS-VL repository requirements, then add the two quantization
backends required by this checkpoint:
```bash
git clone https://github.com/OpenMOSS/MOSS-VL.git
cd MOSS-VL
conda create -n moss_vl_quant python=3.12 pip -y
conda activate moss_vl_quant
pip install -i https://pypi.org/simple --no-build-isolation -r requirements.txt
pip install -i https://pypi.org/simple \
bitsandbytes==0.49.2 \
hqq==0.2.8.post1
python -m pip check
```
The standard release environment uses the following core stack:
| Package | Version |
| --- | --- |
| Python | 3.12 |
| PyTorch | 2.8.0 + CUDA 12.8 |
| Transformers | 4.57.1 |
| Accelerate | 1.12.0 |
| FlashAttention | 2.8.1 |
| bitsandbytes | 0.49.2 |
| HQQ | 0.2.8.post1 |
Video decoding also requires FFmpeg to be available in `PATH`.
## Load the model
Keep `attn_implementation` set to `flash_attention_2` for the 24 GB profile.
```python
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
checkpoint = "/path/to/mossvl_streaming_w4a16_nf4_keep_first4_last4_kv8_hqq"
processor = AutoProcessor.from_pretrained(
checkpoint,
trust_remote_code=True,
frame_extract_num_threads=1,
)
model = AutoModelForCausalLM.from_pretrained(
checkpoint,
trust_remote_code=True,
device_map="auto",
torch_dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
)
model.eval()
```
`generation_config.json` automatically enables the HQQ INT8 KV cache. Do not
override it with a BF16/dynamic cache when using the 24 GB profile.
## Realtime inference
The application supplies PIL-compatible frames with non-decreasing timestamps.
Use `frame_queue_size=1` for the 24 GB realtime profile.
```python
import time
from PIL import Image
session = model.create_realtime_session(
processor,
initial_prompt=(
"Describe important changes in the video as they happen. "
"Stay silent when there is no meaningful update."
),
frame_queue_size=1,
max_tokens_per_turn=12,
max_new_tokens=4096,
do_sample=False,
)
frame_paths = [
"data/frame_0001.jpg",
"data/frame_0002.jpg",
"data/frame_0003.jpg",
]
try:
session.start()
for index, frame_path in enumerate(frame_paths):
image = Image.open(frame_path).convert("RGB")
session.push_frame(image, timestamp=float(index))
while True:
chunk = session.poll_output(timeout=0.0)
if chunk is None:
break
print(chunk, end="", flush=True)
time.sleep(1.0)
session.push_prompt("What changed in the latest frames?")
deadline = time.monotonic() + 5.0
while time.monotonic() < deadline:
chunk = session.poll_output(timeout=0.1)
if chunk is not None:
print(chunk, end="", flush=True)
finally:
session.close()
```
One model instance supports one active realtime session. The model may emit
control tokens such as `<|silence|>`, `<|round_start|>`, and `<|round_end|>`;
applications should filter or render them according to their protocol.
## Offline video inference
```python
text = model.offline_video_generate(
processor,
prompt="Describe this video.",
video="data/example_video.mp4",
shortest_edge=4096,
longest_edge=16777216,
video_max_pixels=201326592,
patch_size=16,
temporal_patch_size=1,
merge_size=2,
video_fps=1.0,
min_frames=1,
max_frames=256,
num_extract_threads=4,
image_mean=[0.5, 0.5, 0.5],
image_std=[0.5, 0.5, 0.5],
max_new_tokens=256,
temperature=1.0,
top_k=50,
top_p=1.0,
repetition_penalty=1.0,
do_sample=False,
vision_chunked_length=64,
)
print(text)
```
## Configuration files
- `config.json`: model and NF4 weight configuration.
- `generation_config.json`: HQQ KV8 configuration.
- `modeling_moss_vl.py`: checkpoint-local MOSS-VL and QuantizedCache code.