CCCCyx's picture
Update final v3 quantization benchmarks
641be01 verified
|
Raw
History Blame Contribute Delete
6.57 kB
---
license: apache-2.0
language:
- en
- zh
library_name: transformers
pipeline_tag: video-text-to-text
base_model: OpenMOSS-Team/MOSS-VL-Instruct-0708
tags:
- MOSS-VL
- image-understanding
- video-understanding
- bitsandbytes
- NF4
- quantized
- custom_code
---
<p align="center">
<img src="assets/logo.png" width="300" alt="MOSS-VL"/>
</p>
<p align="center">
English | <a href="https://huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct-0708-NF4/blob/main/README_zh.md">中文</a>
</p>
# MOSS-VL-Instruct-0708 W4A16 NF4
This is the Transformers NF4 release of
[MOSS-VL-Instruct-0708](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct-0708).
It supports image and video inference through the standard MOSS-VL offline
inference path. This checkpoint is not an SGLang release.
## Architecture
<p align="center">
<img src="assets/architecture.png" alt="MOSS-VL architecture" width="100%"/>
</p>
## Quantization profile
| Component | Format |
| --- | --- |
| 240 eligible Linear layers in language layers 4-43 | bitsandbytes NF4 weight-only quantization with double quantization and BF16 compute |
| First four and last four language layers | BF16 |
| Cross-attention projection modules | BF16 |
| Vision encoder and merger | BF16 |
| Embeddings, norms and `lm_head` | BF16 |
| Transformers KV cache | BF16 |
| Attention backend | FlashAttention 2 |
The checkpoint carries its bitsandbytes configuration. Load it directly and
do not add a second runtime quantization configuration. This variant does not
enable HQQ KV8; `generation_config.json` uses the standard BF16 KV cache.
## Quantization benchmark
The final evaluation compares the original BF16 model with all four release
profiles on their corresponding benchmark suites. This offline NF4 checkpoint
scores 89.53 on DocVQA, 67.30 on VideoMME, 75.86 on MLVU_dev, 51.00/48.17/59.33
on the three TimeLens subsets, and 61.76 on VSIBench.
<p align="center">
<img src="assets/mossvl_quantization_benchmark_comparison_final_v3_en_4k.png" alt="MOSS-VL quantization benchmark comparison" width="100%"/>
</p>
## Hardware requirements
The validated image test peaked at 12,494 MiB of process VRAM. The 1 FPS,
maximum-32-frame video test peaked at 16,708 MiB. A single NVIDIA GPU with
24 GB of VRAM is sufficient for the validated profile.
## Environment
### Installation
```bash
git clone https://github.com/OpenMOSS/MOSS-VL.git
cd MOSS-VL
conda create -n moss_vl_quant python=3.12 pip -y
conda activate moss_vl_quant
pip install -i https://pypi.org/simple --no-build-isolation -r requirements.txt
pip install -i https://pypi.org/simple bitsandbytes==0.49.2
python -m pip check
```
Validated core versions:
| Package | Version |
| --- | --- |
| Python | 3.12.8 |
| PyTorch | 2.8.0 + CUDA 12.8 |
| Transformers | 4.57.1 |
| Accelerate | 1.12.0 |
| FlashAttention | 2.8.1 |
| bitsandbytes | 0.49.2 |
Video decoding also requires FFmpeg in `PATH`.
### Load the model
```python
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
checkpoint = "OpenMOSS-Team/MOSS-VL-Instruct-0708-NF4"
processor = AutoProcessor.from_pretrained(
checkpoint,
trust_remote_code=True,
frame_extract_num_threads=1,
)
model = AutoModelForCausalLM.from_pretrained(
checkpoint,
trust_remote_code=True,
device_map="auto",
torch_dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
)
model.eval()
```
## Image inference
```python
text = model.offline_image_generate(
processor,
prompt="Describe this image.",
image="data/example_image.jpg",
shortest_edge=4096,
longest_edge=16777216,
multi_image_max_pixels=201326592,
patch_size=16,
temporal_patch_size=1,
merge_size=2,
image_mean=[0.5, 0.5, 0.5],
image_std=[0.5, 0.5, 0.5],
max_new_tokens=256,
do_sample=False,
vision_chunked_length=64,
)
print(text)
```
## Video inference
```python
text = model.offline_video_generate(
processor,
prompt="Describe this video.",
video="data/example_video.mp4",
shortest_edge=4096,
longest_edge=16777216,
video_max_pixels=201326592,
patch_size=16,
temporal_patch_size=1,
merge_size=2,
video_fps=1.0,
min_frames=1,
max_frames=32,
num_extract_threads=4,
image_mean=[0.5, 0.5, 0.5],
image_std=[0.5, 0.5, 0.5],
max_new_tokens=256,
do_sample=False,
vision_chunked_length=64,
)
print(text)
```
## Validated reproduction
The official runner passed both the receipt image and the 1 FPS Starbucks
video tests:
```bash
source /inspire/qb-ilm/project/video-understanding/public/train/moss_vl_streaming/8B/quant/venv/bin/activate
/inspire/qb-ilm/project/video-understanding/public/train/moss_vl_streaming/8B/quant/venv/bin/python \
/inspire/qb-ilm/project/video-understanding/public/train/moss_vl_streaming/8B/final_release/mossvl-github/MOSS-VL/inference/run_inference.py \
--checkpoint /inspire/qb-ilm/project/video-understanding/public/train/moss_vl_streaming/8B/final_release/quant/MOSS-VL-0708-Offline-NF4-Keep4-KV16 \
--mode image \
--input /inspire/qb-ilm/project/video-understanding/public/train/moss_vl_streaming/8B/quant/transformers_validation_20260811/inputs/offline_image.json \
--output /inspire/qb-ilm/project/video-understanding/public/train/moss_vl_streaming/8B/quant/transformers_validation_20260811/results/offline_nf4_image_output.json \
--timeout-seconds 300
/inspire/qb-ilm/project/video-understanding/public/train/moss_vl_streaming/8B/quant/venv/bin/python \
/inspire/qb-ilm/project/video-understanding/public/train/moss_vl_streaming/8B/final_release/mossvl-github/MOSS-VL/inference/run_inference.py \
--checkpoint /inspire/qb-ilm/project/video-understanding/public/train/moss_vl_streaming/8B/final_release/quant/MOSS-VL-0708-Offline-NF4-Keep4-KV16 \
--mode video \
--input /inspire/qb-ilm/project/video-understanding/public/train/moss_vl_streaming/8B/quant/transformers_validation_20260811/inputs/offline_video.json \
--output /inspire/qb-ilm/project/video-understanding/public/train/moss_vl_streaming/8B/quant/transformers_validation_20260811/results/offline_nf4_video_output.json \
--timeout-seconds 300
```
Full inputs, commands and raw results:
```text
/inspire/qb-ilm/project/video-understanding/public/train/moss_vl_streaming/8B/quant/transformers_validation_20260811
```
## Configuration files
- `config.json`: model and bitsandbytes NF4 configuration.
- `generation_config.json`: standard generation settings with BF16 KV cache.
- `modeling_moss_vl.py`: checkpoint-local offline MOSS-VL code.