File size: 2,939 Bytes
26cbd31
 
b7f0f2b
 
 
ece3b11
 
 
 
26cbd31
ece3b11
b7f0f2b
ece3b11
 
 
 
 
 
 
 
b7f0f2b
 
 
 
ece3b11
 
 
 
 
 
 
 
 
b7f0f2b
 
ece3b11
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b7f0f2b
ece3b11
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
---
license: mit
base_model:
- Qwen/Qwen2.5-VL-3B-Instruct
- Qwen/Qwen2.5-VL-7B-Instruct
tags:
- vision-language-model
- token-compression
- multimodal
---

# CRISPR — Checkpoints

Checkpoints for **CRISPR: Context-Refined Information Spatial Pooling with
Region-awareness for Efficient Visual Token Compression in VLMs**, accepted
at ACM MM 2026.

- Code: https://github.com/ZuyiZhou/CRISPR
- Paper DOI: https://doi.org/10.1145/3767308.3835007

This repo hosts CRISPR checkpoints for the Qwen2.5-VL-3B-Instruct backbone
(9x and 16x compression) and the Qwen2.5-VL-7B-Instruct backbone (16x
compression). The 7B/9x checkpoint is not currently available (lost prior to
this release) and is not planned unless retraining happens in the future.

## Files

| Path | Compression ratio | Backbone | Notes |
|---|---|---|---|
| `3b_9x/checkpoint.pt` | 9x (3x3 block) | Qwen2.5-VL-3B-Instruct | Stage-2, best checkpoint by val loss |
| `3b_9x/config.json` | | | training config used to produce this checkpoint |
| `3b_16x/checkpoint.pt` | 16x (4x4 block) | Qwen2.5-VL-3B-Instruct | Stage-2, best checkpoint by val loss |
| `3b_16x/config.json` | | | training config used to produce this checkpoint |
| `7b_16x/checkpoint.pt` | 16x (4x4 block) | Qwen2.5-VL-7B-Instruct | Stage-2, best checkpoint by val loss |
| `7b_16x/config.json` | | | training config used to produce this checkpoint |

Each `checkpoint.pt` is a plain `torch.save` dict with keys `config`,
`token_mixer` (TokenMixer state dict), and `local_c3` (LocalC3 state dict,
which also contains the Global Token Fusion sub-module). Only the trainable
CRISPR modules are included — the frozen Qwen2.5-VL vision encoder and
decoder weights are not part of this checkpoint and must be obtained
separately from [Qwen2.5-VL](https://github.com/QwenLM/Qwen2.5-VL). Optimizer/
scheduler state is not included (only the model weights needed for inference
or further fine-tuning are provided).

## Usage

```python
from crispr import create_model_v7

model = create_model_v7(decoder_path="./Qwen/Qwen2.5-VL-3B-Instruct")
model.load_checkpoint("3b_9x/checkpoint.pt")  # see crispr/model_v7.py for the loader
# for the 7B backbone: decoder_path="./Qwen/Qwen2.5-VL-7B-Instruct", checkpoint="7b_16x/checkpoint.pt"
```

See the main repository (https://github.com/ZuyiZhou/CRISPR) for the model
code, training script, and evaluation scripts.

## Citation

```bibtex
@inproceedings{zhou2026crispr,
  author    = {Zhou, Zuyi and Xue, Dizhan and Qian, Shengsheng and Xu, Changsheng},
  title     = {CRISPR: Context-Refined Information Spatial Pooling with
               Region-awareness for Efficient Visual Token Compression in VLMs},
  booktitle = {Proceedings of the 34th ACM International Conference on
               Multimedia (MM '26)},
  year      = {2026},
  publisher = {Association for Computing Machinery},
  address   = {New York, NY, USA},
  doi       = {10.1145/3767308.3835007}
}
```