File size: 3,443 Bytes
3208c6b
5ce5cc6
3a5870b
3208c6b
 
3a5870b
 
 
3208c6b
 
 
3a5870b
 
 
 
 
 
3208c6b
3a5870b
3208c6b
 
3a5870b
792a662
3a5870b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3208c6b
 
 
3a5870b
 
3208c6b
3a5870b
 
 
3208c6b
 
3a5870b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3208c6b
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
---
base_model: Qwen/Qwen3-VL-8B-Instruct
base_model_relation: adapter
library_name: peft
pipeline_tag: image-text-to-text
license: apache-2.0
language:
- en
tags:
- lora
- peft
- cvpd
- self-distillation
- multimodal
- vision-language
- lmm
- ocr
- qwen3-vl
- unsupervised
---

# CVPD: Contrastive Counterfactual Visual Process Distillation

This is the CVPD LoRA adapter for `Qwen/Qwen3-VL-8B-Instruct`, from our paper
[Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots](https://github.com/mbzuai-oryx/CVPD),
accepted to BMVC 2026.

CVPD is a fully self-contained framework for dense, token-level visual self-distillation.
Where prior visual distillation builds its privileged context from outside the model, using
segmentation systems, region proposals, or a stronger annotator, CVPD recovers that context
from the model's own counterfactual behavior. We train on raw, unlabeled images with no
captions, bounding boxes, labels, reward models, or teacher models, by locating **visual blind
spots**, regions the model can perceive but fails to exploit under full-image conditioning:

- **Discovery:** the model writes its own question and probe answer per image and proposes
  candidate regions by self-grounding and by 3x3 and 2x2 partitions. A region is kept when
  cropping to it moves and sharpens the answer distribution while blurring ("ghosting") it
  leaves the full-image behavior unchanged.
- **Contrastive self-distillation:** each retained region instantiates four policies from the
  same backbone. The online student learns from a crop-conditioned positive teacher, is pushed
  away from a ghost-conditioned negative teacher, and is anchored to a frozen reference policy.

We combine these as latent transfer plus contrastive ranking (`lambda_rank=0.5`, margin
`m=0.1`) and optimize under a KL anchor adapted online to a target divergence of `0.03`.

## Usage

This is a LoRA adapter, so load the base model first and attach the adapter:

```python
import torch
from PIL import Image
from transformers import AutoModelForImageTextToText, AutoProcessor
from peft import PeftModel

BASE = "Qwen/Qwen3-VL-8B-Instruct"
ADAPTER = "shravvvv/CVPD"

model = AutoModelForImageTextToText.from_pretrained(BASE, dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, ADAPTER)
processor = AutoProcessor.from_pretrained(ADAPTER)
model.eval()

image = Image.open("example.jpg").convert("RGB")
messages = [{"role": "user", "content": [
    {"type": "image", "image": image},
    {"type": "text", "text": "What is the text written on the sign?"},
]}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)

with torch.inference_mode():
    out = model.generate(**inputs, max_new_tokens=128)
print(processor.batch_decode(out[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])
```

## License

Apache 2.0.

## Citation

```bibtex
@inproceedings{cvpd,
author    = {Venkatraman, Shravan and Thawakar, Omkar and Thawkar, Ritesh and Shaker, Abdelrahman and Muhammad, Rao},
title     = {Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots},
booktitle = {37th British Machine Vision Conference 2026, {BMVC} 2026, Lancaster, UK, November 23-26, 2026},
publisher = {BMVA},
year      = {2026}
}
```