How to use from the
Use from the
PEFT library
from peft import PeftModel
from transformers import AutoModelForCausalLM

base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-VL-8B-Instruct")
model = PeftModel.from_pretrained(base_model, "shravvvv/CVPD")

CVPD: Contrastive Counterfactual Visual Process Distillation

This is the CVPD LoRA adapter for Qwen/Qwen3-VL-8B-Instruct, from our paper Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots, accepted to BMVC 2026.

CVPD is a fully self-contained framework for dense, token-level visual self-distillation. Where prior visual distillation builds its privileged context from outside the model, using segmentation systems, region proposals, or a stronger annotator, CVPD recovers that context from the model's own counterfactual behavior. We train on raw, unlabeled images with no captions, bounding boxes, labels, reward models, or teacher models, by locating visual blind spots, regions the model can perceive but fails to exploit under full-image conditioning:

  • Discovery: the model writes its own question and probe answer per image and proposes candidate regions by self-grounding and by 3x3 and 2x2 partitions. A region is kept when cropping to it moves and sharpens the answer distribution while blurring ("ghosting") it leaves the full-image behavior unchanged.
  • Contrastive self-distillation: each retained region instantiates four policies from the same backbone. The online student learns from a crop-conditioned positive teacher, is pushed away from a ghost-conditioned negative teacher, and is anchored to a frozen reference policy.

We combine these as latent transfer plus contrastive ranking (lambda_rank=0.5, margin m=0.1) and optimize under a KL anchor adapted online to a target divergence of 0.03.

Usage

This is a LoRA adapter, so load the base model first and attach the adapter:

import torch
from PIL import Image
from transformers import AutoModelForImageTextToText, AutoProcessor
from peft import PeftModel

BASE = "Qwen/Qwen3-VL-8B-Instruct"
ADAPTER = "shravvvv/CVPD"

model = AutoModelForImageTextToText.from_pretrained(BASE, dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, ADAPTER)
processor = AutoProcessor.from_pretrained(ADAPTER)
model.eval()

image = Image.open("example.jpg").convert("RGB")
messages = [{"role": "user", "content": [
    {"type": "image", "image": image},
    {"type": "text", "text": "What is the text written on the sign?"},
]}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)

with torch.inference_mode():
    out = model.generate(**inputs, max_new_tokens=128)
print(processor.batch_decode(out[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0])

License

Apache 2.0.

Citation

@inproceedings{cvpd,
author    = {Venkatraman, Shravan and Thawakar, Omkar and Thawkar, Ritesh and Shaker, Abdelrahman and Muhammad, Rao},
title     = {Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots},
booktitle = {37th British Machine Vision Conference 2026, {BMVC} 2026, Lancaster, UK, November 23-26, 2026},
publisher = {BMVA},
year      = {2026}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for shravvvv/CVPD

Adapter
(161)
this model