Add model card and pipeline tag for VisCoP

#1
by nielsr HF Staff - opened
Files changed (1) hide show
  1. README.md +88 -0
README.md ADDED
@@ -0,0 +1,88 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ pipeline_tag: video-text-to-text
3
+ ---
4
+
5
+ # VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
6
+
7
+ This repository contains the model checkpoints for **VisCoP** (Vision Contextualized Probing), a parameter-efficient adaptation framework that augments Vision-Language Models (VLMs) with a compact set of learnable visual probes for robust domain adaptation under distribution shifts.
8
+
9
+ For more details, please refer to:
10
+ * **Paper:** [VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models](https://huggingface.co/papers/2510.13808)
11
+ * **GitHub Repository:** [dominickrei/VisCoP](https://github.com/dominickrei/VisCoP)
12
+
13
+ ---
14
+
15
+ ## ⚙️ Installation
16
+
17
+ To set up the environment, clone the official repository and install the dependencies:
18
+
19
+ ```shell
20
+ git clone https://github.com/dominickrei/VisCoP.git
21
+ cd VisCoP
22
+ pip install -r requirements.txt
23
+ pip install flash-attn --no-build-isolation
24
+ ```
25
+
26
+ ## 💻 Inference
27
+
28
+ You can run video-based inference using the following Python script:
29
+
30
+ ```python
31
+ from viscop import model_init, mm_infer
32
+ from viscop.mm_utils import load_video
33
+
34
+ ## Load model
35
+ model_path = 'dreilly/viscop-models' # Update with your local path or specific checkpoint subfolder
36
+
37
+ model, processor = model_init(
38
+ model_path=model_path,
39
+ device_map={"": "cuda"}
40
+ )
41
+
42
+ ## Load video
43
+ video_path = './assets/ego_cut_carrot.mp4'
44
+
45
+ frames, timestamps = load_video(video_path, fps=1, max_frames=180)
46
+
47
+ ## Create conversation
48
+ conversation = [
49
+ {
50
+ "role": "user",
51
+ "content": [
52
+ {"type": "video", "timestamps": timestamps, "num_frames": len(frames)},
53
+ {"type": "text", "text": "What vegetable is the person cutting in the video?"},
54
+ ]
55
+ }
56
+ ]
57
+
58
+ ## Perform inference
59
+ inputs = processor(
60
+ images=[frames],
61
+ text=conversation,
62
+ merge_size=2,
63
+ return_tensors="pt",
64
+ )
65
+
66
+ prediction = mm_infer(
67
+ inputs,
68
+ model=model,
69
+ tokenizer=processor.tokenizer,
70
+ do_sample=False,
71
+ modal='video'
72
+ )
73
+
74
+ print(prediction)
75
+ ```
76
+
77
+ ## Citation
78
+
79
+ If you find this work helpful, please consider citing our paper:
80
+
81
+ ```bibtex
82
+ @inproceedings{reilly2026viscop,
83
+ title = {VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models},
84
+ author = {Dominick Reilly and Manish Kumar Govind and Le Xue and Srijan Das},
85
+ booktitle = {Proceedings of the European Conference on Computer Vision (ECCV)},
86
+ year = {2026}
87
+ }
88
+ ```