File size: 1,764 Bytes
7fbf753
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
---
license: mit
tags:
- medical
- chest-xray
- visual-question-answering
- difference-vqa
- mimic-cxr
library_name: pytorch
---

# Diff-VQA — Qwen3-0.6B visual-prefix decoder (multiobjective-regsteps encoder)

Online-trained **Difference Visual Question Answering** head for chest X-rays
(MIMIC-CXR / Medical-Diff-VQA). A frozen ViT-L/14 vision encoder produces patch
tokens for a current + reference image pair; a **Qwen3-0.6B** decoder,
conditioned on those tokens as a visual prefix, generates the answer describing
what changed.

This is a **slimmed inference checkpoint**: decoder weights + the frozen vision
encoder weights + architecture metadata (optimizer/scheduler state stripped).

## Contents (`*_best.pt`, `torch.load(..., weights_only=False)`)

| key | description |
|-----|-------------|
| `state_dict` | `DiffVQAQwen3Head` weights (Qwen3-0.6B LLM + `vis_proj` + `frame_emb`) |
| `encoder_state` | frozen ViT-L/14 vision encoder (`objective: multiobjective_caption+regioncontrastive`) |
| `decoder` | `Qwen/Qwen3-0.6B` |
| `tokenizer_name` | `Qwen/Qwen3-0.6B` |
| `vis_dim` | 1024 |
| `encoder_name` | `pretrain-multiobjective-regsteps` |
| `epoch` / `step` / `best_val` | training position + best validation loss |

## Training

- Data: Medical-Diff-VQA `difference` questions over MIMIC-CXR image pairs
  (~100k train / ~16k val / ~16k test).
- Objective: next-token cross-entropy over the answer span (question tokens masked).
- Vision encoder is **frozen**; only the decoder (LLM + visual adapter) is trained.

## Notes

The bundled `encoder_state` is the exact frozen encoder this decoder was trained
against — load both together so the visual features match the space `vis_proj`
expects (a mismatched encoder produces degenerate output).