tinyvlm-vqa / README.md
dhruvpatel93's picture
TinyVLM projector + LoRA adapter
e08de1e verified
|
Raw
History Blame Contribute Delete
2.17 kB
---
license: apache-2.0
base_model: Qwen/Qwen2.5-0.5B-Instruct
tags: [vision-language-model, vqa, image-captioning, lora, siglip2, qwen2.5]
---
# tinyvlm-vqa
A compact LLaVA-style VLM: SigLIP2 vision encoder -> pixel-shuffle -> MLP
projector -> Qwen2.5-0.5B-Instruct with LoRA.
| | |
|---|---|
| Vision encoder | `google/siglip2-base-patch16-224` (frozen) |
| LLM | `Qwen/Qwen2.5-0.5B-Instruct` (LoRA: q,k,v,o,gate,up,down) |
| Visual tokens | 49 per image (14x14 patches, 2x2 pixel-shuffle) |
| Image placeholder | `<|image_pad|>` -- an existing Qwen token, so no vocab resize |
| Stage 1 | projector alignment on Flickr8k + Flickr30k captions |
| Stage 2 | LoRA instruction tuning on VQAv2 + LLaVA-ReCap + captions |
## Answer length is prompt-controlled
Stage 2 mixes one-word VQA answers with multi-sentence descriptions, and VQA
examples carry an explicit hint. Append it for short answers, omit it for prose:
```
What color is the bus?
Answer the question using a single word or phrase. -> "red"
What color is the bus? -> "The bus is red."
```
## Files
- `projector.pt` -- projector weights (`pool_stride` is recoverable from the tensor shapes)
- `lora_adapter/` -- PEFT adapter for the LLM
## Usage
Needs the companion code from the training repo (`Model/model.py`,
`Model/dataset.py`, `Inference/inference.py`):
```python
from huggingface_hub import hf_hub_download, snapshot_download
from inference import load_model, answer
projector = hf_hub_download(repo_id="dhruvpatel93/tinyvlm-vqa", filename="projector.pt")
lora = snapshot_download(repo_id="dhruvpatel93/tinyvlm-vqa", allow_patterns=["lora_adapter/*"]) + "/lora_adapter"
model = load_model(projector, lora, device="cuda")
print(answer(model, "photo.jpg", "What is in this image?"))
print(answer(model, "photo.jpg", "What color is the car?", short_answer=True))
```
## Limitations
224x224 input, 49 visual tokens, and a 0.5B LLM. Expect
everyday-scene captioning and simple VQA -- not OCR, fine detail, counting, or
multi-step visual reasoning.
## Training
- optimizer steps: 10518
- final held-out loss: 0.9565
- batch size 4 x 4 accum, lr 0.0002