File size: 1,001 Bytes
2792967
 
bb2798e
 
 
 
 
 
 
2792967
bb2798e
 
 
9b3f6c2
 
 
bb2798e
9b3f6c2
bb2798e
9b3f6c2
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
---
license: mit
library_name: transformers
pipeline_tag: image-to-text
base_model: Salesforce/blip-image-captioning-base
tags:
- blip
- accessibility
- diagram-captioning
---

# VisAble diagram captioner

BLIP fine-tuned on AI2D-Caption diagrams to produce part-to-whole descriptions for blind and
low-vision students. The vision encoder was frozen and only the text decoder was trained, using a
cross entropy loss that up-weights AI2D entity labels by 5.0x.

Trained for 10 epochs, batch size 8, learning rate 5e-05.

```python
from transformers import BlipForConditionalGeneration, BlipProcessor
from PIL import Image

processor = BlipProcessor.from_pretrained("hkondle/CapstoneMainModel-10ep")
model = BlipForConditionalGeneration.from_pretrained("hkondle/CapstoneMainModel-10ep")

image = Image.open("diagram.png").convert("RGB")
inputs = processor(images=image, return_tensors="pt")
print(processor.decode(model.generate(**inputs, max_length=256, num_beams=4)[0], skip_special_tokens=True))
```