How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
# Warning: Pipeline type "image-to-text" is no longer supported in transformers v5.
# You must load the model directly (see below) or downgrade to v4.x with:
# 'pip install "transformers<5.0.0'
from transformers import pipeline

pipe = pipeline("image-to-text", model="hkondle/CapstoneMainModel")
# Load model directly
from transformers import AutoProcessor, AutoModelForMultimodalLM

processor = AutoProcessor.from_pretrained("hkondle/CapstoneMainModel")
model = AutoModelForMultimodalLM.from_pretrained("hkondle/CapstoneMainModel", device_map="auto")
Quick Links

VisAble diagram captioner

BLIP fine-tuned on AI2D-Caption diagrams to produce part-to-whole descriptions for blind and low-vision students. The vision encoder was frozen and only the text decoder was trained, using a cross entropy loss that up-weights AI2D entity labels by 5.0x.

Trained for 10 epochs, batch size 8, learning rate 5e-05.

from transformers import BlipForConditionalGeneration, BlipProcessor
from PIL import Image

processor = BlipProcessor.from_pretrained("hkondle/CapstoneMainModel-10ep")
model = BlipForConditionalGeneration.from_pretrained("hkondle/CapstoneMainModel-10ep")

image = Image.open("diagram.png").convert("RGB")
inputs = processor(images=image, return_tensors="pt")
print(processor.decode(model.generate(**inputs, max_length=256, num_beams=4)[0], skip_special_tokens=True))
Downloads last month
41
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hkondle/CapstoneMainModel

Finetuned
(59)
this model