Instructions to use hkondle/CapstoneMainModel with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use hkondle/CapstoneMainModel with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "image-to-text" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("image-to-text", model="hkondle/CapstoneMainModel")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("hkondle/CapstoneMainModel") model = AutoModelForMultimodalLM.from_pretrained("hkondle/CapstoneMainModel", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| license: mit | |
| library_name: transformers | |
| pipeline_tag: image-to-text | |
| base_model: Salesforce/blip-image-captioning-base | |
| tags: | |
| - blip | |
| - accessibility | |
| - diagram-captioning | |
| # VisAble diagram captioner | |
| BLIP fine-tuned on AI2D-Caption diagrams to produce part-to-whole descriptions for blind and | |
| low-vision students. The vision encoder was frozen and only the text decoder was trained, using a | |
| cross entropy loss that up-weights AI2D entity labels by 5.0x. | |
| Trained for 10 epochs, batch size 8, learning rate 5e-05. | |
| ```python | |
| from transformers import BlipForConditionalGeneration, BlipProcessor | |
| from PIL import Image | |
| processor = BlipProcessor.from_pretrained("hkondle/CapstoneMainModel-10ep") | |
| model = BlipForConditionalGeneration.from_pretrained("hkondle/CapstoneMainModel-10ep") | |
| image = Image.open("diagram.png").convert("RGB") | |
| inputs = processor(images=image, return_tensors="pt") | |
| print(processor.decode(model.generate(**inputs, max_length=256, num_beams=4)[0], skip_special_tokens=True)) | |
| ``` | |