ciphermosaic's picture
Update README.md
71574c6 verified
|
Raw
History Blame Contribute Delete
3.02 kB
# BLIP Image Captioning LoRA
A LoRA fine tuned version of **Salesforce/blip-image-captioning-base** for image caption generation using the **Flickr8k** dataset. This project demonstrates parameter efficient fine tuning with PEFT while keeping the base model frozen.
## Model Details
* **Model:** `ciphermosaic/blip-image-captioning`
* **Task:** Image Captioning
* **Base Model:** `Salesforce/blip-image-captioning-base`
* **Fine Tuning Method:** LoRA (PEFT)
* **License:** Apache-2.0
## Dataset
The model was fine tuned on the **jxie/flickr8k** dataset.
Dataset split:
| Split | Samples |
| ---------- | ------: |
| Train | 6,000 |
| Validation | 1,000 |
| Test | 1,000 |
## Training Configuration
* Framework: PyTorch
* Transformers
* PEFT (LoRA)
* Hardware: NVIDIA Tesla T4 (Google Colab)
* Epochs: 2
* Batch Size: 4
* Learning Rate: `5e-5`
* Optimizer: AdamW
### LoRA Configuration
```python
LoraConfig(
r=8,
lora_alpha=16,
lora_dropout=0.1,
target_modules=[
"qkv",
"projection"
],
bias="none"
)
```
## Evaluation
| Metric | Value |
| --------------- | ---------: |
| Validation Loss | **8.0478** |
This checkpoint is intended as an educational fine tuning project and baseline implementation for BLIP image captioning with LoRA.
## Usage
### Load with Transformers and PEFT
```python
from transformers import BlipProcessor, BlipForConditionalGeneration
from peft import PeftModel
base_model = BlipForConditionalGeneration.from_pretrained(
"Salesforce/blip-image-captioning-base"
)
model = PeftModel.from_pretrained(
base_model,
"ciphermosaic/blip-image-captioning"
)
processor = BlipProcessor.from_pretrained(
"Salesforce/blip-image-captioning-base"
)
```
### Example Inference
```python
from PIL import Image
import requests
image = Image.open(requests.get(image_url, stream=True).raw).convert("RGB")
inputs = processor(images=image, return_tensors="pt")
outputs = model.generate(**inputs)
caption = processor.decode(outputs[0], skip_special_tokens=True)
print(caption)
```
The model can also be loaded using standard Hugging Face Transformers workflows together with the PEFT adapter.
## Limitations
* Trained only on Flickr8k, therefore generalization to unseen domains may be limited.
* Caption quality depends on image content and may not accurately describe complex scenes.
* Primarily produces English captions.
* Performance may vary for medical, satellite, or highly specialized imagery.
## Intended Uses
* Image caption generation
* Vision Language Model experimentation
* Learning PEFT and LoRA fine tuning
* Educational and research purposes
* Baseline for further image captioning improvements
## Acknowledgements
This work builds upon:
* Salesforce BLIP
* Hugging Face Transformers
* Hugging Face PEFT
* Flickr8k Dataset
Special thanks to the open source community for making these resources available.
## Author
**CipherMosaic**
GitHub: https://github.com/ciphermosaic