baguettotron-internvit-alignment
Baguettotron-VLM is an open, fully-reproducible, multilingual Vision-Language Model in the sub-1B parameter class. It extends PleIAs/Baguettotron β a 321M text-only reasoning SLM β with visual capabilities via the InternViT-300M-448px-V2.5 vision encoder and a lightweight MLP projector, for a total of ~628M parameters. It inherits six European languages (EN, FR, DE, ES, IT, PL) from the Baguettotron backbone.
Two checkpoints are published:
- baguettotron-internvit-alignment β the projector-only warmup. It describes images, and nothing more.
- baguettotron-vision-vqa β instruction-tuned on top of it. It goes past plain description and follows visual instructions, so prefer it for a richer chat experience.
Apache 2.0. Both live in the Baguettotron-VLM collection. Source: github.com/andreagemelli/baguettotron-vlm Β· Write-up: andreagemelli.me/posts/baguettotron-vlm
Alignment checkpoint. Only the MLP projector was trained (on LLaVA-CC3M-Pretrain-595K) β the ViT and the Baguettotron LLM are unmodified base weights. It writes short captions and nothing more. Published so the alignment stage can be reproduced; for actual use take baguettotron-vision-vqa.
Architecture
Image (448Γ448)
β InternViT-300M-448px-V2.5 (304M, frozen) β 1024 tokens Γ 1024d
β Pixel unshuffle (factor=2) β 256 tokens Γ 4096d
β MLP projector (2-layer, ~2.7M) β 256 tokens Γ 576d
β Interleave with text tokens
β Baguettotron (321M, Llama arch, 80L, h=576)
Total: ~628M parameters
Usage
pip install "transformers>=4.56,<5" torch pillow timm einops accelerate
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
from PIL import Image
model = AutoModelForImageTextToText.from_pretrained(
"andreagemelli/baguettotron-internvit-alignment",
trust_remote_code=True,
dtype=torch.bfloat16,
device_map="auto", # also tested on Apple Silicon (mps) and CPU
)
processor = AutoProcessor.from_pretrained(
"andreagemelli/baguettotron-internvit-alignment",
trust_remote_code=True,
)
image = Image.open("photo.jpg").convert("RGB")
inputs = processor(
messages=[{"role": "user", "content": "<image>\nDescribe the image concisely."}],
image=image,
)
inputs = {k: v.to(model.device) for k, v in inputs.items() if v is not None}
print(model.chat(**inputs))
Example output
Greedy, prompt Describe the image concisely. β verbatim output. Images from COCO val2017.
| output | |
|---|---|
![]() |
a cat is sleeping on the couch |
![]() |
the bear is a good friend. |
![]() |
a sign for a stop |
![]() |
the bus is a red double - decoration |
Chat template
Trained on short image captions with no <think> traces. The processor emits a bare assistant prefix (<|im_start|>assistant\n) and the model completes the caption directly. Keep prompts simple ("Describe the image").
Limitations
- Resolution ceiling. One 448Γ448 crop β 256 visual tokens puts document text at roughly 2β4 px/char. OCR, charts and documents are out of reach by architecture, not by budget. Neither published checkpoint reads text in an image.
- Hallucinations, especially on fine-grained or text-heavy questions.
- Multilingual capability is inherited, not verified. The backbone covers six languages; the VLM was never evaluated on non-English benchmarks.
Tested against
transformers4.57. Newer major versions may need adjustments.
Contributions and suggestions are very welcome β issues, PRs, and ideas for better data mixes, training recipes, or evaluation setups are all appreciated. Open an issue or PR on the GitHub repo.
Citation
If you use or extend Baguettotron-VLM in your research, please cite it:
@misc{gemelli2026baguettotronvlm,
title = {Baguettotron-VLM: An Open, Reproducible, Multilingual Sub-1B Vision-Language Model},
author = {Gemelli, Andrea},
year = {2026},
howpublished = {\url{https://huggingface.co/andreagemelli/baguettotron-vision-vqa}},
note = {Code: \url{https://github.com/andreagemelli/baguettotron-vlm}}
}
License
Apache 2.0 β see the GitHub repo.
- Downloads last month
- 18
Model tree for andreagemelli/baguettotron-internvit-alignment
Base model
OpenGVLab/InternViT-300M-448px


