HuggingFaceTB
/

SmolVLM2-2.2B-Instruct

Image-Text-to-Text

video-text-to-text

Model card Files Files and versions

mfarre commited on Mar 3

Commit

f83c1bb

·

verified ·

1 Parent(s): e00155d

Update README.md

Files changed (1) hide show

README.md +11 -0

README.md CHANGED Viewed

@@ -207,6 +207,17 @@ SmolVLM2 is built upon [the shape-optimized SigLIP](https://huggingface.co/googl
 We release the SmolVLM2 checkpoints under the Apache 2.0 license.
 ## Training Data
 SmolVLM2 used 3.3M samples for training originally from ten different datasets: [LlaVa Onevision](https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Data), [M4-Instruct](https://huggingface.co/datasets/lmms-lab/M4-Instruct-Data), [Mammoth](https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M), [LlaVa Video 178K](https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K), [FineVideo](https://huggingface.co/datasets/HuggingFaceFV/finevideo), [VideoStar](https://huggingface.co/datasets/orrzohar/Video-STaR), [VRipt](https://huggingface.co/datasets/Mutonix/Vript), [Vista-400K](https://huggingface.co/datasets/TIGER-Lab/VISTA-400K), [MovieChat](https://huggingface.co/datasets/Enxin/MovieChat-1K_train) and [ShareGPT4Video](https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video).
 In the following plots we give a general overview of the samples across modalities and the source of those samples.

 We release the SmolVLM2 checkpoints under the Apache 2.0 license.
+## Citation information
+You can cite us in the following way:
+```bibtex
+@misc{smolvlm2,
+  title = {SmolVLM2: Bringing Video Understanding to Every Device},
+  author = {Orr Zohar and Miquel Farré and Andi Marafioti and Merve Noyan and Pedro Cuenca and Cyril Zakka and Joshua Lochner},
+  year = {2025},
+  url = {https://huggingface.co/blog/smolvlm2}
+}
+```
 ## Training Data
 SmolVLM2 used 3.3M samples for training originally from ten different datasets: [LlaVa Onevision](https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Data), [M4-Instruct](https://huggingface.co/datasets/lmms-lab/M4-Instruct-Data), [Mammoth](https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M), [LlaVa Video 178K](https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K), [FineVideo](https://huggingface.co/datasets/HuggingFaceFV/finevideo), [VideoStar](https://huggingface.co/datasets/orrzohar/Video-STaR), [VRipt](https://huggingface.co/datasets/Mutonix/Vript), [Vista-400K](https://huggingface.co/datasets/TIGER-Lab/VISTA-400K), [MovieChat](https://huggingface.co/datasets/Enxin/MovieChat-1K_train) and [ShareGPT4Video](https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video).
 In the following plots we give a general overview of the samples across modalities and the source of those samples.