This collection gathers clean multimodal datasets for VLM supervised fine-tuning. Covering visual QA, image dialogue and document reasoning.
Note https://tinyllava-factory.readthedocs.io/en/latest/Prepare%20Datasets.html