A Strong Baseline for Evaluating Vision Encoders
in Multimodal Large Language Models
Yilin Yang1,* · Jun-Tao Tang2,* · Kengyi Wang3 · Siyuan Su3 · Gaoyong Luo4 · Mingda Chen1,†
1School of Artificial Intelligence, Shanghai Jiao Tong University
2Nanjing University ·
3Fudan University ·
4Independent Researcher
*Equal contribution. †Corresponding author.
Overview
This repository hosts the downstream MLLM checkpoints accompanying A Strong Baseline for Evaluating Vision Encoders in Multimodal Large Language Models.
Model Zoo
This Model Zoo covers the 70 visual encoders used in our paper: 43 language-supervised, 22 self-supervised, and 5 discrete tokenizers. The current release provides both pretraining and finetuning checkpoints for all 65 continuous vision encoders with each of the three main language backbones.
Training data
- Pretrain Data — LCS-558K: the image–text alignment dataset from LLaVA-Pretrain, using
blip_laion_cc_sbu_558k.json. This dataset is used for MLLM projector training. - Finetuning Data — LLaVA-v1.5 mix665k (filtered): the LLaVA-v1.5 instruction mixture, using
llava_v1_5_mix665k_drop_ge8kchars.json. The training configuration removes 395 examples with at least 8,000 characters of conversation text.
Downloads
Encoder weightslinks to the original upstream vision encoder or visual tokenizer. Each encoder name links to its model or project page. Use these frozen encoder weights together with the corresponding Stage 1 projector or Stage 2 MLLM checkpoint, preserving the architecture, feature layer and preprocessing specified by that run'sconfig.json.- For Hugging Face models, retain the associated configuration and processor files. Web-SSL MAE 3B has sharded weights; its link opens the complete model repository.
- DINOv3 requires acceptance of the upstream model's terms and authentication. The RAEv2 DINOv3-L (k=7) entry uses the same DINOv3-L/16 backbone with the project's k=7 intermediate-layer readout; its download is the backbone used to construct that representation.
projectordownloads the pretrainingmm_projector.bin;finetunedopens the finetuned checkpoint directory, including model weights, configuration, and language-tokenizer files. The frozen vision encoder must be supplied separately using the matching architecture and weights recorded inconfig.json. Pretraining projector weights alone are not instruction-tuned MLLMs.
Qwen/Qwen2.5-1.5B-Instruct
Qwen3-1.7B
SmolLM2-1.7B-Instruct
Citation
If you use RAVEL or this model zoo, please cite:
@misc{yang2026strong,
title={A Strong Baseline for Evaluating Vision Encoders in Multimodal Large Language Models},
author={Yang, Yilin and Tang, Jun-Tao and Wang, Kengyi and Su, Siyuan and Luo, Gaoyong and Chen, Mingda},
year={2026},
eprint={2610.05413},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.05413}
}