AI & ML interests
Feeling and building the multimodal intelligence.
Recent Activity
View all activity
Papers
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-1.5
-
mvp-lab/LLaVA-OneVision-1.5-Instruct-Data
Viewer β’ Updated β’ 21.9M β’ 63.7k β’ 81 -
mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M
Preview β’ Updated β’ 734k β’ 111 -
lmms-lab/LLaVA-OneVision-1.5-8B-Instruct
Image-Text-to-Text β’ 9B β’ Updated β’ 16.3k β’ 66 -
lmms-lab/LLaVA-OneVision-1.5-4B-Instruct
Image-Text-to-Text β’ 5B β’ Updated β’ 496 β’ 18
MMSearch-R1 is a solution designed to train LMMs to perform on-demand multimodal search in real-world environment.
CVPR 2025 - EgoLife: Towards Egocentric Life Assistant. Homepage: https://egolife-ai.github.io/
The collection of the sae that hooked on llava
-
Multimodal SAE
π¬9Demo for Multimodal-SAE
-
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
Paper β’ 2411.14982 β’ Published β’ 19 -
lmms-lab/llava-sae-explanations-5k
Viewer β’ Updated β’ 9.8k β’ 298 β’ 6 -
lmms-lab/llama3-llava-next-8b-hf-sae-131k
Updated β’ 17 β’ 8
Models focus on video understanding (previously known as LLaVA-NeXT-Video).
-
Video Instruction Tuning With Synthetic Data
Paper β’ 2410.02713 β’ Published β’ 41 -
lmms-lab/LLaVA-Video-178K
Viewer β’ Updated β’ 1.63M β’ 35.3k β’ 202 -
lmms-lab/LLaVA-Video-7B-Qwen2
Video-Text-to-Text β’ 8B β’ Updated β’ 19k β’ 127 -
lmms-lab/LLaVA-Video-72B-Qwen2
Text Generation β’ 73B β’ Updated β’ 236 β’ 21
Dataset Collection of LMMs-Eval
-
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Paper β’ 2407.12772 β’ Published β’ 35 -
lmms-lab-encoder/VQAv2
Viewer β’ Updated β’ 770k β’ 35.9k β’ 38 -
lmms-lab-encoder/MME
Viewer β’ Updated β’ 2.37k β’ 22.2k β’ 37 -
lmms-lab-encoder/DocVQA
Viewer β’ Updated β’ 16.6k β’ 37.3k β’ 87
-
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Paper β’ 2407.07895 β’ Published β’ 42 -
lmms-lab/llava-next-interleave-qwen-7b
Text Generation β’ 8B β’ Updated β’ 358 β’ 27 -
lmms-lab/llava-next-interleave-qwen-7b-dpo
Text Generation β’ 8B β’ Updated β’ 162 β’ 12 -
lmms-lab/M4-Instruct-Data
Updated β’ 1.7k β’ 79
Making Lite version of the dataset to accelerate holistic evaluation during model development!
HEVC-Style Vision Transformer
-
OpenMMReasoner/OpenMMReasoner-ColdStart
Image-Text-to-Text β’ 8B β’ Updated β’ 23 β’ 3 -
OpenMMReasoner/OpenMMReasoner-RL
Image-Text-to-Text β’ 8B β’ Updated β’ 72 β’ 18 -
OpenMMReasoner/OpenMMReasoner-SFT-874K
Viewer β’ Updated β’ 874k β’ 1.02k β’ 8 -
OpenMMReasoner/OpenMMReasoner-RL-74K
Viewer β’ Updated β’ 74.7k β’ 309 β’ 10
as a general evaluator for assessing model performance
a model good at arbitrary types of visual input
-
LLaVA-OneVision: Easy Visual Task Transfer
Paper β’ 2408.03326 β’ Published β’ 61 -
lmms-lab/LLaVA-OneVision-Mid-Data
Viewer β’ Updated β’ 563k β’ 141 β’ 21 -
lmms-lab/LLaVA-OneVision-Data
Viewer β’ Updated β’ 3.94M β’ 21.9k β’ 238 -
lmms-lab/LLaVA-NeXT-Data
Viewer β’ Updated β’ 779k β’ 3.8k β’ 47
Long Context Transfer From Text To Vision: https://lmms-lab.github.io/posts/longva/
Some powerful image models.
-
lmms-lab/llava-next-110b
Text Generation β’ 112B β’ Updated β’ 81 β’ 21 -
lmms-lab/llava-next-72b
Text Generation β’ 73B β’ Updated β’ 271 β’ 14 -
lmms-lab/llava-next-qwen-32b
Text Generation β’ 33B β’ Updated β’ 22 β’ 7 -
lmms-lab/llama3-llava-next-8b
Text Generation β’ 8B β’ Updated β’ 1.73k β’ 106
HEVC-Style Vision Transformer
-
OpenMMReasoner/OpenMMReasoner-ColdStart
Image-Text-to-Text β’ 8B β’ Updated β’ 23 β’ 3 -
OpenMMReasoner/OpenMMReasoner-RL
Image-Text-to-Text β’ 8B β’ Updated β’ 72 β’ 18 -
OpenMMReasoner/OpenMMReasoner-SFT-874K
Viewer β’ Updated β’ 874k β’ 1.02k β’ 8 -
OpenMMReasoner/OpenMMReasoner-RL-74K
Viewer β’ Updated β’ 74.7k β’ 309 β’ 10
https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-1.5
-
mvp-lab/LLaVA-OneVision-1.5-Instruct-Data
Viewer β’ Updated β’ 21.9M β’ 63.7k β’ 81 -
mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M
Preview β’ Updated β’ 734k β’ 111 -
lmms-lab/LLaVA-OneVision-1.5-8B-Instruct
Image-Text-to-Text β’ 9B β’ Updated β’ 16.3k β’ 66 -
lmms-lab/LLaVA-OneVision-1.5-4B-Instruct
Image-Text-to-Text β’ 5B β’ Updated β’ 496 β’ 18
MMSearch-R1 is a solution designed to train LMMs to perform on-demand multimodal search in real-world environment.
CVPR 2025 - EgoLife: Towards Egocentric Life Assistant. Homepage: https://egolife-ai.github.io/
The collection of the sae that hooked on llava
-
Multimodal SAE
π¬9Demo for Multimodal-SAE
-
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
Paper β’ 2411.14982 β’ Published β’ 19 -
lmms-lab/llava-sae-explanations-5k
Viewer β’ Updated β’ 9.8k β’ 298 β’ 6 -
lmms-lab/llama3-llava-next-8b-hf-sae-131k
Updated β’ 17 β’ 8
as a general evaluator for assessing model performance
Models focus on video understanding (previously known as LLaVA-NeXT-Video).
-
Video Instruction Tuning With Synthetic Data
Paper β’ 2410.02713 β’ Published β’ 41 -
lmms-lab/LLaVA-Video-178K
Viewer β’ Updated β’ 1.63M β’ 35.3k β’ 202 -
lmms-lab/LLaVA-Video-7B-Qwen2
Video-Text-to-Text β’ 8B β’ Updated β’ 19k β’ 127 -
lmms-lab/LLaVA-Video-72B-Qwen2
Text Generation β’ 73B β’ Updated β’ 236 β’ 21
a model good at arbitrary types of visual input
-
LLaVA-OneVision: Easy Visual Task Transfer
Paper β’ 2408.03326 β’ Published β’ 61 -
lmms-lab/LLaVA-OneVision-Mid-Data
Viewer β’ Updated β’ 563k β’ 141 β’ 21 -
lmms-lab/LLaVA-OneVision-Data
Viewer β’ Updated β’ 3.94M β’ 21.9k β’ 238 -
lmms-lab/LLaVA-NeXT-Data
Viewer β’ Updated β’ 779k β’ 3.8k β’ 47
Dataset Collection of LMMs-Eval
-
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Paper β’ 2407.12772 β’ Published β’ 35 -
lmms-lab-encoder/VQAv2
Viewer β’ Updated β’ 770k β’ 35.9k β’ 38 -
lmms-lab-encoder/MME
Viewer β’ Updated β’ 2.37k β’ 22.2k β’ 37 -
lmms-lab-encoder/DocVQA
Viewer β’ Updated β’ 16.6k β’ 37.3k β’ 87
Long Context Transfer From Text To Vision: https://lmms-lab.github.io/posts/longva/
-
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Paper β’ 2407.07895 β’ Published β’ 42 -
lmms-lab/llava-next-interleave-qwen-7b
Text Generation β’ 8B β’ Updated β’ 358 β’ 27 -
lmms-lab/llava-next-interleave-qwen-7b-dpo
Text Generation β’ 8B β’ Updated β’ 162 β’ 12 -
lmms-lab/M4-Instruct-Data
Updated β’ 1.7k β’ 79
Some powerful image models.
-
lmms-lab/llava-next-110b
Text Generation β’ 112B β’ Updated β’ 81 β’ 21 -
lmms-lab/llava-next-72b
Text Generation β’ 73B β’ Updated β’ 271 β’ 14 -
lmms-lab/llava-next-qwen-32b
Text Generation β’ 33B β’ Updated β’ 22 β’ 7 -
lmms-lab/llama3-llava-next-8b
Text Generation β’ 8B β’ Updated β’ 1.73k β’ 106
Making Lite version of the dataset to accelerate holistic evaluation during model development!