lmms-lab-ov2

community

AI & ML interests

None defined yet.

Recent Activity

xiangan authored a paper 5 days ago

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

xiangan authored a paper 5 days ago

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

xiangan submitted a paper 7 days ago

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

View all activity

authored 2 papers 5 days ago

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

Paper • 2605.05997 • Published 27 days ago • 18

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Paper • 2605.25979 • Published 9 days ago • 27

submitted a paper to Daily Papers 7 days ago

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Paper • 2605.25979 • Published 9 days ago • 27

authored a paper 8 days ago

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning

Paper • 2605.20342 • Published 15 days ago • 34

published a model 15 days ago

lmms-lab-ov2/LLaVA-OneVision2-8B-Instruct

Image-Text-to-Text • 9B • Updated 20 days ago • 26

updated a model 20 days ago

lmms-lab-ov2/LLaVA-OneVision2-8B-Instruct

Image-Text-to-Text • 9B • Updated 20 days ago • 26

authored a paper 30 days ago

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

Paper • 2604.28185 • Published Apr 30 • 90

authored a paper 4 months ago

OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

Paper • 2602.08683 • Published Feb 9 • 52

submitted a paper to Daily Papers 4 months ago

OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

Paper • 2602.08683 • Published Feb 9 • 52

authored 2 papers 4 months ago

ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder

Paper • 2510.18795 • Published Oct 21, 2025 • 11

DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset

Paper • 2601.10305 • Published Jan 15 • 37

authored 3 papers 6 months ago

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

Paper • 2511.20785 • Published Nov 25, 2025 • 188

UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning

Paper • 2510.13515 • Published Oct 15, 2025 • 12

OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe

Paper • 2511.16334 • Published Nov 20, 2025 • 96

authored 3 papers 7 months ago

ForCenNet: Foreground-Centric Network for Document Image Rectification

Paper • 2507.19804 • Published Jul 26, 2025 • 12

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

Paper • 2509.09118 • Published Sep 11, 2025 • 8

UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning

Paper • 2510.13515 • Published Oct 15, 2025 • 12

authored a paper 8 months ago

ORID: Organ-Regional Information Driven Framework for Radiology Report Generation

Paper • 2411.13025 • Published Nov 20, 2024 • 2

authored 2 papers 8 months ago

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Paper • 2411.15296 • Published Nov 22, 2024 • 21

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Paper • 2501.13826 • Published Jan 23, 2025 • 24