SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Paper • 2608.05137 • Published 3 days ago • 10
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 155
FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs Paper • 2606.22875 • Published Jun 22 • 12
From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion Paper • 2606.12303 • Published Jun 10 • 33
VIA-SD: Verification via Intra-Model Routing for Speculative Decoding Paper • 2606.12243 • Published Jun 10 • 37
CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solving Paper • 2601.01874 • Published Jan 5 • 20
Clear Nights Ahead: Towards Multi-Weather Nighttime Image Restoration Paper • 2505.16479 • Published May 22, 2025 • 12
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs Paper • 2410.12332 • Published Oct 16, 2024 • 3
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems Paper • 2503.16549 • Published Mar 19, 2025 • 15