deepseek-ai/DeepSeek-V4-Flash-Vision-Exp Image-Text-to-Text • 305B • Updated 21 days ago • 803k • • 925
Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding Paper • 2608.28192 • Published 25 days ago • 17
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Paper • 2608.13552 • Published Aug 13 • 47
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images Paper • 2608.06270 • Published Aug 6 • 8
InSight-doc: Agentic Visual Perception for Long-Document Understanding Paper • 2608.10628 • Published Aug 11 • 11
CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing Paper • 2508.06937 • Published Aug 9, 2025 • 7
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Paper • 2507.07562 • Published Jul 10, 2025 • 1
MODNet: Real-Time Trimap-Free Portrait Matting via Objective Decomposition Paper • 2011.11961 • Published Nov 24, 2020
InSight-doc Collection Agentic Visual Perception for Long-Document Understanding • 5 items • Updated Aug 15 • 1
InSight-doc Collection Agentic Visual Perception for Long-Document Understanding • 5 items • Updated Aug 15 • 1