LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs Paper • 2506.05260 • Published Jun 5, 2025
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Paper • 2607.23265 • Published Aug 3 • 1
TimePLE: Rethinking Temporal Representation for Video Temporal Grounding Paper • 2607.23951 • Published Jul 27 • 1
VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents Paper • 2609.38119 • Published 8 days ago • 11
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding? Paper • 2609.38079 • Published 8 days ago • 55
EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory Paper • 2609.37923 • Published 8 days ago • 8
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies Paper • 2609.38155 • Published 8 days ago • 114
MaLiang-Harness: A Programmable Path to Image and Video Generation Paper • 2609.34309 • Published 9 days ago • 412
VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents Paper • 2609.38119 • Published 8 days ago • 11
VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents Paper • 2609.38119 • Published 8 days ago • 11
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification Paper • 2609.06245 • Published Sep 5 • 28
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming Paper • 2608.20958 • Published Aug 21 • 60
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 161
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published Aug 6 • 65