OmniTaskonomy: When Does Visual Generation Improve Visual Understanding? Paper • 2609.38079 • Published 7 days ago • 55
EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory Paper • 2609.37923 • Published 7 days ago • 8
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies Paper • 2609.38155 • Published 7 days ago • 114
MaLiang-Harness: A Programmable Path to Image and Video Generation Paper • 2609.34309 • Published 8 days ago • 412
VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents Paper • 2609.38119 • Published 7 days ago • 11
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification Paper • 2609.06245 • Published Sep 5 • 28
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming Paper • 2608.20958 • Published Aug 21 • 60
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 161
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published Aug 6 • 65
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published Jul 13 • 61
Self-Improving Language Models with Bidirectional Evolutionary Search Paper • 2605.28814 • Published May 27 • 63
AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward Paper • 2605.12495 • Published May 12 • 35
FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization Paper • 2603.19835 • Published Mar 20 • 121
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning Paper • 2603.23483 • Published Mar 24 • 62
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models Paper • 2603.16859 • Published Mar 17 • 48