ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing Paper • 2609.38541 • Published 5 days ago • 30
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies Paper • 2609.38155 • Published 5 days ago • 104
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding? Paper • 2609.38079 • Published 5 days ago • 52
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning Paper • 2609.35767 • Published 6 days ago • 48
TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces Paper • 2609.33295 • Published 7 days ago • 71
OmniTaskonomy Collection OmniTaskonomy probes when generation helps understanding. Check our project page here: https://omni-taskonomy.github.io/ • 3 items • Updated 4 days ago • 2
PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing Paper • 2609.23784 • Published 14 days ago • 14
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 13 days ago • 157
OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation Paper • 2609.22069 • Published 16 days ago • 37
Paint-Anything: Unified Any-Color Control for Image Generation and Editing Paper • 2609.20816 • Published 17 days ago • 57
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Paper • 2607.27380 • Published Jul 29 • 75
Echo-Memory: A Controlled Study of Memory in Action World Models Paper • 2606.09803 • Published Jun 8 • 33
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Paper • 2604.23781 • Published Apr 26 • 34
UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models Paper • 2604.17565 • Published Apr 19 • 10
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details Paper • 2604.06870 • Published Apr 8 • 44
view article Article NEO-unify: Building Native Multimodal Unified Models End to End sensenova • Mar 5 • 182
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation Paper • 2603.08652 • Published Mar 9 • 40
Next-Embedding Prediction Makes Strong Vision Learners Paper • 2512.16922 • Published Dec 18, 2025 • 91