WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 6 days ago • 152
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs Paper • 2609.26796 • Published 5 days ago • 30
Grounded Action Model: 3D Grounding as a Foundation for Robotics Paper • 2609.23863 • Published 7 days ago • 88
Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms Paper • 2609.23658 • Published 7 days ago • 27
astune/text_info_trending_youtube_videos_2019-04-15_to_2020-04-15 Viewer • Updated May 3 • 10.6k • 77 • 5
Rapidata/sora-video-generation-alignment-likert-scoring Viewer • Updated Feb 4, 2025 • 198 • 102 • 16
OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation Paper • 2609.22069 • Published 9 days ago • 35
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue Paper • 2609.21465 • Published 9 days ago • 147
Paint-Anything: Unified Any-Color Control for Image Generation and Editing Paper • 2609.20816 • Published 10 days ago • 55
Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model Paper • 2609.18323 • Published 11 days ago • 132
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation Paper • 2609.20744 • Published 10 days ago • 52
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Paper • 2609.15810 • Published 13 days ago • 50
EventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset Paper • 2609.17189 • Published 12 days ago • 28