WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 4 days ago • 146
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 11 days ago • 246
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models Paper • 2608.27550 • Published 29 days ago • 84
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published Jul 27 • 37
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Paper • 2607.10350 • Published Jul 15 • 80
UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Paper • 2606.21661 • Published Jun 19 • 29
Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching Paper • 2311.12751 • Published Nov 21, 2023 • 2
UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Paper • 2606.21661 • Published Jun 19 • 29
Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories Paper • 2606.11176 • Published Jun 9 • 133
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation Paper • 2605.18739 • Published May 18 • 117
view article Article NEO-unify: Building Native Multimodal Unified Models End to End sensenova • Mar 5 • 182
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond Paper • 2604.22748 • Published Apr 24 • 158 • 5
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond Paper • 2604.22748 • Published Apr 24 • 158 • 5
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing Paper • 2604.04911 • Published Apr 6 • 33
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond Paper • 2604.22748 • Published Apr 24 • 158