Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 6 days ago • 35
ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations Paper • 2606.11188 • Published Jun 9 • 27
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Paper • 2604.26694 • Published Apr 29 • 6
VQ-VA World: Towards High-Quality Visual Question-Visual Answering Paper • 2511.20573 • Published Nov 25, 2025 • 7
LightBagel: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation Paper • 2510.22946 • Published Oct 27, 2025 • 18
MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data Augmentation Paper • 2505.04656 • Published May 7, 2025 • 1
DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion Paper • 2411.04928 • Published Nov 7, 2024 • 56
GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting Paper • 2311.14521 • Published Nov 24, 2023
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 6 days ago • 35
LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models Paper • 2603.07145 • Published Mar 7 • 4
LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models Paper • 2603.07145 • Published Mar 7 • 4
Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss Paper • 2501.07563 • Published Jan 13, 2025 • 1