Abstract
Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through image-plane cues that are ambiguous in depth and rotation or through 3D tracks and blobs that lack complete geometry and lose consistency across viewpoint changes. We introduce 4Director, a video world model conditioned on an explicit 4D scene representation: each object is reconstructed once from the input image as a canonical mesh and moved by one prescribed rigid transformation per frame. This representation provides an intuitive 3D control interface and prevents unobserved geometry from being regenerated independently in every frame. We render the controlled scene as a depth video and introduce a Motion Adapter that transforms this geometric scaffold into video while synthesizing view-consistent appearance, illumination, and non-rigid dynamics. For training, we construct RealCOD-Rigid, a new dataset of 20,774 clips annotated with rigid 3D scenes by our automatic pipeline. We further introduce Identity-Gated IoU (IG-IoU), which jointly evaluates adherence to prescribed object motion and preservation of object identity. Experiments demonstrate that 4Director consistently outperforms prior methods in visual quality and in camera and object control.
Community
4Director turns an image into an editable 3D scene. Every object becomes a complete mesh, placed in the same 3D space as the background and the camera. You place the camera, draw a rigid trajectory for each object, and can bring in new objects from other photos. We render the scene as a depth video, and a Motion Adapter turns it into the final shot, adding appearance, lighting, and non-rigid motion.
Unlike 2D boxes or 3D blobs, a complete mesh already contains each object's hidden side, so it is not reinvented frame by frame when the object turns or the camera orbits.
Interactive 3D scenes and comparisons are on the project page.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- 4DStreamCtrl: Interactive Video Generation with Online 4D Control (2026)
- Manifold4D: Denoising on Point Cloud Rendered Manifolds for Video Re-shooting (2026)
- ScaleVid: Geometry-Aware Video Object Scaling with Mesh-Free Inference (2026)
- Proxy2World: Learning to Generate Worlds From Lightweight Proxies without Seeing Them (2026)
- SpatialCrafter: Single Image World Modeling with Generative 3D Proxies (2026)
- WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model (2026)
- RoGSW4RLD: Feed-Forward 4D Gaussian Lifting for Robot World Model Rollouts (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.02160 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper