4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
Abstract
We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse physical phenomena, including deformation, fluid flow, and fracture. We perform extensive benchmarking of frontier models, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics. 4DCodeBench provides a testbed for tracking progress toward agents that can interpret the dynamics of the world through code. Our benchmark is available at https://github.com/4DCodeBench/4DCodeBench
Community
Can coding agents reconstruct how the world moves—not just how it looks—from video?
4DCodeBench evaluates this through 4D inverse graphics, asking coding agents to reconstruct dynamic scenes as executable graphics programs.
- 200 scenes: 100 real + 100 synthetic, covering diverse physical phenomena.
- 18 frontier agents evaluated on geometry, appearance, and dynamics reconstruction.
- Dynamics are the bottleneck: agents recover static 3D structure much better than motion.
- Agents rarely simulate: 67% of solutions describe motion analytically; stronger reasoning leads to more simulation and better dynamics reconstruction.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender (2026)
- LEGO-Anything: Coding Agents for 3D Scene Reconstruction (2026)
- Text2Sim: Agentic Physics-Based Simulation Generation with Distilled Expertise (2026)
- 3DHarnessBench: Probing Agentic 3D-to-Code Capabilities of Frontier Vision-Language Models (2026)
- CoDimRecon: Agentic Reconstruction of Sim-Ready 3D Scenes with Deformable Curves, Surfaces, and Volumes (2026)
- NeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied Simulation (2026)
- Real2Gym: Building Gyms from Videos, Bringing Skills to Robots (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.03715 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 2
4DCodeBench/Dataset-Real-World
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper