Papers
arxiv:2609.05594

SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution

Published on Sep 4
· Submitted by
Xingjian Ran
on Sep 9
Authors:
,
,
,

Abstract

SceneMosaic combines learned image priors with vision-language agents to efficiently generate diverse, physically valid indoor scenes by evolving local units and composing them globally.

Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet their scalable generation remains challenging. Recent agentic text-to-3D scene pipelines that rely on vision-language models (VLMs) can generate scenes of high fidelity but require costly iterative object placement and refinement. Another mainstream paradigm, parametric image-to-3D scene models, produces scenes efficiently from strong priors learned from 2D images but often leads to imprecise and physically invalid scenes. More importantly, both paradigms struggle to output diverse scenes for a single input, making it hard for them to reflect the dynamically changing nature of real scenes. In this paper we propose SceneMosaic, a framework that combines the merits of both paradigms. It obtains the initial candidate from the learned image-based prior, and subsequently evolves the result through VLM agents, ensuring both efficiency and physical validity. Within the evolution process, SceneMosaic exploits the locality of natural scenes and decomposes a scene into independent local units, allowing separate evolution within each unit before composing the global scene via Cartesian product. On SceneEval-100, SceneMosaic matches the strongest agentic baseline in semantic layout quality with a 24x speedup, substantially reduces physical violations, and receives the highest human ratings. Our code is publicly available at https://github.com/rxjfighting/SceneMosaic.

Community

Paper author Paper submitter

teaser
SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution

Existing agent-based scene generation yields high-quality layouts through iterative refinement, but is slow. Conversely, Image-to-3D methods are fast, but frequently cause physical errors like collisions or floating objects. Crucially, both produce only a single output per input. SceneMosaic solves this by quickly establishing a reliable base scene, then using localized evolution and composition to efficiently generate multiple distinct, physically plausible 3D layouts.

Key Highlights:

  • Hybrid Generation: Uses image priors for fast initialization, followed by agentic iteration to refine spatial positioning—balancing speed, semantic logic, and physical plausibility.
  • Local Evolution to Global Diversity: Decomposes scenes into local sub-units, evolves them independently, and combines them via Cartesian product to create vast layout variants. A perception-aware metric with dynamic Max-Min selection then isolates the most diverse, high-quality scenes.
  • Efficient & Physically Sound: On SceneEval-100, SceneMosaic matches state-of-the-art agent baselines in semantic quality while achieving a 24x speedup and significantly reducing physical violations.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.05594
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.05594 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.05594 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.05594 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.