LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures
Abstract
LLM-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Yet, producing elegant geometry is fundamentally different from producing objects that can be built and perform their intended functions. Existing evaluations largely focus on geometric quality while overlooking physical realizability. We introduce LMBuild, a benchmark for evaluating LLM agents on generating buildable and functional structures. LMBuild represents generated objects as assembled structures comprising part decompositions, joints, materials, and sequences. To support reproducible evaluation, we provide a unified framework consisting of: (1) an interactive environment in which agents can use tools to retrieve, create, and place components to construct objects; (2) a curated benchmark that repurposes established CAD datasets and augments them with knowledge from Wikipedia; and (3) a evaluation framework covering structural soundness, functional affordance, design quality, and physical realization. Evaluations across 30 systems reveal several intriguing findings: (a) Soundness and alignment are no longer the primary bottlenecks for frontier closed-source models, while functional affordance and physical operability remain substantially more challenging; (b) stronger models more effectively create new components, whereas weaker models tend to rely on retrieval; and (c) providing functional specifications substantially improves part completeness, kinematics, and physical operability. These results show that generating real-world structures requires deeper reasoning about functional affordances, mechanics, and designing and creating novel components. We expect LMBuild to provide a foundation for measuring progress and incentivizing research toward agents that generate buildable and functional structures.
Community
HF Daily Papers submission: LMBuild (arXiv 2610.04292)
Links
- GitHub: https://github.com/Lumos-Jiateng/LMBuild
- Project page: https://lumos-jiateng.github.io/LMBuild/
- Dataset: https://huggingface.co/datasets/Lumos-Jiateng/LMBuild
Comment
๐๏ธ LMBuild: can LLM agents build objects that actually work, and not just look right?
LLM agents can already produce good-looking 3D geometry. But a wheelchair whose wheels don't turn, or a desk that can't be assembled, isn't a real object. LMBuild evaluates agents on generating buildable and functional structures: part decompositions, joints, materials and assembly sequences.
๐ง Interactive environment: agents use tools to retrieve catalog parts, create new parts, place and adjust them, then declare joints, materials and an assembly sequence.
๐ฆ Benchmark: we repurpose established CAD datasets (LEGO/BrickComposer, Fusion 360, Artiverse, PartNeXt, open-source product CAD) and ground them in Wikipedia knowledge. LMBuild-Core has 200 objects.
๐ 12 hierarchical metrics in four groups: Soundness (connectivity, collision, stability), Affordance (functional geometry, parts, kinematics), Design (decomposition, aesthetics, alignment; the VLM judges are human-calibrated) and Realization (assembly sequence, material, simulation-based operability).
๐ 30 systems evaluated: frontier closed APIs, open-source (M)LLMs, and domain-specific generators (BrickGPT, PartCrafter, PartPacker, Cube3D, PhysX-Anything โฆ).
Key findings:
1๏ธโฃ Soundness and visual alignment are no longer the main bottleneck for frontier models. Functional affordance and physical operability are: even the best systems score only about 20 on simulation-based operability.
2๏ธโฃ Stronger models create; weaker models retrieve. Allowing part creation helps frontier models but can hurt weaker ones.
3๏ธโฃ Stating functional requirements explicitly substantially improves part completeness, kinematics and operability.
Code, data and evaluation are all open source. Feedback is welcome! ๐
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- HappyWorld-Bench (2026)
- Fysiverse-3D-Vision Technical Report: Generating Executable 3D Worlds from Images through Unified Spatial Reasoning (2026)
- VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? (2026)
- SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes (2026)
- Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction (2026)
- aDSL: Agentic 3D Creation via Joint Agent-Program Design (2026)
- GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel โ
Download all files from Hugging Face
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
