LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures

Jiateng Liu, Rushi Wang, Cheng Qian, Xuejun Zhang, Sun Li, Jiayu Liu, Yifan Shen, Xu Cao, Jiarui Yao, Bingxuan Li, Ruhi Sarikaya, Heng Ji

arXiv:2610.04292 · 2026-10-06 공개 · arXiv · PDF

llm-agents retrieval-tools buildable-structures functional-structures cad-datasets physical-realization structural-soundness functional-affordance

Abstract

LLM-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Yet, producing elegant geometry is fundamentally different from producing objects that can be built and perform their intended functions. Existing evaluations largely focus on geometric quality while overlooking physical realizability. We introduce LMBuild, a benchmark for evaluating LLM agents on generating buildable and functional structures. LMBuild represents generated objects as assembled structures comprising part decompositions, joints, materials, and sequences. To support reproducible evaluation, we provide a unified framework consisting of: (1) an interactive environment in which agents can use tools to retrieve, create, and place components to construct objects; (2) a curated benchmark that repurposes established CAD datasets and augments them with knowledge from Wikipedia; and (3) a evaluation framework covering structural soundness, functional affordance, design quality, and physical realization. Evaluations across 30 systems reveal several intriguing findings: (a) Soundness and alignment are no longer the primary bottlenecks for frontier closed-source models, while functional affordance and physical operability remain substantially more challenging; (b) stronger models more effectively create new components, whereas weaker models tend to rely on retrieval; and (c) providing functional specifications substantially improves part completeness, kinematics, and physical operability. These results show that generating real-world structures requires deeper reasoning about functional affordances, mechanics, and designing and creating novel components. We expect LMBuild to provide a foundation for measuring progress and incentivizing research toward agents that generate buildable and functional structures.

한국어 요약

한 줄 요약

LMBuild는 생성된 3D 구조물이 실제로 제작 가능하고 기능적으로 작동하는지 평가하는 체계적인 벤치마크를 제시한다.

핵심 기여도

핵심 아이디어

기존 3D 생성 평가가 주로 시각적 품질에 집중하는 반면, LMBuild는 생성된 구조물이 실제로 제작 가능하고 기능적으로 작동하는지 평가하는 새로운 접근을 제시한다. 이는 단순히 3D 모델을 생성하는 것을 넘어, 부품 분해, 조인트, 재료, 조립 순서를 포함한 조립 구조로 표현함으로써, 물리적 실현 가능성과 기능적 의도를 평가한다. 특히, LMBuild는 Wikipedia 지식을 활용해 기능적 핵심 부품과 운동학적 관계를 텍스트 기반으로 평가하며, 이를 통해 생성된 구조물이 실제 물리적 제약과 기능적 요구사항을 충족하는지 판단한다.

기술적 접근법

주요 결과

의의 및 한계

LMBuild는 3D 구조 생성 모델이 단순히 시각적으로 매력적인 모델을 생성하는 것을 넘어, 실제로 제작 가능하고 기능적으로 작동하는 구조를 생성하는 능력을 평가할 수 있는 체계적인 기반을 제공한다. 특히, 기능적 의도와 물리적 제약을 고려한 평가 체계는 기존 평가 방법의 한계를 보완하며, 오픈소스와 최첨단 모델 간의 격차를 진단하는 데 유용하다. 그러나, 설계 품질, 사용자 친화적 기하학, 조립 가능성, 작동 가능성 등은 아직 완전히 객관화된 평가가 어려운 영역이며, 제조 및 물리적 테스트를 통한 추가 검증이 필요하다.

실용적 활용

LMBuild는 건축, 제조, 로봇 공학 등에서 실제 제작 가능한 3D 구조를 생성하는 AI 에이전트의 개발과 평가에 활용될 수 있다. 또한, 기능적 명세를 기반으로 한 설계 자동화, 부품 생성 능력 향상, 물리적 제약을 고려한 설계 최적화 등 다양한 연구 분야에 기초 자료로 활용될 수 있다.