BrickBench: Evaluating Agentic Brick Design

Peter Kulits, Yiqing Xu, R. Kenny Jones, Cordelia Schmid, Jiajun Wu

arXiv:2610.12452 · 2026-10-11 공개 · arXiv · PDF

text-conditioned agentic-design agent-environment validity-scoring design-validation part-selection constraint-reasoning brickbench

Abstract

We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs. We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs. We release our benchmark and environment at http://www.brickben.ch

한국어 요약

한 줄 요약

BrickBench는 텍스트 조건에 기반한 에이전트 기반 레고 세트 설계를 평가하는 벤치마크로, 설계의 유효성, 정렬, 디자인 품질을 측정한다.

핵심 기여도

핵심 아이디어

기존 레고 생성 모델은 물리적 제약과 디자인 창의성을 동시에 고려하지 못했다. BrickBench는 이러한 한계를 극복하기 위해 **에이전트가 프로그래밍적으로 조립하고 검증할 수 있는 환경인 BrickAgent**를 도입한다. 이 환경은 연결성, 충돌, 안정성을 검증하며, 에이전트가 자체적으로 해결책을 개발하도록 유도한다.

**VQA**(Visual Question Answering)와 **ELO**(Elo rating system)를 사용하여 정렬과 디자인 품질을 평가한다. 특히, 디자인 품질은 인간 판정과 비교하여 측정되며, 이는 단순한 물리적 유효성과 구분된다.

기술적 접근법

주요 결과

의의 및 한계

BrickBench는 **에이전트가 물리적 제약과 창의적 디자인을 동시에 고려하는 능력을 평가**하는 첫 번째 벤치마크로, 기존 데이터 기반 모델을 대체할 수 있는 잠재력을 보여준다. 그러나 **물리적 안정성 시뮬레이션은 힘의 모델을 고려하지 않아 실제 구조적 안정성과 차이가 있을 수 있다**. 또한, **에이전트는 인간 수준의 디자인 창의성에 도달하지 못하며, 이는 아직 해결되지 않은 핵심 과제**이다.

실용적 활용

BrickBench와 BrickAgent는 **로봇 설계, CAD 자동화, 창의적 AI 연구** 등에서 활용 가능하다. 특히, **에이전트가 복잡한 조립을 프로그래밍적으로 생성하고 검증하는 능력은 산업 설계 자동화에 기여할 수 있다**.