GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation

Tong Wu, Guandao Yang, Zhibing Li, Kai Zhang, Ziwei Liu, Leonidas J. Guibas, Dahua Lin, Gordon Wetzstein

arXiv:2401.04092 · 2026-07-27 공개 · arXiv · PDF

text-to-3d pairwise-comparison gpt-4v evaluation-metric human-aligned elo-rating prompt-generator

Abstract

Despite recent advances in text-to-3D generative methods, there is a notable absence of reliable evaluation metrics. Existing metrics usually focus on a single criterion each, such as how well the asset aligned with the input text. These metrics lack the flexibility to generalize to different evaluation criteria and might not align well with human preferences. Conducting user preference studies is an alternative that offers both adaptability and human-aligned results. User studies, however, can be very ex-pensive to scale. This paper presents an automatic, ver-satile, and human-aligned evaluation metric for text-to-3D generative models. To this end, we first develop a prompt generator using GPT-4V to generate evaluating prompts, which serve as input to compare text-to-3D models. We further design a method instructing GPT-4V to compare two 3D assets according to user-defined crite-ria. Finally, we use these pairwise comparison results to assign these models Elo ratings. Experimental results suggest our metric strongly aligns with human preference across different evaluation criteria. Our code is available at https://github.com/3DTopia/GPTEval3D.

한국어 요약

한 줄 요약

GPT-4V를 활용해 텍스트-3D 생성 모델의 인간 정렬 평가 메트릭을 구축하고, Elo 점수를 통해 모델을 자동 평가하는 시스템을 제안한다.

핵심 기여도

핵심 아이디어

기존 텍스트-3D 생성 모델 평가 메트릭은 단일 기준에만 초점을 맞추고 있으며, 인간 판단과의 정렬성이 낮은 문제가 있었다. 본 연구는 GPT-4V의 다모달 학습 능력을 활용해, 인간 판단과 유사한 방식으로 3D 자산을 평가하는 시스템을 제안한다. 구체적으로, 사용자 정의 평가 기준에 따라 3D 자산을 비교하는 "3D-aware prompt"를 설계하고, 이를 통해 생성된 쌍별 비교 결과를 바탕으로 Elo 점수를 할당함으로써 모델의 종합적 성능을 평가한다. 이는 기존 CLIP 기반 메트릭이 기하학적 세부사항이나 질감을 평가하지 못하는 한계를 보완한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 텍스트-3D 생성 모델 평가에서 인간 판단과의 정렬성을 높이면서도 자동화된 평가를 가능하게 하는 새로운 접근법을 제시한다. 기존 단일 기준 메트릭의 한계를 극복하고, 다양한 평가 기준을 반영할 수 있는 유연한 평가 시스템을 구축한 점에서 학술적·실용적 가치가 있다. 그러나 GPT-4V의 3D 자산 이해 능력이 완벽하지 않기 때문에, 일부 복잡한 기하학적 구조나 세부 텍스처를 정확히 평가하지 못할 수 있는 한계가 존재한다.

실용적 활용

본 연구는 텍스트-3D 생성 모델의 개발 및 비교에 있어 인간 판단과 유사한 방식으로 자동 평가를 수행할 수 있는 도구로 활용될 수 있다. 특히, 연구 개발 과정에서 모델 성능을 빠르게 평가하거나, 산업 현장에서 생성된 3D 자산의 질을 일관되게 평가하는 데 유용하게 사용될 수 있다.