Luce: Relightable Gaussians for 3D Asset Generation

Mayank Singh, Michele Stoppa, Alvise Memo, Rui Yu, Harsha Kalli, Srimanth Gunturi, Muhammad Ahmed Riaz, Behrooz Shahsavari, Waleed Abdulla, David E. Jacobs

arXiv:2608.23943 · 2026-08-30 공개 · arXiv · PDF

latent-space variational-autoencoder image-to-3d rectified-flow-transformer relightable-gaussians pbr-materials toys4k-benchmark texture-mesh

Abstract

High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. To support relighting and integration into standard rendering pipelines, the representation should include physically based rendering (PBR) modalities such as albedo, metallic-roughness, and surface normals. We propose Luce, a 3D representation that unifies geometry and PBR materials within a voxelized multimodal Gaussian cloud, using dedicated Gaussian primitives for each modality. A variational autoencoder compresses this representation into a unified material-aware latent space. A rectified-flow transformer generates this latent from a single image, conditioned on multi-layer features from a pretrained image encoder that preserve both semantic context and fine spatial detail. The latent then decodes into relightable PBR Gaussians and an optional textured mesh with a tangent-space normal map. On Toys4K, Luce achieves state-of-the-art single-image-to-3D generation, improving FID by 28% over the strongest baseline. We further introduce a benchmark of AI-generated images, on which Luce improves the CLIP image-alignment score over the best baseline (0.8519 vs. 0.8299). Luce generates relightable, geometrically accurate, and materially faithful assets that preserve fine details such as text, logos, and inscriptions.

한국어 요약

한 줄 요약

Luce는 이미지에서 재조명 가능한 3D 자산을 생성하기 위한 다중 모달 PBR 가우시안 표현을 제안하며, 기존 기법 대비 28% 개선된 FID 성능을 보인다.

핵심 기여도

핵심 아이디어

기존 3D 생성 모델은 기하학과 재질 표현을 분리하거나, 재조명을 지원하지 않는 한계가 있었다. Luce는 **voxelized multimodal Gaussian cloud**를 통해 기하학과 PBR 재질을 통합 표현하며, 각 PBR 모달(알베도, 메탈릭-러프니스, 노멀)에 대해 별도의 가우시안 원시를 할당한다. 이는 표면의 고주파 세부 정보(예: 인쇄된 텍스트, 로고)를 보존하는 데 기여한다. 또한, **SLatVAE**(Structured Latent VAE)를 통해 이 표현을 압축한 레이턴트 공간을 학습하고, **SLatFlow**(Structured Latent Flow)라는 rectified-flow transformer를 사용하여 단일 이미지에서 이 레이턴트를 생성한다. 이 레이턴트는 재조명 가능한 PBR 가우시안과 선택적 텍스처 메시로 디코딩된다.

기술적 접근법

주요 결과

의의 및 한계

Luce는 기존 3D 생성 모델이 부족했던 **재조명 가능성**과 **표면 세부 정보 보존**을 동시에 달성하며, 산업적 렌더링 파이프라인에 직접 적용 가능한 PBR 표현을 제공한다. 특히, **SLatFlow**와 **SLatVAE**의 결합은 단일 이미지에서 고정밀 3D 자산을 생성하는 데 기여한다. 그러나, **하이퍼파라미터 세부 사항**이나 **대규모 데이터셋에서의 일반화 성능**은 명시되지 않았으며, 이는 향후 연구 주제로 남는다.

실용적 활용

Luce는 **게임 개발**, **영화 VFX**, **상업용 3D 모델링** 등에서 사용할 수 있으며, 특히 **재조명이 필요한 렌더링 작업**이나 **고해상도 세부 정보가 필요한 자산 생성**에 적합하다. 또한, **AI 생성 이미지 기반 3D 자산 생성** 분야에서도 활용 가능하다.