Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models

Jay Zhangjie Wu, Yuxuan Zhang, Haithem Turki, Xuanchi Ren, Jun Gao, M. Shou, Sanja Fidler, Zan Gojcic, Huan Ling

arXiv:2503.01774 · 2026-07-27 공개 · arXiv · PDF

novel-view-synthesis artifact-removal neural-radiance-fields fid-score single-step-diffusion neural-enhancer

Abstract

Neural Radiance Fields and 3D Gaussian Splatting have revolutionized 3D reconstruction and novel-view synthesis task. However, achieving photorealistic rendering from extreme novel viewpoints remains challenging, as artifacts persist across representations. In this work, we introduce Difix3D+, a novel pipeline designed to enhance 3D reconstruction and novel-view synthesis through single-step diffusion models. At the core of our approach is Difix, a single-step image diffusion model trained to enhance and remove artifacts in rendered novel views caused by under-constrained regions of the 3D representation. Difix serves two critical roles in our pipeline. First, it is used during the reconstruction phase to clean up pseudo-training views that are rendered from the reconstruction and then distilled back into 3D. This greatly enhances underconstrained regions and improves the overall 3D representation quality. More importantly, Difix also acts as a neural enhancer during inference, effectively removing residual artifacts arising from imperfect 3D supervision and the limited capacity of current reconstruction models. Difix3D+ is a general solution, a single model compatible with both NeRF and 3DGS representations, and it achieves an average 2× improvement in FID score over baselines while maintaining 3D consistency.

한국어 요약

한 줄 요약

Difix3D+는 단계적 3D 업데이트와 실시간 디퓨전 모델을 활용해 3D 재구성 품질을 2× 향상시키는 새로운 파이프라인이다.

핵심 기여도

핵심 아이디어

Difix3D+는 2D 디퓨전 모델의 시각 지식을 3D 재구성에 효과적으로 전이하는 새로운 접근법을 제시한다. 기존 NeRF나 3DGS는 훈련 데이터에 의존적이며, 미관측 영역에서 아티팩트가 발생한다. Difix는 이러한 아티팩트를 제거하기 위해 단일 단계 디퓨전 모델로 설계되었으며, 훈련 단계에서 생성된 가상 훈련 뷰를 3D로 되돌려 보내면서 품질을 향상시킨다. 또한, 추론 단계에서 실시간으로 아티팩트를 제거하는 역할도 수행한다. Difix는 NeRF와 3DGS 모두에 호환되며, 기존의 다단계 디퓨전 모델과 달리 빠른 추론 속도를 유지한다.

기술적 접근법

주요 결과

의의 및 한계

Difix3D+는 3D 재구성과 새로운 뷰 합성에서 아티팩트 제거 문제를 해결하는 일반적이고 효율적인 솔루션을 제공한다. 특히, 단일 모델로 NeRF와 3DGS 모두를 지원하며, 빠른 추론 속도로 실시간 적용이 가능하다는 점에서 실용적 가치가 높다. 그러나, Difix는 훈련 데이터에 의존적이며, 미관측 영역이 매우 많은 복잡한 장면에서는 한계가 있을 수 있다. 또한, 모델의 훈련 시간은 짧지만, 특정 장면에 대한 미세 조정이 필요할 수 있다.

실용적 활용

Difix3D+는 자동차 산업의 3D 시뮬레이션, 증강현실(AR) 및 가상현실(VR) 환경 구축, 건축 및 도시 설계 분야에서 실시간 3D 재구성 품질 향상에 활용될 수 있다. 특히, 대규모 장면의 빠른 처리가 필요한 산업 현장에서 유용할 것으로 기대된다.