SV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image using Latent Video Diffusion

Vikram S. Voleti, C. Yao, Mark Boss, Adam Letts, David Pankratz, Dmitry Tochilkin, Christian Laforte, Robin Rombach, Varun Jampani

arXiv:2403.12008 · 2026-07-27 공개 · arXiv · PDF

novel-view-synthesis high-resolution camera-control image-to-video multi-view-consistency image-to-3d stablediffusion latent-video-diffusion

Abstract

We present Stable Video 3D (SV3D) -- a latent video diffusion model for high-resolution, image-to-multi-view generation of orbital videos around a 3D object. Recent work on 3D generation propose techniques to adapt 2D generative models for novel view synthesis (NVS) and 3D optimization. However, these methods have several disadvantages due to either limited views or inconsistent NVS, thereby affecting the performance of 3D object generation. In this work, we propose SV3D that adapts image-to-video diffusion model for novel multi-view synthesis and 3D generation, thereby leveraging the generalization and multi-view consistency of the video models, while further adding explicit camera control for NVS. We also propose improved 3D optimization techniques to use SV3D and its NVS outputs for image-to-3D generation. Extensive experimental results on multiple datasets with 2D and 3D metrics as well as user study demonstrate SV3D's state-of-the-art performance on NVS as well as 3D reconstruction compared to prior works.

한국어 요약

한 줄 요약

SV3D는 단일 이미지로부터 높은 해상도의 3D 객체 생성과 다각도 뷰 시너지스를 달성하는 라티언트 비디오 디퓨전 모델이다.

핵심 기여도

핵심 아이디어

SV3D는 2D 생성 모델의 **다각도 일관성 부족** 문제를 해결하기 위해 **비디오 디퓨전 모델**을 활용한다. 기존 연구는 2D 이미지 생성 모델을 3D 생성에 재사용하는 방식이었으나, 이는 **NVS의 일관성 부족**으로 인해 3D 생성 성능이 저하되는 문제가 있었다. SV3D는 **SVD** 모델의 **multi-view consistency**와 **generalization ability**를 활용하여 **단일 이미지로부터 일관된 NVS**를 생성하고, 이를 기반으로 **NeRF와 DMTet mesh**를 최적화하여 3D 객체를 생성한다. 특히, **explicit camera pose conditioning**을 통해 **고해상도 orbital video** 생성이 가능하며, **progressive fine-tuning**을 통해 정적(Static)에서 동적(Dynamic) 뷰 생성까지 확장한다.

기술적 접근법

주요 결과

의의 및 한계

SV3D는 **단일 이미지로부터 높은 해상도의 3D 객체 생성**을 가능하게 하며, **NVS의 일관성**과 **사용자 선호도** 측면에서 기존 방법을 크게 앞선다. 특히, **progressive fine-tuning**을 통해 **정적→동적 뷰 생성**을 효과적으로 처리할 수 있다. 그러나 **mirroring surface**와 같은 **반사 표면**은 **Lambertian reflection model**로 표현이 어렵고, **camera pose의 자유도가 2개 (elevation, azimuth)**로 제한되어 있어 **더 복잡한 NVS 시스템** 구축은 여전히 과제이다.

실용적 활용

SV3D는 **게임 개발, AR/VR, 전자상거래, 로봇 시각** 등에서 **단일 이미지로부터 3D 객체 생성**이 필요한 다양한 산업에 적용 가능하다. 특히, **사용자 생성 콘텐츠 (UGC)**와 **실시간 3D 객체 생성** 시스템에 유용하게 활용될 수 있다.