3DGStream: On-the-Fly Training of 3D Gaussians for Efficient Streaming of Photo-Realistic Free-Viewpoint Videos

Jiakai Sun, Han Jiao, Guangyuan Li, Zhanjie Zhang, Lei Zhao, Wei Xing

arXiv:2403.01444 · 2026-07-27 공개 · arXiv · PDF

real-time-rendering streaming-video dynamic-scenes neural-rendering adaptive-modeling photo-realistic-rendering free-viewpoint-video neural-transformation-cache

Abstract

Constructing photo-realistic Free-Viewpoint Videos (FVVs) of dynamic scenes from multi-view videos remains a challenging endeavor. Despite the remarkable advance- ments achieved by current neural rendering techniques, these methods generally require complete video sequences for offline training and are not capable of real-time rendering. To address these constraints, we introduce 3DGStream, a method designed for efficient FVV streaming of real-world dynamic scenes. Our method achieves fast on-the-fly per- frame reconstruction within 12 seconds and real-time ren- dering at 200 FPS. Specifically, we utilize 3D Gaussians (3DGs) to represent the scene. Instead of the naï ve ap- proach of directly optimizing 3DGs per-frame, we employ a compact Neural Transformation Cache (NTC) to model the translations and rotations of 3DGs, markedly reducing the training time and storage required for each FVV frame. Furthermore, we propose an adaptive 3DG addition strat- egy to handle emerging objects in dynamic scenes. Exper- iments demonstrate that 3DGStream achieves competitive performance in terms of rendering speed, image quality, training time, and model storage when compared with state- of-the-art methods.

한국어 요약

한 줄 요약

3DGStream은 실시간 렌더링과 초당 200프레임(FPS)을 지원하는 3D 가우시안 기반의 실내/실외 동적 장면의 실시간 프리뷰(Free-Viewpoint Video) 스트리밍 기법이다.

핵심 기여도

핵심 아이디어

기존의 NeRF 기반 FVV 생성 방법은 전체 영상 시퀀스에 대한 오프라인 훈련이 필요하며, 실시간 렌더링이 어려운 한계를 가짐. 3DGStream은 3D 가우시안(3DGs)을 장면 표현 기반으로 사용하고, 프레임별 훈련을 통해 실시간 스트리밍을 가능하게 함. 핵심 아이디어는 기존 3DGs의 직접 최적화 대신, NTC를 통해 3DGs의 변환(이동 및 회전)을 모델링함으로써 훈련 시간과 저장 공간을 줄이는 것이다. 또한, 새로운 객체가 등장하는 경우 Adaptive 3DG Addition 전략을 통해 해당 객체에 맞는 3DGs를 추가하고, 주기적인 분할 및 가지치기(pruning)를 통해 정확도를 유지함.

기술적 접근법

주요 결과

의의 및 한계

3DGStream은 동적 장면의 실시간 FVV 스트리밍을 가능하게 하며, 기존 오프라인 훈련 방식의 한계를 극복함. NTC와 Adaptive 3DG Addition의 결합은 장면 변화를 효율적으로 처리할 수 있는 새로운 접근법을 제시. 그러나 초기 3DGs(3DG-S)의 품질에 크게 의존하며, COLMAP 기반 포인트 클라우드의 한계로 인해 원거리 장면 재구성이 어려움. 또한, 훈련 반복 횟수 제한으로 급격한 움직임이나 복잡한 객체 처리가 제한적임.

실용적 활용

3DGStream은 VR/AR/XR 분야에서 실시간 동적 장면 스트리밍을 요구하는 애플리케이션에 적합. 예를 들어, 실내/실외 라이브 이벤트, 게임, 원격 협업 환경에서 사용 가능. 특히, 고해상도와 빠른 렌더링 속도를 요구하는 산업 현장에서 유용함.