HUGS: Holistic Urban 3D Scene Understanding via Gaussian Splatting

Hongyu Zhou, Jiahao Shao, Luxiao Xu, Dongfeng Bai, Weichao Qiu, Bingbing Liu, Yue Wang, Andreas Geiger, Yiyi Liao

arXiv:2403.12722 · 2026-07-27 공개 · arXiv · PDF

novel-view-synthesis gaussian-splatting real-time-rendering semantic-segmentation dynamic-scenes motion-estimation kitti urban-scene-understanding

Abstract

Holistic understanding of urban scenes based on RGB images is a challenging yet important problem. It encompasses understanding both the geometry and appearance to enable novel view synthesis, parsing semantic labels, and tracking moving objects. Despite considerable progress, existing approaches often focus on specific aspects of this task and require additional inputs such as LiDAR scans or manually annotated 3D bounding boxes. In this paper, we introduce a novel pipeline that utilizes 3D Gaussian Splatting for holistic urban scene understanding. Our main idea involves the joint optimization of geometry, appearance, semantics, and motion using a combination of static and dynamic 3D Gaussians, where moving object poses are regularized via physical constraints. Our approach offers the ability to render new viewpoints in real-time, yielding 2D and 3D semantic information with high accuracy, and reconstruct dynamic scenes, even in scenarios where 3D bounding box detection are highly noisy. Experimental results on KITTI, KITTI-360, and Virtual KITTI 2 demonstrate the effectiveness of our approach. Our project page is at https://xdimlab.github.io/hugs_website.

한국어 요약

한 줄 요약

HUGS는 3D 가우시안 스플래팅을 활용해 도시 장면의 기하, 외관, 의미, 운동을 통합 최적화하는 실시간 렌더링 파이프라인을 제안한다.

핵심 기여도

핵심 아이디어

기존 연구는 도시 장면 이해를 위해 LiDAR나 수동 라벨링된 3D 바운딩 박스를 필요로 했으나, HUGS는 단순 RGB 이미지만으로 통합적인 3D 장면 이해를 구현한다. 이는 3D 가우시안 스플래팅을 기반으로, 정적 및 동적 가우시안을 결합하여 기하, 외관, 의미, 운동을 동시에 최적화하는 데 있다. 특히, 동적 객체의 자세는 유니사이클 모델을 통해 물리적 제약을 적용하여 추적 정확도를 높인다. 이는 개별 프레임 최적화보다 안정적이고 노이즈에 강한 결과를 제공한다.

기술적 접근법

주요 결과

의의 및 한계

HUGS는 단순 RGB 이미지만으로도 정확한 3D 장면 이해를 가능하게 하며, 실시간 렌더링과 높은 정확도를 동시에 달성하는 점에서 혁신적이다. 특히, 유니사이클 모델을 통한 동적 객체 추적은 노이즈가 많은 3D 바운딩 박스에서도 안정적인 성능을 보인다. 그러나 현재 모델은 객체 회전에 제한이 있으며, 빛 편집과 같은 추가 자유도를 통제하는 방향으로 발전이 필요하다. 또한, 장면 노출 변동이 큰 경우 노출 모델링이 필수적임을 보여준다.

실용적 활용

자율주행 시뮬레이션, 도시 모델링, AR/VR 등에서 실시간 3D 장면 생성과 의미 정보 추출에 활용 가능하다. 특히, LiDAR가 없는 저비용 시스템에서 유용하며, 노이즈가 많은 환경에서도 안정적인 장면 재구성이 필요한 산업 분야에 적용 가능하다.