Shape of Motion: 4D Reconstruction From a Single Video

Qianqian Wang, Vickie Ye, Hang Gao, Jake Austin, Zhengqi Li, Angjoo Kanazawa

arXiv:2407.13764 · 2026-07-27 공개 · arXiv · PDF

novel-view-synthesis depth-maps motion-estimation monocular-vision dynamic-scene motion-bases se3-structure

Abstract

Monocular dynamic reconstruction is a challenging and long-standing vision problem due to the highly ill-posed nature of the task. Existing approaches depend on templates, are effective only in quasi-static scenes, or fail to model 3D motion explicitly. We introduce a method for reconstructing generic dynamic scenes, featuring explicit, persistent 3D motion trajectories in the world coordinate frame, from casually captured monocular videos. We tackle the problem with two key insights: First, we exploit the low-dimensional structure of 3D motion by representing scene motion with a compact set of $\operatorname{SE}(3)$ motion bases. Each point's motion is expressed as a linear combination of these bases, facilitating soft decomposition of the scene into multiple rigidly-moving groups. Second, we take advantage of off-the-shelf data-driven priors such as monocular depth maps and long-range $2 D$ tracks, and devise a method to effectively consolidate these noisy supervisory signals, resulting in a globally consistent representation of the dynamic scene. Experiments show that our method achieves state-of-the-art performance for both long-range 3D/2D motion estimation and novel view synthesis on dynamic scenes. Project Page: https://shape-o£-motion.github.io/

한국어 요약

한 줄 요약

단일 비디오에서 3D 운동 궤적을 포함한 4D 재구성을 위한 새로운 방법을 제안한다.

핵심 기여도

핵심 아이디어

기존 방법은 장거리 3D 운동을 효과적으로 모델링하지 못하거나, 템플릿에 의존하거나, 2D 기반의 짧은 운동만 추정하는 한계가 있었다. 본 연구는 3D 운동이 $\mathbb{SE}(3)$ 기저로 표현될 수 있다는 통찰을 바탕으로, 장거리 3D 운동 궤적을 재구성하는 새로운 접근법을 제안한다. 각 3D 가우시안 포인트는 $\mathbb{SE}(3)$ 기저의 선형 결합으로 표현되어, 장거리 운동을 소프트하게 분해할 수 있다. 또한, 모노카메라 깊이 추정과 2D 트랙 추정을 결합하여 노이즈가 있는 신호를 전역적으로 통합함으로써, 전역 일관성 있는 3D 운동 표현을 구축한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 단일 비디오에서 장거리 3D 운동 궤적을 재구성하는 데 중요한 기여를 하며, 모노카메라 기반의 4D 재구성 문제 해결에 기초를 제공한다. 특히, $\mathbb{SE}(3)$ 기저를 사용한 운동 표현은 기존의 단거리 운동 추정 방식을 벗어나, 장거리 추적의 정확도를 크게 향상시킨다. 그러나 장거리 카메라 시점 변화나 텍스처 없는 영역에서는 실패할 수 있으며, 장거리 추적을 위해 여전히 테스트 타임 최적화가 필요하다는 한계가 있다.

실용적 활용

본 연구는 드론, 로봇, 자율주행 등에서 장거리 3D 운동 추적과 새로운 뷰 합성이 필요한 상황에 적용 가능하다. 또한, 콘텐츠 생성, 증강현실(AR), 영상 분석 등에서 4D 재구성 기술의 활용 범위를 확장할 수 있다.