ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

Xiaozhong Lyu, Gen Li, Zhiyin Qian, Xucong Zhang, Marc Pollefeys, Siyu Tang

arXiv:2607.17790 · 2026-07-21 공개 · arXiv · PDF

egocentric-video hot3d masked-generative-transformer holoassist aria-digital-twin taco egocentric-depth camera-tracking

Abstract

Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desirable. However, existing approaches often rely on auxiliary inputs such as pre-computed camera trajectories, treat scene perception and human ego-motion modeling as separate problems despite their strong interdependency, and suffer from slow inference time. To address these limitations, we present ReViV, the first unified framework for holistic egocentric 4D reconstruction that extracts both viewer and view dynamics from a single monocular RGB video. We formulate the task as learning the full joint probability distribution over multimodal signals, including RGB video, camera trajectory, gaze direction, full-body motion, hand motion, and depth. Powered by a Masked Generative Egocentric Transformer, ReViV operates within a single feed-forward architecture to simultaneously reconstruct the temporally consistent 4D reconstruction across the viewer and the view with fast inference speed. Extensive experiments on diverse benchmarks, including HoloAssist, HOT3D, ARCTIC, Aria Digital Twin, and TACO, demonstrate that ReViV achieves state-of-the-art accuracy and efficiency across holistic ego-body, hand, and gaze reconstruction, camera tracking, while maintaining highly competitive egocentric depth estimation without relying on heavy task-specific priors. Code and models are fully open-sourced: https://reviv4d.github.io/.

한국어 요약

한 줄 요약

ReViV는 단일 모노카메라 영상에서 시청자와 환경의 4D 재구성을 동시에 수행하는 통합 프레임워크로, 다양한 벤치마크에서 최고 성능을 달성한다.

핵심 기여도

핵심 아이디어

기존 연구는 환경 재구성과 시청자 운동 추정을 독립적으로 처리하여 시간적 일관성을 잃는 문제가 있었다. ReViV는 이들을 **결합 확률 분포**로 모델링하여 **동시 재구성**을 가능하게 한다. 이는 시청자와 환경 간의 강한 상호의존성을 고려한 통합 접근법이다. 핵심 아이디어는 **모노카메라 영상에서 시청자(ego-body, hand, gaze)와 환경(카메라 궤적, 깊이)을 동시에 재구성**하는 데 있다. 이는 **Masked Generative Egocentric Transformer (MGET)**를 통해 달성되며, 이 모델은 **랜덤 마스킹된 토큰을 예측하는 방식**으로 다중 모달 신호 간의 관계를 학습한다.

기술적 접근법

주요 결과

의의 및 한계

ReViV는 **단일 모노카메라 영상에서 시청자와 환경을 동시에 재구성**하는 첫 번째 통합 프레임워크로, **4D 환경 이해와 인간 운동 추정**의 통합을 가능하게 한다. 이는 **보조 입력 없이도 높은 정확도와 효율성**을 달성하며, **보조 장비나 사전 계산된 SLAM에 의존하지 않아 실용성**이 높다. 그러나 **모노카메라 영상의 깊이 추정은 스케일 불확실성**이 존재하며, **시청자의 몸은 심각하게 가려져 있어 결정론적 추정이 어려움**. 또한, **복잡한 환경에서는 크로스-모달 의존성이 제한될 수 있음**.

실용적 활용

ReViV는 **증강현실(AR), 가상현실(VR), 보조 로봇, 인간-컴퓨터 상호작용** 등에서 시청자와 환경을 실시간으로 이해해야 하는 시스템에 적용 가능하다. 특히 **단일 카메라 기반의 저비용 장비**에서 **4D 환경 재구성과 시청자 운동 추정**이 필요한 경우에 유용하다.