ChatPaper.aiChatPaper

ReViV: 단안 일인칭 비디오로부터 관찰자와 관찰 장면의 4D 재구성

ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

July 20, 2026
저자: Xiaozhong Lyu, Gen Li, Zhiyin Qian, Xucong Zhang, Marc Pollefeys, Siyu Tang
cs.AI

초록

착용형 전방 카메라와 같은 에고센트릭(자기중심적) 기기는 인간 시청자와 주변 환경 간의 연속적인 상호작용을 포착하는 독특한 시각을 제공합니다. 따라서 이 4D 표현을 재구성할 수 있는 전체적이고 효율적인 다중 모달 모델이 매우 요구됩니다. 그러나 기존 접근법은 사전 계산된 카메라 궤적과 같은 보조 입력에 의존하는 경우가 많으며, 강한 상호 의존성에도 불구하고 장면 인식과 인간의 자기 움직임 모델링을 별개의 문제로 다루고, 느린 추론 시간이라는 문제를 안고 있습니다. 이러한 한계를 해결하기 위해, 우리는 단일 단안 RGB 비디오에서 시청자와 시점 동역학을 모두 추출하는 최초의 통합 프레임워크인 ReViV를 제안합니다. 우리는 이 과제를 RGB 비디오, 카메라 궤적, 시선 방향, 전신 동작, 손 동작, 깊이를 포함한 다중 모달 신호에 대한 전체 결합 확률 분포를 학습하는 것으로 정식화합니다. 마스크 생성 에고센트릭 트랜스포머(Masked Generative Egocentric Transformer)를 기반으로, ReViV는 단일 피드포워드 아키텍처 내에서 작동하여 시청자와 시점에 걸쳐 시간적으로 일관된 4D 재구성을 빠른 추론 속도로 동시에 수행합니다. HoloAssist, HOT3D, ARCTIC, Aria Digital Twin, TACO를 포함한 다양한 벤치마크에 대한 광범위한 실험을 통해, ReViV가 전체적인 에고-바디, 손, 시선 재구성 및 카메라 추적에서 최첨단 정확도와 효율성을 달성하면서도, 무거운 작업별 사전 정보에 의존하지 않고 매우 경쟁력 있는 에고센트릭 깊이 추정을 유지함을 보여줍니다. 코드와 모델은 완전히 오픈소스로 제공됩니다: https://reviv4d.github.io/.
English
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desirable. However, existing approaches often rely on auxiliary inputs such as pre-computed camera trajectories, treat scene perception and human ego-motion modeling as separate problems despite their strong interdependency, and suffer from slow inference time. To address these limitations, we present ReViV, the first unified framework for holistic egocentric 4D reconstruction that extracts both viewer and view dynamics from a single monocular RGB video. We formulate the task as learning the full joint probability distribution over multimodal signals, including RGB video, camera trajectory, gaze direction, full-body motion, hand motion, and depth. Powered by a Masked Generative Egocentric Transformer, ReViV operates within a single feed-forward architecture to simultaneously reconstruct the temporally consistent 4D reconstruction across the viewer and the view with fast inference speed. Extensive experiments on diverse benchmarks, including HoloAssist, HOT3D, ARCTIC, Aria Digital Twin, and TACO, demonstrate that ReViV achieves state-of-the-art accuracy and efficiency across holistic ego-body, hand, and gaze reconstruction, camera tracking, while maintaining highly competitive egocentric depth estimation without relying on heavy task-specific priors. Code and models are fully open-sourced: https://reviv4d.github.io/.