ReViV: 単眼の一人称視点動画から視聴者と視野を4Dで再構築する
ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
July 20, 2026
著者: Xiaozhong Lyu, Gen Li, Zhiyin Qian, Xucong Zhang, Marc Pollefeys, Siyu Tang
cs.AI
要旨
エゴセントリックデバイス(例えば、ウェアラブルな前方カメラ)は、人間の観察者と周囲環境との間の継続的なインタラクションを捉える独自の視点を提供する。この4次元表現を再構築可能な、包括的かつ効率的なマルチモーダルモデルが強く望まれている。しかし既存手法は、事前計算されたカメラ軌跡などの補助入力を必要とすることが多く、シーン認識と人間のエゴモーション(自己運動)モデリングを強い相互依存性にもかかわらず別々の問題として扱い、推論時間が遅いという欠点を抱える。これらの制約に対処するため、我々はReViVを提案する。これは単一の単眼RGBビデオから観察者と視野の両方のダイナミクスを抽出する、初の統合型エゴセントリック4次元再構築フレームワークである。本タスクを、RGBビデオ、カメラ軌跡、視線方向、全身動作、手動作、深度を含むマルチモーダル信号に対する完全な同時確率分布の学習として定式化する。マスク生成エゴセントリックトランスフォーマー(Masked Generative Egocentric Transformer)を基盤とし、ReViVは単一のフィードフォワードアーキテクチャ内で動作し、観察者と視野にわたる時間的一貫性のある4次元再構築を高速な推論速度で同時に実現する。HoloAssist、HOT3D、ARCTIC、Aria Digital Twin、TACOなど多様なベンチマークにおける広範な実験により、ReViVは包括的な自己身体、手、視線の再構築、カメラトラッキングにおいて最先端の精度と効率を達成し、重いタスク固有の事前知識に依存することなく、非常に競争力のあるエゴセントリック深度推定を維持することを示す。コードとモデルは完全にオープンソース化されている:https://reviv4d.github.io/。
English
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desirable. However, existing approaches often rely on auxiliary inputs such as pre-computed camera trajectories, treat scene perception and human ego-motion modeling as separate problems despite their strong interdependency, and suffer from slow inference time. To address these limitations, we present ReViV, the first unified framework for holistic egocentric 4D reconstruction that extracts both viewer and view dynamics from a single monocular RGB video. We formulate the task as learning the full joint probability distribution over multimodal signals, including RGB video, camera trajectory, gaze direction, full-body motion, hand motion, and depth. Powered by a Masked Generative Egocentric Transformer, ReViV operates within a single feed-forward architecture to simultaneously reconstruct the temporally consistent 4D reconstruction across the viewer and the view with fast inference speed. Extensive experiments on diverse benchmarks, including HoloAssist, HOT3D, ARCTIC, Aria Digital Twin, and TACO, demonstrate that ReViV achieves state-of-the-art accuracy and efficiency across holistic ego-body, hand, and gaze reconstruction, camera tracking, while maintaining highly competitive egocentric depth estimation without relying on heavy task-specific priors. Code and models are fully open-sourced: https://reviv4d.github.io/.