ChatPaper.aiChatPaper

長期ストリーミング3D再構成のための局所的文脈の再考

Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction

August 27, 2026
著者: Jiarong Han, Jincheng Xiong, Yuzhou Liu, Linzhe Shi, Changjie Wu, Ning Guo, Mu Xu, Hang Zhang, Ming Qian
cs.AI

要旨

極めて長いビデオからのストリーミング3次元再構成では、制限されたメモリと計算量の下でカメラ運動とシーン形状をオンラインで推定する必要がある。初期のストリーミングモデルは、有限コンテキストバッファやコンパクトなリカレント状態を用いて因果的かつコスト有界な推論を実現するが、その推定結果はシーケンスが長くなるにつれてしばしば劣化する。最近の手法は、短距離コンテキストと永続的または多レベルの長距離メモリを組み合わせることで、長期的な安定性を向上させている。我々は別のアプローチを取る。学習された時間状態を厳密に局所に保ち、そのターゲットがシーケンス長に依存しない予測を定式化する。本稿では、直前の11フレームのみからKV特徴をキャッシュするシンプルなストリーミングモデルABot-Reconを提案する。これは、現在のカメラ座標系におけるポイントマップと、隣接フレーム間の相対姿勢を予測する。これらの予測は基準座標系の変更に対して等変であり、グローバルな姿勢と形状は逐次合成によって復元される。蓄積ドリフトを低減するため、軽量な時間リファイナーが最近の視覚・動きコンテキストを用いて相対回転を改善し、さらに合成を考慮したポーズ損失が多段階のポーズ合成を監視する。挑戦的な長いシーケンスのベンチマークにおける広範な評価により、我々の局所コンテキストアプローチが優れた長期的性能を持つことが実証される。Oxford Spiresでは、ABot-ReconはATE 4.35 m、RPE-R 0.12°を達成し、従来の最良結果と比較して両方の誤差を約40%削減する。
English
Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online under bounded memory and computation. Early streaming models achieve causal, bounded-cost inference using finite context buffers or compact recurrent states, yet their estimates often deteriorate as sequences grow. Recent methods improve long-horizon stability by coupling short-range context with persistent or multi-level long-range memory. We pursue a different route: we keep the learned temporal state strictly local and formulate predictions whose targets remain independent of sequence length. We present ABot-Recon, a simple streaming model that caches KV features from only the preceding 11 frames. It predicts a point map in the current camera coordinate system together with an adjacent-frame relative pose. These predictions remain equivariant under changes of reference frame, and global poses and geometry are recovered through sequential composition. To reduce accumulated drift, a lightweight temporal refiner improves relative rotations using recent visual and motion context, while a composition-aware pose loss supervises multi-step pose composition. Extensive evaluations on challenging long-sequence benchmarks demonstrate the superior long-horizon performance of our local-context approach. On Oxford Spires, ABot-Recon achieves an ATE of 4.35 m and an RPE-R of 0.12^circ, reducing both errors by approximately 40\% relative to the best prior results.