ChatPaper.aiChatPaper

重新審視局部上下文以實現長程流式三維重建

Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction

August 27, 2026
作者: Jiarong Han, Jincheng Xiong, Yuzhou Liu, Linzhe Shi, Changjie Wu, Ning Guo, Mu Xu, Hang Zhang, Ming Qian
cs.AI

摘要

從極長影片進行串流式3D重建需要在有界記憶體與運算下,線上估計相機運動與場景幾何。早期的串流模型利用有限的上下文緩衝區或緊湊的循環狀態,實現因果且有界成本的推論,但其估計往往隨著序列增長而退化。近期方法透過將短程上下文與持久性或多層級長程記憶結合,提升了長時程穩定性。我們採取不同的途徑:我們保持學習的時序狀態嚴格局部化,並建構目標與序列長度無關的預測。我們提出ABot-Recon,一個簡單的串流模型,僅快取前11幀的KV特徵。它預測當前相機座標系中的點圖以及相鄰幀的相對姿態。這些預測在參考幀變換下保持等變,全局姿態與幾何透過序列組合恢復。為了減少累積漂移,輕量化的時序精化器利用近期視覺與運動上下文改善相對旋轉,同時組合感知姿態損失監督多步姿態組合。在具挑戰性的長序列基準上的廣泛評估,證明了我們局部上下文方法優越的長時程效能。在Oxford Spires上,ABot-Recon達到了4.35公尺的ATE與0.12度的RPE-R,相較於先前最佳結果,兩項誤差均降低了約40%。
English
Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online under bounded memory and computation. Early streaming models achieve causal, bounded-cost inference using finite context buffers or compact recurrent states, yet their estimates often deteriorate as sequences grow. Recent methods improve long-horizon stability by coupling short-range context with persistent or multi-level long-range memory. We pursue a different route: we keep the learned temporal state strictly local and formulate predictions whose targets remain independent of sequence length. We present ABot-Recon, a simple streaming model that caches KV features from only the preceding 11 frames. It predicts a point map in the current camera coordinate system together with an adjacent-frame relative pose. These predictions remain equivariant under changes of reference frame, and global poses and geometry are recovered through sequential composition. To reduce accumulated drift, a lightweight temporal refiner improves relative rotations using recent visual and motion context, while a composition-aware pose loss supervises multi-step pose composition. Extensive evaluations on challenging long-sequence benchmarks demonstrate the superior long-horizon performance of our local-context approach. On Oxford Spires, ABot-Recon achieves an ATE of 4.35 m and an RPE-R of 0.12^circ, reducing both errors by approximately 40\% relative to the best prior results.