ChatPaper.aiChatPaper

장기간 스트리밍 3D 재구성을 위한 로컬 컨텍스트의 재검토

Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction

August 27, 2026
저자: Jiarong Han, Jincheng Xiong, Yuzhou Liu, Linzhe Shi, Changjie Wu, Ning Guo, Mu Xu, Hang Zhang, Ming Qian
cs.AI

초록

매우 긴 비디오로부터의 스트리밍 3D 재구성은 제한된 메모리와 연산 조건에서 카메라 모션과 장면 기하를 온라인으로 추정할 것을 요구한다. 초기 스트리밍 모델들은 유한 컨텍스트 버퍼나 간결한 순환 상태를 사용하여 인과적이고 비용이 제한된 추론을 달성하지만, 시퀀스가 길어질수록 추정치가 종종 저하된다. 최근 방법들은 단기 컨텍스트를 영구적 또는 다중 수준의 장기 메모리와 결합함으로써 장기 안정성을 개선한다. 우리는 다른 경로를 추구한다: 학습된 시간적 상태를 엄격히 지역적으로 유지하고, 그 목표가 시퀀스 길이와 무관하게 유지되는 예측을 공식화한다. 우리는 이전 11개 프레임의 KV 특징만 캐시하는 간단한 스트리밍 모델인 ABot-Recon을 제시한다. 이 모델은 현재 카메라 좌표계의 포인트 맵과 인접 프레임 간 상대 포즈를 예측한다. 이러한 예측은 기준 좌표계의 변경 하에서 등변량을 유지하며, 전역 포즈와 기하는 순차적 합성을 통해 복원된다. 누적 드리프트를 줄이기 위해, 경량의 시간적 리파이너가 최근 시각 및 모션 컨텍스트를 사용하여 상대 회전을 개선하고, 합성 인식 포즈 손실이 다단계 포즈 합성을 지도한다. 까다로운 장기 시퀀스 벤치마크에 대한 광범위한 평가는 우리의 지역 컨텍스트 접근 방식의 우수한 장기 성능을 입증한다. Oxford Spires에서 ABot-Recon은 ATE 4.35m와 RPE-R 0.12°를 달성하여, 최고의 기존 결과 대비 두 오차를 모두 약 40% 줄인다.
English
Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online under bounded memory and computation. Early streaming models achieve causal, bounded-cost inference using finite context buffers or compact recurrent states, yet their estimates often deteriorate as sequences grow. Recent methods improve long-horizon stability by coupling short-range context with persistent or multi-level long-range memory. We pursue a different route: we keep the learned temporal state strictly local and formulate predictions whose targets remain independent of sequence length. We present ABot-Recon, a simple streaming model that caches KV features from only the preceding 11 frames. It predicts a point map in the current camera coordinate system together with an adjacent-frame relative pose. These predictions remain equivariant under changes of reference frame, and global poses and geometry are recovered through sequential composition. To reduce accumulated drift, a lightweight temporal refiner improves relative rotations using recent visual and motion context, while a composition-aware pose loss supervises multi-step pose composition. Extensive evaluations on challenging long-sequence benchmarks demonstrate the superior long-horizon performance of our local-context approach. On Oxford Spires, ABot-Recon achieves an ATE of 4.35 m and an RPE-R of 0.12^circ, reducing both errors by approximately 40\% relative to the best prior results.