ChatPaper.aiChatPaper

Self-Geometry:面向幾何一致3D視覺基礎模型的免真值、即插即用測試時自適應

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

August 11, 2026
作者: Seokhyun Youn, Dahyeon Kye, Sung-Ho Bae, Jihyong Oh
cs.AI

摘要

近期的視覺基礎模型(VFM)能夠在單次前向傳播中預測深度、相機姿態與點圖,無需逐場景優化,並展現出強大的泛化能力。然而,強制執行顯式的多視圖幾何一致性(例如透過光束法平差)計算成本高昂,因此在 VFM 預訓練階段並未施加此約束,從而可能導致不一致性的產生。為了解決此問題,先前研究雖在測試時施加源自模型輸出(如點圖、特徵)的隱式自一致性,但其帶來的效能提升本質上有限,尤其當預訓練 VFM 在特定場景中的預測極不準確時更是如此。與此類隱式信號不同,我們提出 Self-Geometry——一個即插即用的測試時自適應流程,直接利用二維像素對應關係作為偽真值,施加顯式的多視圖幾何約束。我們提出的 Self-Geometry 包含三個組成部分:幾何解耦優化(Geometric Disentanglement Optimization),結合多視圖一致性損失與對極一致性損失,並透過梯度解耦(Gradient Disentanglement)避免梯度衝突;幀角度鄰域(Frame Angular-Neighbor),一種基於 SO(3) 測地距離的視圖取樣器,用於輕量地施加上述約束;以及輕量級測試時自適應(Lightweight TTA),透過 LoRA 對 VFM 進行適配。我們的方法在六個 VFM(VGGT、π³、DA3-Giant/Large/Base/Small)與四個基準測試(7Scenes、ETH3D、ScanNet++、HiRoom)上,於姿態估計與幾何估計兩個面向均取得一致的效能提升。
English
Recent Vision Foundation Models (VFMs) predict depth, camera pose, and pointmap in a single forward pass without per-scene optimization, achieving strong generalization. However, enforcing explicit multi-view geometric consistency, e.g., through bundle adjustment, is computationally costly and is thus not imposed during VFM pretraining, so such inconsistency can arise. To address this, implicit self-consistency derived from model outputs (e.g., pointmaps, features), though enforced at test-time in prior work, delivers inherently limited performance gain, especially on scenes where the pretrained VFM is highly inaccurate. In contrast to this implicit signal, we propose Self-Geometry, a plug-and-play test-time adaptation pipeline that directly imposes explicit multi-view geometric constraints using 2D pixel correspondences as pseudo ground-truth. Our proposed Self-Geometry consists of Geometric Disentanglement Optimization, which combines Multi-View Consistency and Epipolar Consistency losses with Gradient Disentanglement to prevent gradient conflict; Frame Angular-Neighbor, a view sampler based on SO(3) geodesic distances for lightly imposing these constraints; and Lightweight TTA, which adapts VFMs via LoRA. Our method achieves consistent improvements in both pose and geometry estimation across six VFMs (VGGT, π^3, DA3-Giant/Large/Base/Small) and four benchmarks (7Scenes, ETH3D, ScanNet++, HiRoom).