ChatPaper.aiChatPaper

自几何:用于几何一致3D视觉基础模型的无真值即插即用测试时自适应

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

August 11, 2026
作者: Seokhyun Youn, Dahyeon Kye, Sung-Ho Bae, Jihyong Oh
cs.AI

摘要

近期视觉基础模型(VFMs)在单次前向传播中预测深度、相机位姿和点图,无需逐场景优化,实现了强大的泛化能力。然而,强制显式的多视图几何一致性(例如通过光束法平差)计算成本高昂,因此在VFM预训练期间未施加,从而可能出现这种不一致性。为了解决此问题,从模型输出(如点图、特征)导出的隐式自一致性,尽管先前工作是在测试时强制执行,但带来的性能提升本质上有限,尤其是在预训练VFM高度不准确的场景中。与这种隐式信号相比,我们提出了自几何(Self-Geometry),一种即插即用的测试时自适应流水线,直接利用2D像素对应关系作为伪真值施加显式的多视图几何约束。我们提出的Self-Geometry包括:几何解耦优化(Geometric Disentanglement Optimization),结合多视图一致性和对极一致性损失,并通过梯度解耦防止梯度冲突;帧角度邻域(Frame Angular-Neighbor),一种基于SO(3)测地距离的视图采样器,用于轻量施加这些约束;以及轻量级TTA,通过LoRA自适应VFMs。我们的方法在六个VFM(VGGT、π³、DA3-Giant/Large/Base/Small)和四个基准(7Scenes、ETH3D、ScanNet++、HiRoom)上的位姿和几何估计均实现了一致的改进。
English
Recent Vision Foundation Models (VFMs) predict depth, camera pose, and pointmap in a single forward pass without per-scene optimization, achieving strong generalization. However, enforcing explicit multi-view geometric consistency, e.g., through bundle adjustment, is computationally costly and is thus not imposed during VFM pretraining, so such inconsistency can arise. To address this, implicit self-consistency derived from model outputs (e.g., pointmaps, features), though enforced at test-time in prior work, delivers inherently limited performance gain, especially on scenes where the pretrained VFM is highly inaccurate. In contrast to this implicit signal, we propose Self-Geometry, a plug-and-play test-time adaptation pipeline that directly imposes explicit multi-view geometric constraints using 2D pixel correspondences as pseudo ground-truth. Our proposed Self-Geometry consists of Geometric Disentanglement Optimization, which combines Multi-View Consistency and Epipolar Consistency losses with Gradient Disentanglement to prevent gradient conflict; Frame Angular-Neighbor, a view sampler based on SO(3) geodesic distances for lightly imposing these constraints; and Lightweight TTA, which adapts VFMs via LoRA. Our method achieves consistent improvements in both pose and geometry estimation across six VFMs (VGGT, π^3, DA3-Giant/Large/Base/Small) and four benchmarks (7Scenes, ETH3D, ScanNet++, HiRoom).