Self-Geometry:幾何的に一貫した3Dビジョン基盤モデルのためのGTフリーかつプラグアンドプレイなテスト時適応
Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models
August 11, 2026
著者: Seokhyun Youn, Dahyeon Kye, Sung-Ho Bae, Jihyong Oh
cs.AI
要旨
近年の視覚基盤モデル(VFM)は、シーンごとの最適化を必要とせず、単一の順伝播で深度・カメラ姿勢・ポイントマップを予測し、強い汎化性能を達成している。しかし、バンドル調整などを用いた明示的な多視点幾何整合性の強制は計算コストが高く、VFMの事前学習では課されないため、そのような不整合が生じ得る。この問題に対処するため、モデル出力(例:ポイントマップ、特徴量)から導出される暗黙的な自己整合性は、先行研究ではテスト時に強制されるものの、特に事前学習済みVFMの精度が低いシーンでは、本質的に限られた性能向上しかもたらさない。この暗黙的なシグナルとは対照的に、我々は2次元ピクセル対応を擬似グラウンドトゥルースとして用いて、明示的な多視点幾何制約を直接課すプラグアンドプレイ型のテスト時適応パイプラインであるSelf-Geometryを提案する。提案するSelf-Geometryは、勾配競合を防ぐために勾配分離を併用し、多視点整合性損失とエピポーラ整合性損失を組み合わせた幾何学的分離最適化、SO(3)測地線距離に基づいてこれらの制約を軽く課す視点サンプラーであるFrame Angular-Neighbor、そしてLoRAを介してVFMを適応させる軽量TTAから構成される。本手法は、6つのVFM(VGGT、π^3、DA3-Giant/Large/Base/Small)と4つのベンチマーク(7Scenes、ETH3D、ScanNet++、HiRoom)にわたり、姿勢推定と幾何推定の両方で一貫した改善を達成する。
English
Recent Vision Foundation Models (VFMs) predict depth, camera pose, and pointmap in a single forward pass without per-scene optimization, achieving strong generalization. However, enforcing explicit multi-view geometric consistency, e.g., through bundle adjustment, is computationally costly and is thus not imposed during VFM pretraining, so such inconsistency can arise. To address this, implicit self-consistency derived from model outputs (e.g., pointmaps, features), though enforced at test-time in prior work, delivers inherently limited performance gain, especially on scenes where the pretrained VFM is highly inaccurate. In contrast to this implicit signal, we propose Self-Geometry, a plug-and-play test-time adaptation pipeline that directly imposes explicit multi-view geometric constraints using 2D pixel correspondences as pseudo ground-truth. Our proposed Self-Geometry consists of Geometric Disentanglement Optimization, which combines Multi-View Consistency and Epipolar Consistency losses with Gradient Disentanglement to prevent gradient conflict; Frame Angular-Neighbor, a view sampler based on SO(3) geodesic distances for lightly imposing these constraints; and Lightweight TTA, which adapts VFMs via LoRA. Our method achieves consistent improvements in both pose and geometry estimation across six VFMs (VGGT, π^3, DA3-Giant/Large/Base/Small) and four benchmarks (7Scenes, ETH3D, ScanNet++, HiRoom).