ChatPaper.aiChatPaper

Self-Geometry: 기하학적으로 일관된 3D 비전 파운데이션 모델을 위한 GT-Free 및 플러그 앤 플레이 테스트-타임 적응

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

August 11, 2026
저자: Seokhyun Youn, Dahyeon Kye, Sung-Ho Bae, Jihyong Oh
cs.AI

초록

최근 비전 파운데이션 모델(VFM)은 장면별 최적화 없이 단일 순전파만으로 깊이, 카메라 포즈, 포인트맵을 예측하며 강력한 일반화 성능을 보인다. 그러나 번들 조정(bundle adjustment)과 같은 명시적 다중 시점 기하 일관성의 강제는 계산 비용이 높아 VFM 사전학습 중에는 적용되지 않으므로, 예측 결과에 기하적 비일관성이 발생할 수 있다. 이를 해결하기 위해 기존 연구에서는 모델 출력(예: 포인트맵, 특징)에서 도출된 암시적 자기 일관성을 테스트 시점에 부과하지만, 특히 사전학습된 VFM의 정확도가 낮은 장면에서는 성능 향상이 근본적으로 제한적이다. 이와 같은 암시적 신호와 달리, 본 논문은 2D 픽셀 대응점을 의사 지상 실측으로 사용하여 명시적 다중 시점 기하 제약을 직접 부과하는 플러그 앤 플레이(plug-and-play) 테스트 시 적응 파이프라인인 Self-Geometry를 제안한다. 제안하는 Self-Geometry는 기울기 충돌을 방지하기 위해 기울기 분리(Gradient Disentanglement)를 적용하여 다중 시점 일관성(Multi-View Consistency) 손실과 에피폴라 일관성(Epipolar Consistency) 손실을 결합하는 기하 분리 최적화(Geometric Disentanglement Optimization), SO(3) 측지 거리에 기반하여 이러한 제약을 경량으로 부과하는 뷰 샘플러인 프레임 각도 이웃(Frame Angular-Neighbor), 그리고 LoRA를 통해 VFM을 적응시키는 경량 TTA(Lightweight TTA)로 구성된다. 본 방법은 6개의 VFM(VGGT, π³, DA3-Giant/Large/Base/Small)과 4개의 벤치마크(7Scenes, ETH3D, ScanNet++, HiRoom)에서 포즈 및 지오메트리 추정 성능을 일관되게 향상시킨다.
English
Recent Vision Foundation Models (VFMs) predict depth, camera pose, and pointmap in a single forward pass without per-scene optimization, achieving strong generalization. However, enforcing explicit multi-view geometric consistency, e.g., through bundle adjustment, is computationally costly and is thus not imposed during VFM pretraining, so such inconsistency can arise. To address this, implicit self-consistency derived from model outputs (e.g., pointmaps, features), though enforced at test-time in prior work, delivers inherently limited performance gain, especially on scenes where the pretrained VFM is highly inaccurate. In contrast to this implicit signal, we propose Self-Geometry, a plug-and-play test-time adaptation pipeline that directly imposes explicit multi-view geometric constraints using 2D pixel correspondences as pseudo ground-truth. Our proposed Self-Geometry consists of Geometric Disentanglement Optimization, which combines Multi-View Consistency and Epipolar Consistency losses with Gradient Disentanglement to prevent gradient conflict; Frame Angular-Neighbor, a view sampler based on SO(3) geodesic distances for lightly imposing these constraints; and Lightweight TTA, which adapts VFMs via LoRA. Our method achieves consistent improvements in both pose and geometry estimation across six VFMs (VGGT, π^3, DA3-Giant/Large/Base/Small) and four benchmarks (7Scenes, ETH3D, ScanNet++, HiRoom).