밀집 공간 지각을 위한 비전 사전 학습
Vision Pretraining for Dense Spatial Perception
July 6, 2026
저자: Zelin Fu, Bin Tan, Changjiang Sun, Shaohui Liu, Kecheng Zheng, Yinghao Xu, Xing Zhu, Yujun Shen, Nan Xue
cs.AI
초록
밀집 공간 지각은 물리적 지능에 필수적이며, 시각 시스템은 픽셀 관측값으로부터 구조화되고 정량적이며 실행 가능한 표현을 복원할 것으로 기대된다. 현대의 시각 기반 모델은 종종 상세한 공간 이해를 희생하면서 의미 불변성을 우선시하는 경향이 있다. 본 연구에서는 경계와 형상 불연속성이 기하학적 속성을 인지하는 데 필수적인 단서를 제공한다는 전제에 기반하여, 경계 중심 렌즈를 통해 시각 사전 학습을 조사한다. 구체적으로, 서브픽셀 경계 표현을 동적으로 학습하고, 이후 발견된 경계 포함 토큰을 마스크 대상으로 활용하여 밀집 시각 토큰 학습을 촉진하는 자기지도 패러다임인 마스크 경계 모델링을 제안한다. 이 프레임워크를 확장하여 LingBot-Vision을 개발하고, DINOv3를 강력한 기준선으로 삼아 다양한 하위 시각 과제에서 효용성을 입증한다. 주목할 점은, LingBot-Vision이 깊이 완성을 위한 LingBot-Depth 1.0에서 LingBot-Depth 2.0으로의 발전을 주도하여, 구현 인공지능의 핵심 축인 향상된 깊이 추정을 제공한다는 것이다. 우리의 발견은 경계 모델링이 단순한 선분을 넘어 공간적으로 구조화된 시각 표현을 학습하기 위한 확장 가능한 사전 학습 원리로 기능함을 보여준다.
English
Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observations. Modern visual foundation models tend to prioritize semantic invariance, often at the expense of detailed spatial understanding. In this work, we study vision pretraining through a boundary-centric lens, motivated by the premise that boundaries and shape discontinuities offer essential cues for perceiving geometric properties. Concretely, we propose masked boundary modeling, a self-supervised paradigm that dynamically learns sub-pixel boundary representations and subsequently leverages the discovered boundary-bearing tokens as masked targets to facilitate dense visual token learning. By scaling this framework, we develop LingBot-Vision and demonstrate its efficacy across a diverse set of downstream vision tasks with DINOv3 as a strong baseline. Remarkably, LingBot-Vision drives the progression from LingBot-Depth 1.0 to LingBot-Depth 2.0 for depth completion, and thereby yields enhanced depth estimation, a key pillar for embodied artificial intelligence. Our findings reveal that boundary modeling goes beyond simple line segments and instead serves as a scalable pretraining principle for learning spatially structured visual representations.