ChatPaper.aiChatPaper

密集空間感知的視覺預訓練

Vision Pretraining for Dense Spatial Perception

July 6, 2026
作者: Zelin Fu, Bin Tan, Changjiang Sun, Shaohui Liu, Kecheng Zheng, Yinghao Xu, Xing Zhu, Yujun Shen, Nan Xue
cs.AI

摘要

密集空間感知對物理智慧至關重要,視覺系統需從像素觀測中恢復結構化、公制且可操作的表徵。現代視覺基礎模型傾向於優先確保語義不變性,卻常以犧牲細緻空間理解為代價。本研究從邊界視角切入探討視覺預訓練,其核心假設在於:邊界與形狀不連續性能提供感知幾何屬性的關鍵線索。具體而言,我們提出遮罩邊界建模(Masked Boundary Modeling)——一種自我監督學習範式,透過動態學習次像素邊界表徵,並將此過程發現的邊界標記作為遮罩目標,促進密集視覺標記學習。透過擴展此框架,我們開發出LingBot-Vision,並以強基線DINOv3為對照,驗證其於多樣下游視覺任務中的效能。值得注意的是,LingBot-Vision推動了深度補償任務從LingBot-Depth 1.0到LingBot-Depth 2.0的躍進,從而實現增強的深度估計——此為具身人工智慧的關鍵支柱。我們的研究揭示:邊界建模不僅止於簡單線段,更能作為學習空間結構化視覺表徵的可擴展預訓練原則。
English
Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observations. Modern visual foundation models tend to prioritize semantic invariance, often at the expense of detailed spatial understanding. In this work, we study vision pretraining through a boundary-centric lens, motivated by the premise that boundaries and shape discontinuities offer essential cues for perceiving geometric properties. Concretely, we propose masked boundary modeling, a self-supervised paradigm that dynamically learns sub-pixel boundary representations and subsequently leverages the discovered boundary-bearing tokens as masked targets to facilitate dense visual token learning. By scaling this framework, we develop LingBot-Vision and demonstrate its efficacy across a diverse set of downstream vision tasks with DINOv3 as a strong baseline. Remarkably, LingBot-Vision drives the progression from LingBot-Depth 1.0 to LingBot-Depth 2.0 for depth completion, and thereby yields enhanced depth estimation, a key pillar for embodied artificial intelligence. Our findings reveal that boundary modeling goes beyond simple line segments and instead serves as a scalable pretraining principle for learning spatially structured visual representations.