ChatPaper.aiChatPaper

密な空間知覚のための視覚事前学習

Vision Pretraining for Dense Spatial Perception

July 6, 2026
著者: Zelin Fu, Bin Tan, Changjiang Sun, Shaohui Liu, Kecheng Zheng, Yinghao Xu, Xing Zhu, Yujun Shen, Nan Xue
cs.AI

要旨

密な空間認識は物理的知能にとって不可欠であり、視覚システムはピクセル観測から構造化された計測可能で行動可能な表現を復元することが期待される。現代の視覚基盤モデルは意味的不変性を優先する傾向があり、その代償として詳細な空間理解が損なわれることが多い。本研究では、境界と形状の不連続性が幾何学的特性を知覚するための本質的な手がかりを提供するという前提に基づき、境界中心の視点から視覚事前学習を考察する。具体的には、自己教師あり学習パラダイムであるマスク境界モデリングを提案する。これは、動的にサブピクセル境界表現を学習し、その後発見された境界を含むトークンをマスク対象として活用することで、密な視覚トークン学習を促進するものである。この枠組みをスケールすることで、我々はLingBot-Visionを開発し、その有効性をDINOv3を強力なベースラインとする多様な下流視覚タスクにおいて実証する。特筆すべきは、LingBot-Visionが深度補完のためのLingBot-Depth 1.0からLingBot-Depth 2.0への進化を促進し、それによって具現化人工知能の重要な基盤である深度推定を向上させた点である。本研究の成果は、境界モデリングが単なる線分を超え、空間的に構造化された視覚表現を学習するためのスケーラブルな事前学習原理として機能することを示している。
English
Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observations. Modern visual foundation models tend to prioritize semantic invariance, often at the expense of detailed spatial understanding. In this work, we study vision pretraining through a boundary-centric lens, motivated by the premise that boundaries and shape discontinuities offer essential cues for perceiving geometric properties. Concretely, we propose masked boundary modeling, a self-supervised paradigm that dynamically learns sub-pixel boundary representations and subsequently leverages the discovered boundary-bearing tokens as masked targets to facilitate dense visual token learning. By scaling this framework, we develop LingBot-Vision and demonstrate its efficacy across a diverse set of downstream vision tasks with DINOv3 as a strong baseline. Remarkably, LingBot-Vision drives the progression from LingBot-Depth 1.0 to LingBot-Depth 2.0 for depth completion, and thereby yields enhanced depth estimation, a key pillar for embodied artificial intelligence. Our findings reveal that boundary modeling goes beyond simple line segments and instead serves as a scalable pretraining principle for learning spatially structured visual representations.