ChatPaper.aiChatPaper

SPARGen: ネイティブなマルチモーダル生成による空間知覚と推論の統合

SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

August 14, 2026
著者: Jinsheng Quan, Jianhua Li, Siyi Xie, Xuanke Shi, Kewang Deng, Zukai Chen, Feifei Shao, Lei Yang, Quan Wang, Yawei Luo
cs.AI

要旨

視覚的観察からの空間知覚と推論には、幾何学的構造の復元、対応関係の確立、空間関係の理解が必要である。既存のアプローチは通常、タスク固有のアーキテクチャや外部の幾何学モジュールを用いてこれらの能力を別々に扱っており、同一の物理的シーンの相補的な表現間での知識転送を制限している。本稿では、3D再構成、密な対応関係、空間推論を命令条件付き生成タスクとして捉える統一マルチモーダルフレームワークであるSPARGenを紹介する。SPARGenは、コンパクトな構造化出力と言語出力をトークン列として直列化するとともに、画像整列形式で密な幾何学フィールドを生成し、ネイティブなマルチモーダル生成モデル内で共有表現を空間的教師信号が共同して形成することを可能にする。3D再構成、対応関係、空間推論のベンチマークにわたる実験により、SPARGenが単一のネイティブなマルチモーダル生成フレームワーク内で異種の空間タスクにおいて競争力のある性能を達成することを示す。
English
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specific architectures or external geometric modules, limiting knowledge transfer among complementary representations of the same physical scene. We introduce SPARGen, a unified multimodal framework that casts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation tasks. SPARGen serializes compact structured and linguistic outputs as token sequences while generating dense geometric fields in image-aligned forms, enabling spatial supervision to jointly shape shared representations within a native multimodal generative model. Experiments across benchmarks for 3D reconstruction, correspondence, and spatial reasoning show that SPARGen achieves competitive performance across heterogeneous spatial tasks within a single native multimodal generative framework.