SPARGen:透過原生多模態生成統一空間感知與推理
SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation
August 14, 2026
作者: Jinsheng Quan, Jianhua Li, Siyi Xie, Xuanke Shi, Kewang Deng, Zukai Chen, Feifei Shao, Lei Yang, Quan Wang, Yawei Luo
cs.AI
摘要
從視覺觀察中進行空間感知與推理,需要恢復幾何結構、建立對應關係並理解空間關係。現有方法通常使用特定任務的架構或外部幾何模組來分別處理這些能力,限制了同一物理場景中互補表徵之間的知識遷移。我們提出 SPARGen,一個統一的多元模態框架,將三維重建、稠密對應與空間推理重新定義為指令條件化的生成任務。SPARGen 將緊湊的結構化與語言輸出序列化為詞元序列,同時以影像對齊的形式生成稠密幾何場,使空間監督能夠在原生多元模態生成模型中共同塑造共享表徵。在三維重建、對應關係與空間推理等基準上的實驗顯示,SPARGen 在單一原生多元模態生成框架內,於異質空間任務上達到了具有競爭力的表現。
English
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specific architectures or external geometric modules, limiting knowledge transfer among complementary representations of the same physical scene. We introduce SPARGen, a unified multimodal framework that casts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation tasks. SPARGen serializes compact structured and linguistic outputs as token sequences while generating dense geometric fields in image-aligned forms, enabling spatial supervision to jointly shape shared representations within a native multimodal generative model. Experiments across benchmarks for 3D reconstruction, correspondence, and spatial reasoning show that SPARGen achieves competitive performance across heterogeneous spatial tasks within a single native multimodal generative framework.