ChatPaper.aiChatPaper

SPARGen:通过原生多模态生成统一空间感知与推理

SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

August 14, 2026
作者: Jinsheng Quan, Jianhua Li, Siyi Xie, Xuanke Shi, Kewang Deng, Zukai Chen, Feifei Shao, Lei Yang, Quan Wang, Yawei Luo
cs.AI

摘要

从视觉观察中进行空间感知与推理,需要恢复几何结构、建立对应关系并理解空间关系。现有方法通常使用任务特定架构或外部几何模块分别处理这些能力,从而限制了对同一物理场景的互补表示之间的知识迁移。我们提出SPARGen,一个统一的多模态框架,将三维重建、稠密对应和空间推理转化为指令条件生成任务。SPARGen将紧凑的结构化输出和语言输出序列化为token序列,同时以图像对齐的形式生成稠密几何场,使得空间监督能够在原生多模态生成模型内共同塑造共享表示。在三维重建、对应和空间推理基准上的实验表明,SPARGen在单一原生多模态生成框架内,能够在异构空间任务上取得具有竞争力的性能。
English
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specific architectures or external geometric modules, limiting knowledge transfer among complementary representations of the same physical scene. We introduce SPARGen, a unified multimodal framework that casts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation tasks. SPARGen serializes compact structured and linguistic outputs as token sequences while generating dense geometric fields in image-aligned forms, enabling spatial supervision to jointly shape shared representations within a native multimodal generative model. Experiments across benchmarks for 3D reconstruction, correspondence, and spatial reasoning show that SPARGen achieves competitive performance across heterogeneous spatial tasks within a single native multimodal generative framework.