SPARGen: 네이티브 멀티모달 생성을 통한 공간 지각과 추론의 통합
SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation
August 14, 2026
저자: Jinsheng Quan, Jianhua Li, Siyi Xie, Xuanke Shi, Kewang Deng, Zukai Chen, Feifei Shao, Lei Yang, Quan Wang, Yawei Luo
cs.AI
초록
시각적 관측으로부터의 공간 지각과 추론은 기하학적 구조를 복원하고, 대응 관계를 수립하며, 공간 관계를 이해하는 것을 요구한다. 기존 접근법은 일반적으로 이러한 능력을 작업별 아키텍처나 외부 기하 모듈을 통해 개별적으로 처리하여, 동일한 물리적 장면의 상보적 표현 간 지식 전이를 제한한다. 우리는 3D 재구성, 밀집 대응, 공간 추론을 지시 조건부 생성 작업으로 설정하는 통합 멀티모달 프레임워크인 SPARGen을 제안한다. SPARGen은 간결한 구조적·언어적 출력을 토큰 시퀀스로 직렬화하는 동시에, 밀집 기하 필드를 이미지 정렬 형태로 생성함으로써, 공간 감독이 자체 멀티모달 생성 모델 내에서 공유 표현을 공동으로 형성하게 한다. 3D 재구성, 대응, 공간 추론 벤치마크에 걸친 실험은 SPARGen이 단일 자체 멀티모달 생성 프레임워크 내에서 이질적인 공간 작업들에 대해 경쟁력 있는 성능을 달성함을 보여준다.
English
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specific architectures or external geometric modules, limiting knowledge transfer among complementary representations of the same physical scene. We introduce SPARGen, a unified multimodal framework that casts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation tasks. SPARGen serializes compact structured and linguistic outputs as token sequences while generating dense geometric fields in image-aligned forms, enabling spatial supervision to jointly shape shared representations within a native multimodal generative model. Experiments across benchmarks for 3D reconstruction, correspondence, and spatial reasoning show that SPARGen achieves competitive performance across heterogeneous spatial tasks within a single native multimodal generative framework.