ChatPaper.aiChatPaper

보여주기, 말하지 않기: 생성형 픽셀에서의 공간 인지 평가 - LLM 텍스트보다는

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

July 23, 2026
저자: Xu Wang, Kaixiang Yao, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Wenqi Zhang, Xuhong Zhang
cs.AI

초록

공간 지능은 에이전트가 정적 의미 이해에서 물리적 세계와 상호작용하는 방향으로 나아가는 데 필수적이다. 많은 공간적 과제는 연속적인 시각적 장면에 기반하며, 이러한 장면에서 위치, 영역 및 경로는 정밀한 좌표나 불연속적인 텍스트 기호를 보고하는 것보다는 손가락으로 가리키거나 표시하거나 그림을 그리는 방식으로 더 자연스럽게 표현된다. 그러나 기존의 공간 추론 벤치마크는 일반적으로 좌표, 선택지 또는 텍스트를 요구하므로, 이미지 생성 모델에 대해 답변 인터페이스 불일치를 초래한다. 이는 텍스트 출력 VLM과 동일한 과제 의미론 하에서 이미지 생성 모델을 평가하기 어렵게 만든다. 특히 이미지 생성 모델이 픽셀 공간에서 직접 공간적 판단을 외부화할 수 있는 능력이 있음에도 불구하고 그러하다. 본 연구에서는 ProVisE(Protocolized Visual Evaluation)를 제안한다. 이는 벤치마크에 구애받지 않는 프레임워크로, 이미지 생성 모델로부터 프로토콜에 제약된 시각적 답변을 도출하고, 이를 원래 지표와 호환되는 구조화된 예측으로 파싱한다. ProVisE는 또한 새로운 벤치마크에 대해 과제별 프로토콜을 구축하고 검증하는 에이전틱 빌더(Agentic Builder)를 포함한다. 또한 SpatialGen-Bench를 소개하는데, 이는 14개의 공간적 하위 과제, 4가지 능력 수준, 다양한 답변 형식에 걸쳐 470개의 샘플로 구성된 엄선된 진단 벤치마크이다. 대표적인 텍스트 출력 VLM과 이미지 생성 모델을 통일된 설정에서 평가하고, 6개의 외부 공간 벤치마크에 대해 에이전틱 프로토콜 구축을 검증한다. 결과는 이미지 생성 모델이 공간적 답변을 픽셀 공간에서 직접 외부화할 수 있을 때 경쟁력이 있는 반면, 텍스트 출력 VLM은 구성적 공간 추론에서 명확한 우위를 유지한다는 것을 보여준다. 이러한 발견은 픽셀 공간 표현과 텍스트 기반 추론의 상호 보완적 강점을 드러내며, 이미지 생성 모델의 공간 인지 연구를 위한 지표 호환 테스트베드를 구축한다.
English
Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models. This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space. We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks. We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.