展示,而非告知:基于生成像素而非LLM文本的空间认知评估
Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
July 23, 2026
作者: Xu Wang, Kaixiang Yao, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Wenqi Zhang, Xuhong Zhang
cs.AI
摘要
空间智能对智能体从静态语义理解转向与物理世界交互至关重要。许多空间任务植根于连续视觉场景,其中位置、区域和路径通过指向、标记或绘制来表达比报告精确坐标或离散文本符号更自然。然而,现有空间推理基准通常需要坐标、选项或文本,这为图像生成模型造成了答案接口不匹配的问题。这使得在相同任务语义下评估图像生成模型与文本输出视觉语言模型变得困难,尽管前者能够直接在像素空间中外化空间判断。我们提出ProVisE(协议化视觉评估),这是一个与基准无关的框架,它从图像生成模型中引出受协议约束的视觉答案,并将其解析为与原始指标兼容的结构化预测。ProVisE还包含一个智能体式构建器,可为新基准构建并验证任务特定协议。我们进一步引入SpatialGen-Bench,这是一个精心策划的诊断性基准,包含470个样本,涵盖14个空间子任务、四个能力层级和多种答案形式。我们在统一设置下评估了代表性文本输出视觉语言模型和图像生成模型,并在六个外部空间基准上验证了智能体式协议构建的有效性。结果表明,当空间答案可在像素空间中直接外化时,图像生成模型具有竞争力,而文本输出视觉语言模型在组合空间推理方面仍保持明显优势。这些发现揭示了像素空间表达与文本推理的互补优势,并为研究图像生成模型中的空间认知建立了与指标兼容的测试平台。
English
Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models. This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space. We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks. We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.