展示,而非講述:以生成像素而非LLM文本評估空間認知
Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
July 23, 2026
作者: Xu Wang, Kaixiang Yao, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Wenqi Zhang, Xuhong Zhang
cs.AI
摘要
空間智能對於智慧體從靜態語義理解邁向與物理世界互動至關重要。許多空間任務根植於連續的視覺場景,其中位置、區域與路徑透過指點、標記或繪製來表達,遠比報告精確座標或離散文字符號更為自然。然而,現有的空間推理基準通常要求座標、選項或文字,為圖像生成模型造成答案介面不匹配的問題。這使得即便圖像生成模型能直接在像素空間中外部化空間判斷,仍難以在與文字輸出視覺語言模型(VLM)相同的任務語義下進行評估。我們提出 ProVisE(Protocolized Visual Evaluation,協議化視覺評估),這是一個與基準無關的框架,能從圖像生成模型中引出受協議約束的視覺答案,並將其解析為與原始度量相容的結構化預測。ProVisE 還包含一個智能代理建構器,可為新基準建構並驗證任務專屬協議。我們進一步引入 SpatialGen-Bench,這是一個精心策劃的診斷性基準,包含 470 個樣本,涵蓋 14 個空間子任務、四個能力層級及多樣化的答案形式。我們在統一設定下評估具代表性的文字輸出 VLM 與圖像生成模型,並在六個外部空間基準上驗證智能代理協議的建構。結果顯示,當空間答案可直接在像素空間中外部化時,圖像生成模型具有競爭力;而文字輸出 VLM 在組合空間推理上仍保有明顯優勢。這些發現揭示了像素空間表達與文字推理的互補優勢,並為研究圖像生成模型中的空間認知建立了度量相容的測試平台。
English
Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models. This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space. We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks. We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.