ChatPaper.aiChatPaper

見せよ、語るな:LLMテキストではなく生成ピクセルにおける空間認知の評価

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

July 23, 2026
著者: Xu Wang, Kaixiang Yao, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Wenqi Zhang, Xuhong Zhang
cs.AI

要旨

空間知能は、エージェントが静的な意味理解から物理世界との相互作用へと移行するために不可欠である。多くの空間タスクは連続的な視覚シーンに基づいており、位置、領域、経路は、正確な座標や離散的なテキスト記号を報告するよりも、指差し、マーキング、描画によってより自然に表現される。しかしながら、既存の空間推論ベンチマークは通常、座標、選択肢、またはテキストを必要とし、画像生成モデルにとって解答インターフェースの不一致を生み出している。そのため、画像生成モデルが空間的判断をピクセル空間に直接外在化できるにもかかわらず、テキスト出力の視覚言語モデル(VLM)と同じタスク意味論の下で評価することが困難となっている。我々は、ProVisE(Protocolized Visual Evaluation:プロトコル化視覚評価)を提案する。これは、ベンチマークに依存しないフレームワークであり、画像生成モデルからプロトコルに制約された視覚的解答を引き出し、元の評価指標と互換性のある構造化予測に解析する。ProVisEはまた、新しいベンチマークのためのタスク固有のプロトコルを構築・検証するAgentic Builderを含む。さらに、14の空間サブタスク、4つの能力レベル、多様な解答形式にわたる470サンプルからなる厳選された診断ベンチマークであるSpatialGen-Benchを導入する。我々は、代表的テキスト出力VLMと画像生成モデルを統一設定で評価し、6つの外部空間ベンチマークでAgenticプロトコル構築を検証する。結果は、空間的解答がピクセル空間に直接外在化できる場合には画像生成モデルが競争力を持つ一方、テキスト出力VLMは構成的空間推論において明確な優位性を維持することを示している。これらの知見は、ピクセル空間表現とテキストベース推論の相補的な強みを明らかにし、画像生成モデルにおける空間認知研究のための評価指標互換テストベッドを確立する。
English
Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models. This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space. We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks. We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.