ChatPaper.aiChatPaper

空間における自己:UAV身体化知能における自己認識と空間認知のベンチマーキング

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

July 14, 2026
著者: Zhishan Zou, Guoyan Sun, Zhiwei Wei, Jiancheng Pan, Yujie Li, Mugen Peng, Wenjia Xu
cs.AI

要旨

自律型UAVシステムは、複雑な実世界環境での動作実現のために、多様なマルチモーダル大規模言語モデル(MLLM)への依存を強めている。このような身体化シナリオでは、周囲空間の理解に加え、エージェント自身の一貫した表現を維持することが求められる。しかし、既存のUAV向けアプローチやベンチマークは環境中心の設計が大半であり、主に空間理解タスクに焦点を当てており、エージェントの自己認識は暗黙的に扱われるに留まっている。このギャップを埋めるため、本稿では統一的な「空間内の自己」形式に基づくUAVシナリオ向け身体化空間知能評価ベンチマークであるSIS-Benchを導入する。SIS-Benchは、空間と自己という補完的な二つの次元、ならびに知覚、記憶、推論からなる三層の階層に沿って評価を構成する。本ベンチマークは、専門家検証を伴うタスク条件付き構築パイプラインを通じて1,646本の実世界UAV映像から派生した13タスク、4,856組の質問回答対を収録する。 広範な評価から、現行のMLLMは動的かつエージェント中心のプロセスをモデル化する上で根本的な限界を持つことが明らかとなった。特に、空間認知と自己認識の間には顕著な不均衡が認められ、認知レベルが上がるにつれて性能が段階的に低下する傾向が観察された。これらの知見に動機づけられ、我々はさらに、オプティカルフローと視覚特徴融合を通じて自己関連の動的要素を組み込んだ動作認識表現を探求する。実験結果は、エージェントの動作をモデル化することで、空間認知のみならず自己認識においても知覚と記憶の性能が一貫して向上し、下流のUAV意思決定タスクへと一般化されることを示す。 本研究の結果は、身体化空間知能の進展における自己認識の重要性を強調するとともに、動作認識型「空間内の自己」モデリングのための新たなベンチマークと実証的証拠を提供するものである。
English
Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a coherent representation of the agent itself. However, existing UAV-oriented approaches and benchmarks remain largely environment-centric, primarily focusing on spatial understanding tasks, with the agent's self-awareness remaining implicit. To address this gap, we introduce SIS-Bench, a benchmark for evaluating embodied spatial intelligence in UAV scenarios under a unified self-in-space formulation. SIS-Bench organizes evaluation along two complementary dimensions, space and self, and a three-level hierarchy of perception, memory, and reasoning. It contains 4,856 question--answer pairs across 13 tasks derived from 1,646 real-world UAV videos through a task-conditioned construction pipeline with expert verification.Extensive evaluations reveal that current MLLMs exhibit fundamental limitations in modeling dynamic and agent-centered processes. In particular, we observe a clear imbalance between spatial cognition and self-awareness, as well as a progressive performance degradation across cognitive levels.Motivated by these findings, we further explore a motion-aware representation that incorporates self-related dynamics through optical flow and visual feature fusion. Experimental results show that modeling agent motion consistently improves perception and memory performance, not only in spatial cognition but also in self-awareness, and generalizes to downstream UAV decision-making tasks.Our results highlight the importance of self-awareness for advancing embodied spatial intelligence, and provide both a new benchmark and empirical evidence for motion-aware self-in-space modeling.