空间中的自我:无人机具身智能中自我意识与空间认知的基准评估
Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
July 14, 2026
作者: Zhishan Zou, Guoyan Sun, Zhiwei Wei, Jiancheng Pan, Yujie Li, Mugen Peng, Wenjia Xu
cs.AI
摘要
自主无人机系统日益依赖多模态大语言模型(MLLMs)在复杂真实环境中运行。这类具身场景不仅需要理解周围空间,还需维持对代理自身的连贯表征。然而,现有面向无人机的方法与基准仍以环境为中心,主要聚焦空间理解任务,而代理的自我意识则长期隐而未现。为填补这一空白,我们提出SIS-Bench——一个基于统一“空间中的自我”框架、评估无人机场景中具身空间智能的基准。SIS-Bench沿空间与自我两个互补维度,及感知、记忆与推理三层认知层级组织评估体系。它包含从1,646段真实无人机视频中通过任务条件构建管道与专家验证生成的4,856个问答对,覆盖13项任务。广泛评估表明,当前MLLM在建模动态与以代理为中心的过程方面存在根本性局限。具体而言,我们观察到空间认知与自我意识之间明显失衡,以及认知层级间性能逐步退化。基于这些发现,我们进一步探索了一种融合光流与视觉特征的自我相关动态运动感知表征。实验结果显示,对代理运动进行建模能持续提升感知与记忆性能——不仅在空间认知方面,更在自我意识层面——并泛化至下游无人机决策任务。我们的研究结果凸显了自我意识对推进具身空间智能的重要性,并为运动感知的“空间中的自我”建模提供了新基准与实证证据。
English
Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a coherent representation of the agent itself. However, existing UAV-oriented approaches and benchmarks remain largely environment-centric, primarily focusing on spatial understanding tasks, with the agent's self-awareness remaining implicit. To address this gap, we introduce SIS-Bench, a benchmark for evaluating embodied spatial intelligence in UAV scenarios under a unified self-in-space formulation. SIS-Bench organizes evaluation along two complementary dimensions, space and self, and a three-level hierarchy of perception, memory, and reasoning. It contains 4,856 question--answer pairs across 13 tasks derived from 1,646 real-world UAV videos through a task-conditioned construction pipeline with expert verification.Extensive evaluations reveal that current MLLMs exhibit fundamental limitations in modeling dynamic and agent-centered processes. In particular, we observe a clear imbalance between spatial cognition and self-awareness, as well as a progressive performance degradation across cognitive levels.Motivated by these findings, we further explore a motion-aware representation that incorporates self-related dynamics through optical flow and visual feature fusion. Experimental results show that modeling agent motion consistently improves perception and memory performance, not only in spatial cognition but also in self-awareness, and generalizes to downstream UAV decision-making tasks.Our results highlight the importance of self-awareness for advancing embodied spatial intelligence, and provide both a new benchmark and empirical evidence for motion-aware self-in-space modeling.