ChatPaper.aiChatPaper

空間中的自我:無人機具身智能中的自我意識與空間認知基準測試

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

July 14, 2026
作者: Zhishan Zou, Guoyan Sun, Zhiwei Wei, Jiancheng Pan, Yujie Li, Mugen Peng, Wenjia Xu
cs.AI

摘要

自主无人机系统日益依賴多模態大型語言模型(MLLMs)在複雜的真實環境中運作。這類具身場景不僅需要理解周圍空間,還需維持對代理自身的一致性表徵。然而,現有的無人機導向方法與基準仍以環境為中心,主要聚焦於空間理解任務,代理的自我意識始終處於隱含狀態。為填補此空缺,我們提出SIS-Bench,一個基於「空間中的自我」統一框架來評估無人機場景下具身空間智能的基準。SIS-Bench沿著「空間」與「自我」兩個互補維度,以及「感知」、「記憶」、「推理」三層級結構進行評估。透過任務條件式建構流程及專家驗證,從1,646部真實無人機影片中提取13項任務,共包含4,856組問答對。廣泛評估顯示,當前MLLMs在建模動態及以代理為中心的過程上存在根本限制。尤其,我們觀察到空間認知與自我意識之間明顯失衡,以及認知層級間的漸進式效能衰退。基於這些發現,我們進一步探索融入自我相關動態的動作感知表徵,該表徵融合光流與視覺特徵。實驗結果顯示,建模代理動作持續提升感知與記憶效能,不僅在空間認知上,也在自我意識上有所改善,並可泛化至下游無人機決策任務。我們的結果突顯自我意識對於推動具身空間智能的重要性,並為動作感知的「空間中的自我」建模提供了新的基準與實證依據。
English
Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a coherent representation of the agent itself. However, existing UAV-oriented approaches and benchmarks remain largely environment-centric, primarily focusing on spatial understanding tasks, with the agent's self-awareness remaining implicit. To address this gap, we introduce SIS-Bench, a benchmark for evaluating embodied spatial intelligence in UAV scenarios under a unified self-in-space formulation. SIS-Bench organizes evaluation along two complementary dimensions, space and self, and a three-level hierarchy of perception, memory, and reasoning. It contains 4,856 question--answer pairs across 13 tasks derived from 1,646 real-world UAV videos through a task-conditioned construction pipeline with expert verification.Extensive evaluations reveal that current MLLMs exhibit fundamental limitations in modeling dynamic and agent-centered processes. In particular, we observe a clear imbalance between spatial cognition and self-awareness, as well as a progressive performance degradation across cognitive levels.Motivated by these findings, we further explore a motion-aware representation that incorporates self-related dynamics through optical flow and visual feature fusion. Experimental results show that modeling agent motion consistently improves perception and memory performance, not only in spatial cognition but also in self-awareness, and generalizes to downstream UAV decision-making tasks.Our results highlight the importance of self-awareness for advancing embodied spatial intelligence, and provide both a new benchmark and empirical evidence for motion-aware self-in-space modeling.