공간 내 자아: UAV 체화 지능에서의 자기 인식 및 공간 인지 벤치마킹
Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
July 14, 2026
저자: Zhishan Zou, Guoyan Sun, Zhiwei Wei, Jiancheng Pan, Yujie Li, Mugen Peng, Wenjia Xu
cs.AI
초록
자율 UAV 시스템은 점차 다중 모달 대규모 언어 모델(MLLM)에 의존하여 복잡한 실제 환경에서 작동하고 있다. 이러한 체화된 시나리오는 주변 공간을 이해할 뿐만 아니라, 에이전트 자체에 대한 일관된 표현을 유지하는 것도 요구한다. 그러나 기존의 UAV 중심 접근 방식과 벤치마크는 주로 공간 이해 과제에 초점을 맞춘 환경 중심으로 남아 있으며, 에이전트의 자기 인식은 암묵적으로 처리된다. 이러한 격차를 해소하기 위해, 우리는 통일된 자기-공간(self-in-space) 정식화 하에서 UAV 시나리오의 체화된 공간 지능을 평가하기 위한 벤치마크인 SIS-Bench를 소개한다. SIS-Bench는 공간과 자기라는 두 가지 상호 보완적 차원과 지각, 기억, 추론의 3단계 계층 구조에 따라 평가를 구성한다. 이 벤치마크는 전문가 검증을 포함한 과제 조건화된 구축 파이프라인을 통해 1,646개의 실제 UAV 비디오에서 도출된 13개 과제에 걸쳐 4,856개의 질문-답변 쌍을 포함한다.
광범위한 평가 결과, 현재의 MLLM은 동적이고 에이전트 중심적인 과정을 모델링하는 데 근본적인 한계를 나타냄이 드러났다. 특히, 공간 인지와 자기 인식 사이의 명확한 불균형과 더불어 인지 수준에 따른 점진적인 성능 저하가 관찰되었다. 이러한 발견에 동기 부여되어, 우리는 광학 흐름과 시각 특징 융합을 통해 자기 관련 역학을 통합하는 움직임 인식 표현을 추가로 탐구한다. 실험 결과는 에이전트 움직임을 모델링하는 것이 공간 인지뿐만 아니라 자기 인식에서도 지각 및 기억 성능을 일관되게 향상시키며, 하위 UAV 의사 결정 과제로 일반화됨을 보여준다. 우리의 결과는 체화된 공간 지능을 발전시키기 위한 자기 인식의 중요성을 강조하며, 움직임 인식 자기-공간 모델링을 위한 새로운 벤치마크와 경험적 증거를 함께 제공한다.
English
Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a coherent representation of the agent itself. However, existing UAV-oriented approaches and benchmarks remain largely environment-centric, primarily focusing on spatial understanding tasks, with the agent's self-awareness remaining implicit. To address this gap, we introduce SIS-Bench, a benchmark for evaluating embodied spatial intelligence in UAV scenarios under a unified self-in-space formulation. SIS-Bench organizes evaluation along two complementary dimensions, space and self, and a three-level hierarchy of perception, memory, and reasoning. It contains 4,856 question--answer pairs across 13 tasks derived from 1,646 real-world UAV videos through a task-conditioned construction pipeline with expert verification.Extensive evaluations reveal that current MLLMs exhibit fundamental limitations in modeling dynamic and agent-centered processes. In particular, we observe a clear imbalance between spatial cognition and self-awareness, as well as a progressive performance degradation across cognitive levels.Motivated by these findings, we further explore a motion-aware representation that incorporates self-related dynamics through optical flow and visual feature fusion. Experimental results show that modeling agent motion consistently improves perception and memory performance, not only in spatial cognition but also in self-awareness, and generalizes to downstream UAV decision-making tasks.Our results highlight the importance of self-awareness for advancing embodied spatial intelligence, and provide both a new benchmark and empirical evidence for motion-aware self-in-space modeling.