ChatPaper.aiChatPaper

RoboSPA: VLA 모델은 단순한 장면과 단기 과제를 넘어설 수 있는가?

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

September 4, 2026
저자: Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang
cs.AI

초록

비전-언어-행동(Vision-Language-Action, VLA) 모델은 언어로 조건화된 로봇 조작에서 유망한 진전을 보여 왔다. 그러나 기존 데이터셋과 벤치마크는 주로 미리 정의된 설정에서 작업 완료를 평가하므로, 공간적 및 절차적 복잡성이 증가하는 상황에서의 모델 추론에 대한 통찰을 제한적으로만 제공한다. 우리는 VLA 모델의 체화된 추론을 진단하기 위한 대규모 로봇 조작 데이터셋 및 벤치마크인 RoboSPA(Robot Spatial-Procedural Assessment)를 제안한다. RoboSPA는 세밀한 공간 추론(Fine-Grained Spatial Reasoning)과 장기 지평 절차 계획(Long-Horizon Procedural Planning)이라는 두 가지 핵심 차원에 초점을 맞추며, 10개 작업 범주와 56개 기본 작업을 포괄한다. 각 작업은 다섯 가지 난이도 수준에 걸쳐 인스턴스화되어, 공간적 모호성과 절차적 복잡성이 점차 증가하는 280개 변형을 생성한다. 우리는 여러 구현체와 다양한 장면에 걸쳐 527K개의 궤적을 수집한다. 이진 성공률을 넘어, RoboSPA는 보다 세부적인 평가를 위한 진단 지표를 도입한다. 대표적인 VLA 모델에 대한 실험은 현재 시스템이 복잡한 공간 관계, 정밀한 저수준 실행, 메모리 집약적 계획에 여전히 어려움을 겪고 있음을 보여준다. 이러한 결과는 RoboSPA가 더 유능하고 신뢰할 수 있으며 일반화 가능한 체화된 에이전트를 개발하기 위한 도전적인 진단 벤치마크임을 입증한다. 우리의 데이터와 코드는 https://github.com/fanzhenxuan/RoboSPA에서 확인할 수 있다.
English
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce RoboSPA (Robot Spatial-Procedural Assessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. RoboSPA focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, RoboSPA introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https://github.com/fanzhenxuan/RoboSPA.