ChatPaper.aiChatPaper

RoboSPA:VLA 模型能否超越簡單場景與短時程任務?

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

September 4, 2026
作者: Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang
cs.AI

摘要

視覺-語言-動作(VLA)模型在語言條件式機器人操控方面已展現出可觀進展。然而,現有資料集與基準主要評估預定義設定下的任務完成情形,對於模型在空間與程序複雜度日益提高時的推理能力,所能提供的洞見仍相當有限。我們提出 RoboSPA(Robot Spatial-Procedural Assessment),一個大規模機器人操控資料集與基準,用於診斷 VLA 模型中的具身推理。RoboSPA 聚焦於兩個核心面向:細粒度空間推理與長時程程序規劃,涵蓋 10 個任務類別與 56 個基礎任務。每個任務皆實例化為五個難度等級,產生 280 個變體,其空間模糊性與程序複雜度逐步提高。我們收集了橫跨多種具身形態與多樣場景的 527K 條軌跡。除了二元成功率之外,RoboSPA 亦引入診斷性指標,以進行更細緻的評估。對代表性 VLA 模型進行的實驗顯示,現有系統仍在複雜空間關係、精確低階執行與記憶密集型規劃方面面臨困難。這些結果確立 RoboSPA 作為一個具挑戰性的診斷基準,可用於開發能力更強、更可靠且更具泛化性的具身代理。我們的資料與程式碼可在 https://github.com/fanzhenxuan/RoboSPA 取得。
English
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce RoboSPA (Robot Spatial-Procedural Assessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. RoboSPA focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, RoboSPA introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https://github.com/fanzhenxuan/RoboSPA.