RoboSPA:VLAモデルは単純なシーンと短期的タスクを超えられるか?
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
September 4, 2026
著者: Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang
cs.AI
要旨
Vision-Language-Action(VLA)モデルは、言語条件付きロボットマニピュレーションにおいて有望な進展を示している。しかし、既存のデータセットとベンチマークは主に事前定義された設定下でのタスク完了を評価しており、空間的・手続き的複雑性が増大する状況下でのモデル推論についての知見は限られている。我々は、VLAモデルにおける身体性推論を診断するための大規模ロボットマニピュレーション用データセット兼ベンチマークであるRoboSPA(Robot Spatial-Procedural Assessment)を提案する。RoboSPAは、細粒度空間推論と長ホライズン手続き計画という2つの中核次元に焦点を当て、10のタスクカテゴリと56の基本タスクを対象とする。各タスクは5つの難易度レベルにわたってインスタンス化され、空間的曖昧性と手続き的複雑性が増大する280の変種を生み出す。我々は、複数のエンボディメントと多様なシーンにわたる527K件の軌道を収集する。二値的成功率を超えて、RoboSPAはより詳細な評価のための診断指標を導入する。代表的なVLAモデルに関する実験は、現在のシステムが依然として複雑な空間関係、精密な低レベル実行、およびメモリ集約的な計画に苦慮していることを示している。これらの結果は、RoboSPAが、より有能で信頼性が高く、汎化可能な身体性エージェントを開発するための困難な診断ベンチマークであることを確立する。我々のデータとコードは https://github.com/fanzhenxuan/RoboSPA で利用可能である。
English
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce RoboSPA (Robot Spatial-Procedural Assessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. RoboSPA focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, RoboSPA introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https://github.com/fanzhenxuan/RoboSPA.