ChatPaper.aiChatPaper

SIEVE: VLAモデルを用いた模倣学習のための構造認識データ選択

SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models

July 7, 2026
著者: Changti Wu, Bin Yu, Zhaolong Shen, Shijie Lian, Xiaopeng Lin, Cong Huang, Zhirui Zhang, Lei Zhang, Kai Chen
cs.AI

要旨

ビジョン・言語・行動(VLA)モデルは通常、大規模なロボット実演データセットを用いた模倣学習によって訓練される。しかし、データ量を増やしても、冗長性、ノイズ、不均一なカバレッジのために必ずしもより優れた方策が得られるわけではない。既存のデータ選択手法は、実演データを軌跡レベルまたは状態行動レベルで評価することが多く、長期行動を構成する再利用可能な構造を見落としている。本論文では、VLA模倣学習のための構造認識型データ選択手法であるSIEVEを提案する。SIEVEは、実演データを再利用可能なプリミティブと遷移インターフェースの構成とみなす。まず、セグメント化された軌跡から視覚運動プリミティブを発見し、次に収穫逓減の条件下で再利用を考慮した構造的露出を最大化することにより、選択予算を構成パターンに割り当てる。最後に、各構成パターンバケット内でメドイド軌跡を選択し、中心的で安定した模倣に適した実演データを保持する。複数のデータセット、ベンチマーク、VLAモデルにわたる実験により、SIEVEが競合するデータ選択ベースラインを一貫して上回ることが示された。注目すべきことに、SIEVEは実演データの50%と訓練ステップの50%のみを使用しながら、全データ訓練を上回ることができる。これは、プリミティブと遷移を通じて捉えられた再利用可能な構造が、効率的なVLA模倣学習にとって重要な信号であることを示唆している。
English
Vision-Language-Action (VLA) models are typically trained by imitation learning on large-scale robot demonstration datasets, but more data does not necessarily yield better policies due to redundancy, noise, and uneven coverage. Existing data selection methods often assess demonstrations at either the trajectory or state-action level, missing the reusable structures that compose long-horizon behaviors. In this paper, we propose SIEVE, a structure-aware data selection method for VLA imitation learning. SIEVE views demonstrations as compositions of reusable primitives and transition interfaces. It first discovers visuo-motor primitives from segmented trajectories, then allocates selection budgets to composition patterns by maximizing reuse-aware structural exposure under diminishing returns. Finally, it selects medoid trajectories within each composition-pattern bucket to retain central, stable, and imitation-friendly demonstrations. Experiments across multiple datasets, benchmarks, and VLA models show that SIEVE consistently outperforms competitive data selection baselines. Notably, SIEVE can surpass full-data training while using only 50% of demonstrations and 50% of training steps, suggesting that reusable structure, captured through primitives and transitions, is an important signal for efficient VLA imitation learning.