ChatPaper.aiChatPaper

SIEVE:面向VLA模型的模仿学习中的结构感知数据选择

SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models

July 7, 2026
作者: Changti Wu, Bin Yu, Zhaolong Shen, Shijie Lian, Xiaopeng Lin, Cong Huang, Zhirui Zhang, Lei Zhang, Kai Chen
cs.AI

摘要

视觉-语言-动作(VLA)模型通常通过大规模机器人演示数据集的模仿学习进行训练,但由于数据存在冗余、噪声和覆盖不均,数据量增加并不一定能带来更优的策略。现有数据选择方法通常从轨迹或状态-动作层面评估演示,忽略了构成长期行为结构的可重用模块。本文提出SIEVE,一种面向VLA模仿学习的结构感知数据选择方法。SIEVE将演示视为可重用基元与过渡接口的组合,首先从分段轨迹中发现视觉运动基元,然后在收益递减条件下通过最大化结构重用覆盖度来为组合模式分配选择预算,最后在每个组合模式桶内选取medoid轨迹,保留核心、稳定且利于模仿的演示。在多个数据集、基准测试和VLA模型上的实验表明,SIEVE持续优于竞争性数据选择基线。值得注意的是,SIEVE仅使用50%的演示数据和50%的训练步数即可超越全数据训练效果,这表明通过基元和过渡捕获的可重用结构,是实现高效VLA模仿学习的重要信号。
English
Vision-Language-Action (VLA) models are typically trained by imitation learning on large-scale robot demonstration datasets, but more data does not necessarily yield better policies due to redundancy, noise, and uneven coverage. Existing data selection methods often assess demonstrations at either the trajectory or state-action level, missing the reusable structures that compose long-horizon behaviors. In this paper, we propose SIEVE, a structure-aware data selection method for VLA imitation learning. SIEVE views demonstrations as compositions of reusable primitives and transition interfaces. It first discovers visuo-motor primitives from segmented trajectories, then allocates selection budgets to composition patterns by maximizing reuse-aware structural exposure under diminishing returns. Finally, it selects medoid trajectories within each composition-pattern bucket to retain central, stable, and imitation-friendly demonstrations. Experiments across multiple datasets, benchmarks, and VLA models show that SIEVE consistently outperforms competitive data selection baselines. Notably, SIEVE can surpass full-data training while using only 50% of demonstrations and 50% of training steps, suggesting that reusable structure, captured through primitives and transitions, is an important signal for efficient VLA imitation learning.