ChatPaper.aiChatPaper

SIEVE: VLA 모델을 활용한 모방 학습을 위한 구조 인식 데이터 선택

SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models

July 7, 2026
저자: Changti Wu, Bin Yu, Zhaolong Shen, Shijie Lian, Xiaopeng Lin, Cong Huang, Zhirui Zhang, Lei Zhang, Kai Chen
cs.AI

초록

비전-언어-행동(VLA) 모델은 일반적으로 대규모 로봇 시연 데이터셋에 대한 모방 학습을 통해 훈련되지만, 데이터가 많다고 반드시 더 나은 정책으로 이어지지는 않는데, 이는 중복성, 노이즈, 불균일한 커버리지 때문이다. 기존 데이터 선택 방법은 종종 시연을 궤적 또는 상태-행동 수준에서 평가하여, 장기적 행동을 구성하는 재사용 가능한 구조를 간과한다. 본 논문에서는 VLA 모방 학습을 위한 구조 인식 데이터 선택 방법인 SIEVE를 제안한다. SIEVE는 시연을 재사용 가능한 기본 요소(primitives)와 전환 인터페이스(transition interfaces)의 조합으로 본다. 먼저 분할된 궤적에서 시각-운동 기본 요소를 발견한 후, 체감 수익 하에서 재사용 인식 구조적 노출을 최대화하는 방식으로 구성 패턴에 선택 예산을 할당한다. 마지막으로 각 구성 패턴 버킷 내에서 메도이드 궤적을 선택하여 중심적이고 안정적이며 모방에 적합한 시연을 유지한다. 여러 데이터셋, 벤치마크, VLA 모델에 걸친 실험 결과, SIEVE는 경쟁력 있는 데이터 선택 기준 방법들을 일관되게 능가함을 보여준다. 특히, SIEVE는 시연의 50%와 훈련 단계의 50%만 사용하면서도 전체 데이터 훈련을 능가할 수 있는데, 이는 기본 요소와 전환을 통해 포착된 재사용 가능한 구조가 효율적인 VLA 모방 학습을 위한 중요한 신호임을 시사한다.
English
Vision-Language-Action (VLA) models are typically trained by imitation learning on large-scale robot demonstration datasets, but more data does not necessarily yield better policies due to redundancy, noise, and uneven coverage. Existing data selection methods often assess demonstrations at either the trajectory or state-action level, missing the reusable structures that compose long-horizon behaviors. In this paper, we propose SIEVE, a structure-aware data selection method for VLA imitation learning. SIEVE views demonstrations as compositions of reusable primitives and transition interfaces. It first discovers visuo-motor primitives from segmented trajectories, then allocates selection budgets to composition patterns by maximizing reuse-aware structural exposure under diminishing returns. Finally, it selects medoid trajectories within each composition-pattern bucket to retain central, stable, and imitation-friendly demonstrations. Experiments across multiple datasets, benchmarks, and VLA models show that SIEVE consistently outperforms competitive data selection baselines. Notably, SIEVE can surpass full-data training while using only 50% of demonstrations and 50% of training steps, suggesting that reusable structure, captured through primitives and transitions, is an important signal for efficient VLA imitation learning.