ChatPaper.aiChatPaper

모방 학습은 정교한 조작에서 시간적 강건성을 보존하는가? 다양한 작업 실행 속도에서의 전문가-학습자 비교

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds

September 1, 2026
저자: Clinton Enwerem, John S. Baras, Calin Belta
cs.AI

초록

모방 학습으로 습득된 정교한 조작 정책은 일반적으로 장면, 객체 또는 지시의 변동에 대한 견고성을 기준으로 평가되지만, 작업 실행 속도에 따른 성능은 덜 자주 검토된다. 이는 학습자가 모방하는 전문가에 비해 얼마나 많은 시간적 견고성을 유지하는지에 대한 의문을 남긴다. 우리는 동일한 작업 조건, 초기 조건 샘플링, 그리고 속도 향상 요인 아래에서 전문가와 학습자를 비교한다. 우리는 로봇이 소포를 획득하고 방향을 바꾼 뒤 삽입하는 접촉이 많은 작업인 ParcelStow에서 이 평가를 구현한다. 시연은 소포 획득 이후의 조작 단계에 대한 속도 향상 범위를 포괄한다. 스크립트 기반 전문가와 전문가의 시연으로 훈련된 트랜스포머 기반 액션 청킹(ACT) 정책은 모두 공칭 속도에서 100% 작업 성공률을 달성한다. 그러나 이들의 성공률은 시연된 범위 내에서 차이를 보인다. 최대 속도에서 전문가의 성공률은 84%이고 ACT의 성공률은 53%이다. 서로 다른 매개변수 초기화를 가진 두 개의 ACT 정책은 유사한 성능 저하를 보인다. 공칭 속도에서 최대 시연 속도까지 각각 34퍼센트 포인트와 48퍼센트 포인트 감소하며, 이는 전문가의 16퍼센트 포인트 감소와 대조된다. 단계별 분석에 따르면, 최대 시연 속도에서 발생한 ACT의 실패 47건 중 35건이 삽입 정렬 오류이다. 상대 운동 핸드오프 조건에서 모든 ACT 획득은 자유 공간에서의 방향 전환 및 이동 과정을 거치는 동안 소포를 유지하지만, 전체 작업을 완료하는 비율은 64%에 불과하며, 이는 전문가 획득 후의 95%와 비교된다. 평가된 모든 정책과 속도에 걸쳐, 힘 폐쇄가 없는 414건의 획득 중 어느 것도 작업을 완료하지 못했다. 따라서 동일한 공칭 작업 성공률이 실행 속도 전반에 걸친 전문가 성능의 보존을 의미하지는 않는다. 코드, 데이터, 평가 스크립트는 https://github.com/coenwerem/parcelstow에서 제공된다.
English
Dexterous manipulation policies learned by imitation are typically evaluated for robustness to variation in scenes, objects, or instructions, but their performance across task execution speeds is less often examined. This leaves open how much temporal robustness a learner retains relative to the expert it imitates. We compare an expert and learner under the same task conditions, initial-condition draws, and speedup factors. We instantiate the evaluation in ParcelStow, a contact-rich task in which the robot acquires, reorients, and inserts a parcel. The demonstrations span the speedup range for the manipulation phases after parcel acquisition. A scripted expert and an Action Chunking with Transformers (ACT) policy trained from the expert's demonstrations both achieve 100 percent task success at nominal speed. Their success rates diverge within the demonstrated range: at its maximum, expert success is 84 percent and ACT success is 53 percent. Two ACT policies with different parameter initializations show similar degradation, decreasing by 34 and 48 percentage points from nominal speed to the maximum demonstrated speed, compared with 16 points for the expert. Stage-level analysis shows that 35 of ACT's 47 failures at the maximum demonstrated speed are insertion misalignments. Under the relative-motion handoff, every ACT acquisition retains the parcel through reorientation and transfer in free space, but only 64 percent complete the overall task, compared with 95 percent after expert acquisition. Across all evaluated policies and speeds, none of the 414 acquisitions without force closure completes the task. Equal nominal task success therefore does not imply preservation of expert performance across execution speeds. Code, data, and evaluation scripts are available at https://github.com/coenwerem/parcelstow.