로봇 학습을 위한 진행 보상 모델링: 종합 서베이
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey
July 22, 2026
저자: Jianshu Zhang, Keliang Wu, Haoran Lu, Anbang Liu, Ce Zhang, Weijie Yin, Chengxuan Qian, Xiyuan Yang, Zhenyu Pan, Guo Ye, Han Liu
cs.AI
초록
로봇 학습은 넓은 행동 공간을 가진 동적 환경에서 이루어집니다. 종료 성공 신호는 로봇에게 과제 완료 여부만을 알려줄 뿐, 현재 행동이 진행 중인지, 변화가 없는지, 아니면 이전 진행 상황을 되돌리고 있는지 설명하지 않습니다. 이러한 이유로 최근 연구들은 작업 실행 중 피드백을 제공하는 진행 보상을 점차 탐구하고 있습니다. 그러나 현재 문헌에는 공유된 프레임워크가 부재합니다. 기존 방법들은 서로 다른 관찰, 목표 명세, 출력 신호, 감독 소스, 평가 프로토콜을 사용합니다. 이로 인해 방법들을 비교하고 그 결과가 실제로 무엇을 검증하는지 이해하기 어렵습니다. 본 서베이에서는 로봇 학습을 위한 진행 보상 모델링에 대한 통합된 관점을 제시합니다. 우리는 이 분야를 세 가지 연결된 단계로 구성합니다. 먼저 진행 모델의 인터페이스를 연구합니다. 이는 모델이 어떤 정보를 수신하고 어떤 형태의 진행 신호를 생성하는지 질문함으로써 문제를 외부에서 정의합니다. 그런 다음 모델 내부로 이동하여 이 신호를 구성하는 데 사용되는 방법들을 연구합니다. 이를 통해 진행 추정 및 보상 생성 이면의 다양한 가정과 메커니즘이 드러납니다. 마지막으로 이러한 방법들을 뒷받침하는 데이터와 벤치마크를 검토합니다. 이를 통해 진행 감독이 어떻게 획득되는지와 서로 다른 평가가 실제로 무엇을 측정하는지 확인할 수 있습니다. 함께, 이 세 가지 관점은 진행 모델이 무엇인지, 어떻게 구축되는지, 그리고 그 품질이 어떻게 검증되는지를 연결합니다. 또한 현재 접근 방식의 주요 한계를 요약하고 향후 연구 방향을 논의합니다.
English
Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current behavior is making progress, remaining unchanged, or undoing earlier progress. For this reason, recent studies have increasingly explored progress rewards that provide feedback during task execution. However, the current literature lacks a shared framework. Existing methods use different observations, goal specifications, output signals, supervision sources, and evaluation protocols. This makes it difficult to compare them and understand what their results actually validate. In this survey, we provide a unified view of progress reward modeling for robotic learning. We organize the field in three connected steps. We first study the interface of a progress model. This defines the problem from the outside by asking what information the model receives and what form of progress signal it produces. We then move inside the model and study the methods used to construct this signal. This reveals the different assumptions and mechanisms behind progress estimation and reward generation. Finally, we examine the data and benchmarks that support these methods. This shows how progress supervision is obtained and what different evaluations actually measure. Together, these three perspectives connect what a progress model is, how it is built, and how its quality is validated. We further summarize the main limitations of current approaches and discuss future research directions.