ロボット学習のための進歩報酬モデリング:包括的サーベイ
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey
July 22, 2026
著者: Jianshu Zhang, Keliang Wu, Haoran Lu, Anbang Liu, Ce Zhang, Weijie Yin, Chengxuan Qian, Xiyuan Yang, Zhenyu Pan, Guo Ye, Han Liu
cs.AI
要旨
ロボット学習は、大規模な行動空間を持つ動的環境で行われる。終端成功信号は、タスクが完了したかどうかのみをロボットに伝える。現在の行動が進展しているのか、変化がないのか、あるいは以前の進展を台無しにしているのかについては説明しない。このため、近年の研究では、タスク実行中にフィードバックを提供する進捗報酬(progress rewards)がますます探求されている。しかし、現在の文献には共通の枠組みが欠如している。既存の手法は、異なる観測、目標仕様、出力信号、教師信号のソース、評価プロトコルを使用している。そのため、それらを比較し、結果が実際に何を検証しているのかを理解することが困難になっている。本サーベイでは、ロボット学習のための進捗報酬モデリングに関する統一的な視点を提供する。この分野を三つの相互に関連するステップで整理する。まず、進捗モデルのインターフェースを検討する。これは、モデルがどのような情報を受け取り、どのような形式の進捗信号を生成するのかを問うことで、外部から問題を定義する。次に、モデルの内部に進み、この信号を構築するために用いられる手法を研究する。これにより、進捗推定と報酬生成の背後にある異なる仮定やメカニズムが明らかになる。最後に、これらの手法を支えるデータとベンチマークを精査する。これにより、進捗の教師信号がどのように得られ、異なる評価が実際に何を測定しているのかが示される。これらの三つの視点を合わせることで、進捗モデルとは何か、どのように構築されるか、その品質がどのように検証されるかが結びつく。さらに、現在のアプローチの主な限界をまとめ、将来の研究の方向性について議論する。
English
Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current behavior is making progress, remaining unchanged, or undoing earlier progress. For this reason, recent studies have increasingly explored progress rewards that provide feedback during task execution. However, the current literature lacks a shared framework. Existing methods use different observations, goal specifications, output signals, supervision sources, and evaluation protocols. This makes it difficult to compare them and understand what their results actually validate. In this survey, we provide a unified view of progress reward modeling for robotic learning. We organize the field in three connected steps. We first study the interface of a progress model. This defines the problem from the outside by asking what information the model receives and what form of progress signal it produces. We then move inside the model and study the methods used to construct this signal. This reveals the different assumptions and mechanisms behind progress estimation and reward generation. Finally, we examine the data and benchmarks that support these methods. This shows how progress supervision is obtained and what different evaluations actually measure. Together, these three perspectives connect what a progress model is, how it is built, and how its quality is validated. We further summarize the main limitations of current approaches and discuss future research directions.