ChatPaper.aiChatPaper

RynnValue: 時間的距離によるロボティクス価値基盤モデルのスケーリング

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

August 10, 2026
著者: Dongchi Huang, Hongyin Zhang, Bohan Hou, Siteng Huang, Zhian Su, Hang Guo, Tong Lu, Zhaofeng Xu, Jiahao Tang, Jianfei Yang, Donglin Wang, Peixi Peng, Mingxiu Chen, Deli Zhao, Xin Li
cs.AI

要旨

汎用報酬モデルはロボット学習のスケーリングにおけるボトルネックとしてますます重要になっているが、大規模な異種コーパスから価値関連能力を学習する方法は依然として十分に研究されていない。既存手法は教師信号を選好や正規化進捗といったタスク内部のアンカーに結び付けており、それらのいずれもエンボディメントやデータソースをまたいでうまく転移しない。我々は、このアンカーを時間的距離(観測から言語指定ゴールまでの有向コスト・トゥ・ゴー)に置き換える、ロボット操作のためのオープンソース価値基盤モデル RynnValue を導入する。時間的距離ラベルはタイムスタンプから直接導出できるため、RynnValue は選好や進捗のアノテーションなしで、7,000時間以上、約300万本の指示条件付きクリップにスケールする。時間的価値学習を大規模に適用しても信頼性があるものにするため、ランダムな時間的サンプリング、時間順序シャッフリング、価値分離アテンションを組み合わせ、予測を失敗や後退に鈍感にしてしまうショートカットを抑制する。選好ラベルなしで学習した RynnValue は、RBM-EVAL-OOD において平均 Kendall's tau_a 0.675 を達成し、完全に選好ラベルで教師された最先端手法(0.655)を上回り、進捗のみの比較モデル(0.292)の2倍以上を達成する。また、未見のタスク、エンボディメント、視点に対してゼロショット汎化する。ポテンシャルベースの報酬整形によって密報酬へ変換すると、実世界での方策成功率はオンラインで52.5%から72.5%へ、オフラインで63.8%から82.5%へ向上する。これらの結果は、時間的距離がスケーラブルな教師信号として、また汎用ロボット方策に対する実用的な報酬インターフェースとして機能することを確立する。
English
General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal anchors such as preferences or normalized progress, none of which transfer cleanly across embodiments and data sources. We introduce RynnValue, an open-source value foundation model for robotic manipulation that replaces these anchors with temporal distance, the directed cost-to-go from an observation to the language-specified goal. Because temporal-distance labels can be derived directly from timestamps, RynnValue scales to over 7,000 hours and roughly 3M instruction-conditioned clips without preference or progress annotations. To make temporal-value learning reliable at scale, we combine random temporal sampling, temporal-order shuffling, and value-isolation attention, suppressing shortcuts that would leave predictions insensitive to failures and regressions. Trained without preference labels, RynnValue attains an average Kendall's tau_a of 0.675 on RBM-EVAL-OOD, surpassing the fully preference-supervised state of the art (0.655) and more than doubling a progress-only counterpart (0.292), while generalizing zero-shot to unseen tasks, embodiments, and viewpoints. Converted into dense rewards via potential-based shaping, it raises real-world policy success from 52.5% to 72.5% online and from 63.8% to 82.5% offline. These results establish temporal distance as a scalable supervision target and practical reward interface for generalist robot policies.