ChatPaper.aiChatPaper

「重みか、それともスキルか?ロボット学習技術のサーベイ:行動予測の重みから、自らスキルを書くロボットまで」

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

August 3, 2026
著者: Gaytri Jena, Kapil Wanaskar, Vinija Jain, Aman Chadha, Vasu Sharma, Amitava Das
cs.AI

要旨

ロボット学習は、凍結された重みに能力を焼き込むポリシー(視覚・言語・行動モデル、すなわちVLAモデル)と、自身の実行可能なスキルをコードとして書き、洗練させるエージェントという二つの賭けに分かれつつある。本サーベイは、この「重み対スキル」という軸に沿って当該分野を整理するものである。その中心的な分析上の貢献は、コード・アズ・ポリシー手法を自己改善の度合いに応じて配置した深掘り調査であり、ゼロショットプログラム合成から、閉ループ自己修復、永続的スキルメモリ、そして実行フィードバック・スキルメモリ・進化的探索が単一のオープンエンドなループに組み合わさる、まだ疎らにしか埋められていないセルにまで及ぶ。このセルを占めるのは、ごく最近の少数のシステム(例えばASPIRE、ENPIRE、RoboClaw)のみである。我々は、教師なし強化学習によるスキル発見から大規模言語モデルによるスキルライブラリに至る、相補的な「スキル」側の極もマッピングし、「スキル」という語が少なくとも五つの異なる意味で使用されており、そのうち勾配更新なしで自己改善するのはコードの意味だけであることを示す。続いて、このタクソノミーを新興のスキル経済へと結び付ける。商用ロボットスキルマーケットプレイスは現在、ワンタップスキルをロボット間で配布しているが、それは静的な再生のみを出荷しており、適応、異なる身体間の移植性、来歴、安全性検証、合成、標準化といった未解決の問題を浮き彫りにしている。これは意図的に焦点を絞ったサーベイである。分野を網羅的にカタログ化するのではなく、一つのタクソノミーと一連の比較表を通じて、六つの手法ファミリーにわたる77の代表的なシステムを検討し、自己改善メカニズムの操作的定義を提供するとともに、各ファミリーが何をできないかを明記するものである。
English
Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights versus skills. Its central analytical contribution is a deep-dive that arranges code-as-policy methods by their degree of self-improvement, from zero-shot program synthesis, through closed-loop self-repair and persistent skill memory, to the sparsely populated cell in which execution feedback, skill memory, and evolutionary search combine into one open-ended loop; only a few very recent systems (for example ASPIRE, ENPIRE, and RoboClaw) occupy that cell. We map the complementary "skills" pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and show that the word "skill" is used in at least five distinct senses, of which only the code sense self-improves without gradient updates. We then connect the taxonomy to the emerging skill economy: commercial robot-skill marketplaces now distribute one-tap skills across robots but ship only static playback, which surfaces open problems of adaptation, cross-embodiment portability, provenance, safety verification, composition, and standardisation. This is a deliberately focused survey. Rather than cataloguing the field exhaustively, it examines 77 representative systems across six technique families through one taxonomy and a set of contrast tables, and it supplies operational definitions of the self-improvement mechanisms together with a statement of what each family cannot do.