权重还是技能?机器人学习技术综述:从动作预测权重到自主编写技能的机器人
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
August 3, 2026
作者: Gaytri Jena, Kapil Wanaskar, Vinija Jain, Aman Chadha, Vasu Sharma, Amitava Das
cs.AI
摘要
机器人学习正在分化为两大方向:一类是将能力固化进冻结权重的策略(视觉-语言-动作模型,即VLA模型),另一类是以代码形式编写并不断优化自身可执行技能的智能体。本综述围绕"权重vs技能"这一主轴来组织该领域。其核心分析贡献在于深入梳理了代码即策略方法按自我改进程度的分布——从零样本程序合成,经闭环自修复和持久技能记忆,直至执行反馈、技能记忆与进化搜索融合为一个开放式循环的象限,而该象限目前研究者寥寥;仅有少数非常新的系统(如ASPIRE、ENPIRE和RoboClaw)占据此位置。我们同时梳理了互补的"技能"一极,涵盖从无监督强化学习技能发现到大语言模型技能库的各类方法,并表明"技能"一词至少存在五种不同的含义,其中只有代码意义上的技能能在无需梯度更新的情况下实现自我改进。随后,我们将该分类体系与新兴的技能经济联系起来:商用机器人技能市场目前可跨机器人分发一键技能,但仅提供静态回放,这暴露了适应性、跨具身可移植性、溯源、安全验证、组合与标准化等方面的开放性问题。这是一篇刻意聚焦的综述。它并非穷尽式地罗列该领域,而是通过一个分类体系和一组对比表,考察了六大技术类别中的77个代表性系统,并提供了自我改进机制的操作性定义,同时阐明了各类别的能力边界。
English
Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights versus skills. Its central analytical contribution is a deep-dive that arranges code-as-policy methods by their degree of self-improvement, from zero-shot program synthesis, through closed-loop self-repair and persistent skill memory, to the sparsely populated cell in which execution feedback, skill memory, and evolutionary search combine into one open-ended loop; only a few very recent systems (for example ASPIRE, ENPIRE, and RoboClaw) occupy that cell. We map the complementary "skills" pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and show that the word "skill" is used in at least five distinct senses, of which only the code sense self-improves without gradient updates. We then connect the taxonomy to the emerging skill economy: commercial robot-skill marketplaces now distribute one-tap skills across robots but ship only static playback, which surfaces open problems of adaptation, cross-embodiment portability, provenance, safety verification, composition, and standardisation. This is a deliberately focused survey. Rather than cataloguing the field exhaustively, it examines 77 representative systems across six technique families through one taxonomy and a set of contrast tables, and it supplies operational definitions of the self-improvement mechanisms together with a statement of what each family cannot do.