ChatPaper.aiChatPaper

權重還是技能?機器人學習技術綜述:從預測動作的權重到能自行編寫技能的機器人

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

August 3, 2026
作者: Gaytri Jena, Kapil Wanaskar, Vinija Jain, Aman Chadha, Vasu Sharma, Amitava Das
cs.AI

摘要

機器人學習正分裂為兩條路線:一是將能力固化於凍結權重中的策略(視覺-語言-行動模型,即 VLA 模型),二是以程式碼形式撰寫並完善自身可執行技能的代理。本綜述圍繞「權重 vs. 技能」這條軸線來組織該領域。其核心分析貢獻在於一項深入探討,依照自我改進程度排列「程式碼即策略」方法:從零樣本程式合成,經由閉迴路自我修復與持久技能記憶,再到一個鮮少有人佔據的格位——在其中,執行回饋、技能記憶與演化搜尋結合成為一個開放式迴圈;目前僅有少數非常新近的系統(如 ASPIRE、ENPIRE 及 RoboClaw)位居此格。我們亦描繪了互補的「技能」一極,從無監督強化學習的技能發現,到大語言模型的技能函式庫,並指出「技能」一詞至少承載五種不同涵義,其中唯有「程式碼」的涵義能在不經梯度更新的情況下自我改進。接著,我們將這套分類法連結至正在興起的技能經濟:商業機器人技能市集如今能跨機器人分發一鍵技能,但所出貨的僅是靜態播放,因而凸顯出適應、跨本體可攜性、來源溯源、安全驗證、組合與標準化等開放問題。這是一份刻意聚焦的綜述。它並未窮舉式地編列整個領域,而是透過一套分類法與一組對照表,檢視六個技術家族中的 77 個代表性系統,並為各種自我改進機制提供操作性定義,同時說明每個家族所不能做到的事。
English
Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights versus skills. Its central analytical contribution is a deep-dive that arranges code-as-policy methods by their degree of self-improvement, from zero-shot program synthesis, through closed-loop self-repair and persistent skill memory, to the sparsely populated cell in which execution feedback, skill memory, and evolutionary search combine into one open-ended loop; only a few very recent systems (for example ASPIRE, ENPIRE, and RoboClaw) occupy that cell. We map the complementary "skills" pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and show that the word "skill" is used in at least five distinct senses, of which only the code sense self-improves without gradient updates. We then connect the taxonomy to the emerging skill economy: commercial robot-skill marketplaces now distribute one-tap skills across robots but ship only static playback, which surfaces open problems of adaptation, cross-embodiment portability, provenance, safety verification, composition, and standardisation. This is a deliberately focused survey. Rather than cataloguing the field exhaustively, it examines 77 representative systems across six technique families through one taxonomy and a set of contrast tables, and it supplies operational definitions of the self-improvement mechanisms together with a statement of what each family cannot do.