ChatPaper.aiChatPaper

가중치인가, 기술인가? 로봇 학습 기법에 대한 조사: 행동 예측 가중치부터 자신의 기술을 스스로 작성하는 로봇까지

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

August 3, 2026
저자: Gaytri Jena, Kapil Wanaskar, Vinija Jain, Aman Chadha, Vasu Sharma, Amitava Das
cs.AI

초록

로봇 학습은 두 가지 방향으로 분기하고 있다: 고정된 가중치에 능력을 구워 넣는 정책(비전-언어-행동, 즉 VLA 모델)과, 에이전트가 자신의 실행 가능한 스킬을 코드로 직접 작성하고 개선해 나가는 접근법이 그것이다. 본 서베이는 가중치 대(對) 스킬이라는 축을 기준으로 해당 분야를 체계화한다. 이 논문의 핵심적 분석 기여는 코드-기반 정책 방법들을 자기 개선의 정도에 따라 배열한 심층 분석으로, 제로샷 프로그램 합성에서부터 폐루프 자가 수리, 지속적 스킬 메모리를 거쳐, 실행 피드백과 스킬 메모리, 진화 탐색이 하나의 개방형 루프로 결합되는 매우 드문 영역에 이르기까지를 다룬다. 이 마지막 영역을 차지하는 시스템은 ASPIRE, ENPIRE, RoboClaw 등 극히 최근의 소수에 불과하다. 또한 본 논문은 비지도 강화학습 기반 스킬 발견에서 대규모 언어 모델 스킬 라이브러리에 이르는 상보적 '스킬' 축을 매핑하고, '스킬'이라는 용어가 최소한 다섯 가지 서로 다른 의미로 사용되며, 그중 코드 의미에서만 경사 업데이트 없이 자기 개선이 가능함을 보여준다. 이후 본 논문은 이 분류체계를 부상하는 스킬 경제와 연결한다: 상업용 로봇 스킬 마켓플레이스는 이제 다양한 로봇에 원탭 스킬을 배포하지만 정적 재생만을 제공할 뿐이며, 이는 적응, 교차 구현 이식성, 출처 추적, 안전성 검증, 구성, 표준화라는 미해결 과제를 표면화한다. 이는 의도적으로 초점을 좁힌 서베이다. 본 논문은 분야를 망라하여 목록화하는 대신, 여섯 가지 기법 계열에 걸친 77개의 대표 시스템을 하나의 분류체계와 대조 표들을 통해 검토하고, 자기 개선 메커니즘에 대한 작동적 정의와 각 계열이 수행할 수 없는 기능에 대한 진술을 함께 제시한다.
English
Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights versus skills. Its central analytical contribution is a deep-dive that arranges code-as-policy methods by their degree of self-improvement, from zero-shot program synthesis, through closed-loop self-repair and persistent skill memory, to the sparsely populated cell in which execution feedback, skill memory, and evolutionary search combine into one open-ended loop; only a few very recent systems (for example ASPIRE, ENPIRE, and RoboClaw) occupy that cell. We map the complementary "skills" pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and show that the word "skill" is used in at least five distinct senses, of which only the code sense self-improves without gradient updates. We then connect the taxonomy to the emerging skill economy: commercial robot-skill marketplaces now distribute one-tap skills across robots but ship only static playback, which surfaces open problems of adaptation, cross-embodiment portability, provenance, safety verification, composition, and standardisation. This is a deliberately focused survey. Rather than cataloguing the field exhaustively, it examines 77 representative systems across six technique families through one taxonomy and a set of contrast tables, and it supplies operational definitions of the self-improvement mechanisms together with a statement of what each family cannot do.