reinforcement-learning robot-learning vla-models self-improvement skill-discovery code-as-policy enpire skill-economy
Abstract
Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights versus skills. Its central analytical contribution is a deep-dive that arranges code-as-policy methods by their degree of self-improvement, from zero-shot program synthesis, through closed-loop self-repair and persistent skill memory, to the sparsely populated cell in which execution feedback, skill memory, and evolutionary search combine into one open-ended loop; only a few very recent systems (for example ASPIRE, ENPIRE, and RoboClaw) occupy that cell. We map the complementary "skills" pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and show that the word "skill" is used in at least five distinct senses, of which only the code sense self-improves without gradient updates. We then connect the taxonomy to the emerging skill economy: commercial robot-skill marketplaces now distribute one-tap skills across robots but ship only static playback, which surfaces open problems of adaptation, cross-embodiment portability, provenance, safety verification, composition, and standardisation. This is a deliberately focused survey. Rather than cataloguing the field exhaustively, it examines 77 representative systems across six technique families through one taxonomy and a set of contrast tables, and it supplies operational definitions of the self-improvement mechanisms together with a statement of what each family cannot do.
한국어 요약
한 줄 요약
로봇 학습은 고정된 가중치와 자가 개선 가능한 스킬이라는 두 축으로 나뉘고 있다.
핵심 기여도
- 77개 대표 시스템을 6가지 기술 패밀리로 분류하고, 대조 테이블을 통해 체계적으로 정리.
- 코드 기반 정책 방법을 자가 개선 정도에 따라 3단계로 구분 (zero-shot, closed-loop, open-ended).
- "스킬"이라는 용어가 최소 5가지 의미로 사용되며, 그 중 코드 기반 스킬만 경사 하강 없이 자가 개선 가능.
- ASPIRE, ENPIRE, RoboClaw 등 최신 시스템이 open-ended 루프에 속함.
핵심 아이디어
로봇 학습 분야는 "가중치"와 "스킬"이라는 두 가지 접근 방식으로 나뉘고 있다. "가중치" 기반 접근은 시각-언어-작업(VLA) 모델을 통해 고정된 정책을 학습하는 반면, "스킬" 기반 접근은 로봇이 실행 가능한 코드 형태로 스킬을 생성하고 개선한다. 이 논문은 "자가 개선"이라는 개념을 중심으로 코드 기반 정책을 세 단계로 구분한다. 특히, 실행 피드백, 스킬 메모리, 진화적 탐색이 결합된 open-ended 루프는 현재 매우 적은 시스템만이 달성하고 있다. 이는 로봇이 스스로 학습하고 적응하는 능력을 갖추는 데 중요한 단계로, ASPIRE, ENPIRE, RoboClaw 등이 대표적이다.
기술적 접근법
- 코드 기반 정책 방법은 zero-shot 프로그램 합성, closed-loop 자가 수리, persistent skill memory, open-ended 루프로 구분.
- VLA 모델은 고정된 가중치를 사용해 시각-언어-작업을 통합.
- 스킬 탐색 기법에는 비지도 강화 학습 기반 스킬 발견과 대규모 언어 모델 기반 스킬 라이브러리 포함.
- 77개 시스템을 6가지 기술 패밀리로 분류하고 대조 테이블을 통해 분석.
주요 결과
- open-ended 루프는 ASPIRE, ENPIRE, RoboClaw 등 최신 시스템만이 달성.
- 현재 상용 로봇 스킬 마켓플레이스는 정적 재생만 제공해 적응, 이식성, 안전 검증 등 문제 제기.
- 코드 기반 스킬은 경사 하강 없이 자가 개선 가능하며, 다른 스킬 정의와 구분됨.
의의 및 한계
이 논문은 로봇 학습 분야를 체계적으로 분류하고, 자가 개선 메커니즘을 명확히 정의함으로써 학술적 기반을 제공한다. 특히, 코드 기반 스킬이 경사 하강 없이 자가 개선할 수 있다는 점은 기존 접근과 차별화된다. 그러나 open-ended 루프를 구현한 시스템은 여전히 제한적이며, 실용화를 위한 이식성, 표준화, 안전 검증 등의 문제는 여전히 미해결 상태이다. 또한, 모든 시스템을 포괄적으로 다루기보다는 대표적인 77개 시스템에 집중했기 때문에 일부 접근법이 누락될 수 있다.
실용적 활용
이 연구는 로봇 스킬 마켓플레이스, 자율 로봇 시스템, 산업 자동화 등에서 활용 가능하다. 특히, 코드 기반 스킬은 로봇이 실시간으로 환경에 적응하고 스스로 개선할 수 있는 기반을 제공하며, 이는 서비스 로봇, 제조, 물류 분야에서 실용적 가치가 크다.