vision-language-models robot-control agent-harness recursive-self-improvement proprioception state-representation robodojo bimanual-tasks
Abstract
Most robot policies keep a model in the control loop: a VLA maps observations to actions, and an Agent Harness, such as Agent-as-Policy or Harness VLA queries a VLM for decision making at run time. We propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy. If this state can be represented accurately, the decision making can be written entirely in code. We therefore propose Code-Only-as-Policy (COAP): code measures and tracks the robot, environment, and task state from camera images and proprioception, and makes every decision from it. The same code applies across episodes, and different tasks share one library without a VLM or VLA in the loop. Compared with VLAs and Agent Harnesses, we analyze three advantages of COAP: (i) Explicit State: the state can be stored in code; (ii) Execution: code makes decision making controllable, recovers from failures flexibly, and runs fast and cheaply online; (iii) Extensibility: new tasks reuse, inherit, or extend the shared library, so capabilities can accumulate over tasks. These advantages make COAP a suitable medium for recursive self-improvement (RSI): coding agents develop the library in a closed loop, and each change is explicit and controllable. On RoboDojo's 42 bimanual tasks, the resulting library reaches a success rate of 70.24% without a model at test time. The upper bound of COAP lies in how accurately the state is represented for decision making and how robust the code logic is. We thus propose COAP as a new paradigm for embodied tasks; since it applies across episodes, it can also serve as an efficient data engine for VLAs and Agent Harnesses.
한국어 요약
한 줄 요약
COAP는 로봇 정책을 코드로만 구현하여 재귀적 자기 개선(RSI)을 가능하게 하는 새로운 패러다임이다.
핵심 기여도
- COAP는 로봇과 환경 상태를 코드 변수에 저장하고, 정책을 코드 로직으로 구현하여 VLM/VLA 없이 실행 가능.
- RoboDojo의 42개 이족 태스크에서 70.24% 성공률 달성 (기존 최고 기법 대비 +38.9%).
- 코드 정책은 확장성과 실행 효율성, 실패 복구 능력이 뛰어나 RSI에 적합.
- COAP는 VLA 및 Agent Harness에 대한 데이터 엔진으로도 활용 가능.
핵심 아이디어
COAP는 로봇과 환경을 "Embodied Turing Machine"으로 모델링한다. 이 모델에서 로봇의 정책은 코드로 구현되며, 상태는 코드 변수에 저장된다. 기존 VLA나 Agent Harness는 정책을 학습 모델에 의존하지만, COAP는 코드로 상태를 추적하고, 정확한 상태 정보를 기반으로 명시적인 결정을 내린다. 이는 코드가 투명하고 제어 가능하며, 재사용 가능한 라이브러리로 확장될 수 있음을 의미한다. COAP는 각 태스크에 맞는 코드 로직이 17% 미만으로 제한되며, 83%는 공유 라이브러리에서 사용된다. 이는 정책의 재사용성과 확장성을 높인다.
기술적 접근법
- **COAP 정책**: 로봇의 관찰(카메라 이미지, 프로피오셉션)을 기반으로 상태를 추적하고, 코드로 정책을 실행.
- **상태 표현**: 객체의 자세와 공간 관계, 로봇 팔의 자세 등을 코드 변수에 저장.
- **실행 방식**: 코드는 정확한 상태 정보를 기반으로 결정을 내리며, 실패 시 백트래킹과 재시도 가능.
- **확장성**: 새로운 태스크는 공유 라이브러리의 코드를 재사용, 상속, 확장하여 구현.
- **RSI 구조**: 코드 작성 에이전트(Opus 5.5 Max)가 라이브러리를 반복적으로 개선하며, 각 변경 사항은 명시적이고 제어 가능.
주요 결과
- RoboDojo의 42개 이족 태스크에서 COAP 정책은 70.24% 성공률 달성.
- 기존 최고 기법 대비 +38.9% 개선.
- 태스크별 코드 로직은 전체 코드의 17% 미만, 83%는 공유 라이브러리에서 사용.
- 코드 정책은 VLM/VLA 없이 실행되며, 실행 속도와 비용 효율성이 뛰어남.
의의 및 한계
COAP는 로봇 정책을 코드로 명시적으로 표현함으로써 정책 개선 과정을 투명하고 제어 가능하게 만든다. 이는 재귀적 자기 개선(RSI)에 이상적인 구조이며, 코드의 확장성과 재사용성도 높인다. 또한, COAP는 VLA 및 Agent Harness에 대한 데이터 생성 엔진으로 활용 가능하다. 그러나 COAP의 한계는 상태 표현의 정확성과 코드 로직의 안정성에 있다. 일부 태스크에서는 여전히 VLM/VLA가 필요할 수 있으며, 일반화 능력이 제한적일 수 있다.
실용적 활용
COAP는 로봇이 반복적인 태스크를 수행하는 산업 환경에서 유용하며, 코드 기반 정책 개선이 필요한 연구 분야에도 적용 가능하다. 특히, 코드 라이브러리를 기반으로 한 RSI는 로봇 학습의 효율성을 높일 수 있다.