RobotUse: Allocating Computation, Context, and Decisions

Junhoo Lee, Injun Baek, Seungyeon Kim, Suhyun Jeon, Minkyu Kim, Baekseung Kim, Nojun Kwak

arXiv:2610.04929 · 2026-10-06 공개 · arXiv · PDF

task-success motion-planning subagent feedback-learning context-retention physical-action decision-harnessing robot-use

Abstract

Robot agents must connect their intended actions to observed outcomes while retaining the context needed to revise their choices over repeated attempts. Existing interfaces often leave these choices inside predefined tools or require agents to manage detailed execution code and its growing history. We introduce RobotUse, a robot agent harness that organizes computation, context, and decisions around specifying and revising physical actions. Agents visually select targets and poses, while the backend handles geometry, motion planning, and control. Subagents retain detailed interactions within each subgoal and return the information needed for subsequent decisions. Continual harnessing lets agents learn from execution by updating a persistent playbook. On RoboLab, RobotUse achieves 45% task success, outperforming CaP-X by 6.7 percentage points while maintaining compact decision contexts and reducing reliance on predefined action abstractions. Furthermore, we show that RobotUse learns from real-world execution despite imperfect feedback and transfers what it learns to subsequent tasks. Project page is available at https://robotuse-team.github.io/.

한국어 요약

한 줄 요약

RobotUse는 물리적 행동을 지정하고 수정하는 데 초점을 맞춘 로봇 에이전트 허네스를 제시하며, RoboLab에서 45%의 작업 성공률을 달성한다.

핵심 기여도

핵심 아이디어

RobotUse는 로봇 에이전트가 물리적 행동을 지정하고 수정할 수 있도록 **intent-based** 허네스를 설계한 것이 핵심이다. 주 에이전트는 언어로 작업 의도를 유지하고, **subagent**가 시각적 관찰을 기반으로 작업 관련 대상, 위치, 방향을 선택하고 정제한다. 이는 **geometry, motion planning, control**을 포함한 백엔드가 실행 가능한 동작으로 변환한다.

**Playbook**은 작업 분해와 위임을 안내하며, 어떤 조건에서 무엇이 효과적이었는지, 무엇을 다르게 시도해야 하는지를 경험적으로 기록한다. 이는 **continual harnessing**을 통해 백엔드는 고정된 상태에서 경험을 업데이트하고, 후속 결정을 개선한다. 이 접근법은 **predefined action abstraction에 대한 의존도를 줄이며**, 로봇의 내부 결정과 물리적 결과 간의 연결을 명확히 유지한다.

기술적 접근법

주요 결과

의의 및 한계

RobotUse는 로봇 에이전트의 **geometric computation과 control을 위임**하면서도, **물리적 선택을 수정할 수 있는 context를 유지**하는 새로운 허네스 설계를 제시한다. 이는 **long-horizon, multi-step 작업**을 지원하며, **실행 경험을 기반으로 지속적으로 개선**할 수 있는 구조를 제공한다.

그러나, **imperfect feedback** 상황에서도 학습이 가능하다는 점은 한계이자 강점이다. 즉, **정확한 피드백 없이도 학습이 가능하지만**, 이는 **정확도 향상에 한계가 있을 수 있다**. 또한, **playbook의 초기 구성이 성능에 큰 영향을 미칠 수 있다**는 점도 한계로 지적된다.

실용적 활용

RobotUse는 **로봇 조작 시스템**, 특히 **복잡한 물리적 작업을 반복적으로 수행해야 하는 산업 현장**에 적용 가능하다. 예를 들어, **로봇 팔이 물체를 정확히 조작하거나 장애물을 회피하는 작업**에서 유용하며, **자율 주행 차량의 경로 계획 및 실행**에도 활용될 수 있다.