co-evolution world-modeling training-optimization execution-outcomes skill-proxies reward-proxies in-context-guidance dynamic-feedback
Abstract
Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to parallelize. World modeling offers a natural intermediate proxy that allows agents to query lower-cost, more controllable feedback before committing to real actions. Classical world models instantiate this proxy primarily through future physical-state prediction, a formulation useful yet narrow for agents that require actionable feedback beyond raw state transitions. In this work, we conceptualize Agent-Centric Interactive World Proxies, shifting the fundamental paradigm from physical state transitions to agent-usable information transitions, such as execution outcomes, retrieved experiences or skills, and verification signals, broadening the scope of world modeling to provide versatile feedback for continually improving agents. To systematically map this design space, we organize world proxies into six functional forms based on their feedback modalities: dynamics, spatial, execution, memory/experience, skill, and reward/verification proxies, which together characterize the primary ways world modeling serves agent improvement. We further analyze how these proxies empower agents across three progressive levels: L.1 Inference-Time Guidance, where proxy outputs enrich in-context information for superior decisions; L.2 Training-Time Optimization, where proxy outputs yield rewards, critiques, or synthetic rollouts for policy learning; and L.3 Agent-Proxy Co-Evolution, where real-environment evidence continuously updates both the proxy and the agent for co-evolution. Ultimately, this work recasts world modeling into an agent-centric paradigm, establishing a roadmap for building world proxies that empower agents to plan better, learn faster, and evolve continually.
한국어 요약
한 줄 요약
세계 모델링을 에이전트 중심의 정보 전환 프록시로 확장하여 지속적 개선을 지원한다.
핵심 기여도
- 에이전트 중심의 Interactive World Proxies 개념 도입.
- 6가지 기능적 형태(동역학, 공간, 실행, 메모리/경험, 스킬, 보상/검증)로 세계 프록시 분류.
- 3단계 에이전트-프록시 상호작용 수준(L.1~L.3) 제시.
- 기존 물리 상태 예측 중심 모델에서 벗어나 실행 결과, 경험, 검증 신호 등으로 확장.
핵심 아이디어
기존 세계 모델링은 주로 미래 물리 상태 예측에 초점을 맞추어 왔으나, 이는 에이전트가 행동 결정에 필요한 정보를 제공하는 데 한계가 있다. 본 연구는 에이전트가 사용 가능한 정보 전환(예: 실행 결과, 경험, 검증)을 기반으로 세계 프록시를 구성함으로써, 세계 모델링의 범위를 확장한다. 이는 에이전트가 더 유용한 피드백을 받아 지속적으로 개선할 수 있도록 한다. 특히, 정보 전환에 기반한 프록시는 에이전트의 실행 시 판단, 학습 시 최적화, 그리고 프록시와의 공진화를 가능하게 한다는 점에서 혁신적이다.
기술적 접근법
- 6가지 기능적 형태의 프록시: dynamics, spatial, execution, memory/experience, skill, reward/verification.
- 3단계 상호작용 수준: Inference-Time Guidance, Training-Time Optimization, Agent-Proxy Co-Evolution.
- 에이전트가 실행 결과, 경험, 검증 신호 등을 기반으로 행동 결정 및 학습을 수행하도록 구조화.
- 기존의 future physical-state prediction 중심 접근에서 벗어난 정보 기반 프록시 설계.
주요 결과
- 기존 세계 모델링 접근에서 벗어난 새로운 프레임워크 제시.
- 6가지 프록시 형태와 3단계 상호작용 수준을 통해 에이전트 개선의 구조적 기반 제공.
- 명시적인 수치 결과는 제공되지 않으나, 프레임워크의 확장성과 유연성 강조.
의의 및 한계
본 연구는 세계 모델링의 핵심 개념을 에이전트 중심의 정보 전환으로 재정의함으로써, 에이전트가 더 유연하고 지속적으로 학습할 수 있는 기반을 제공한다. 특히, 실행 결과, 경험, 검증 신호 등을 활용한 프록시는 실제 환경과의 상호작용을 줄이면서도 성능을 유지할 수 있는 가능성을 제시한다. 그러나 구체적인 실험 수치나 비교 대상이 명시되지 않아, 실제 성능 개선 여부는 추가 연구가 필요하다.
실용적 활용
본 프레임워크는 로봇, 자율 주행, 게임 AI 등 에이전트가 실제 환경과 상호작용을 최소화하면서도 지속적으로 학습해야 하는 분야에 적용 가능하다. 특히, 비용이 높거나 위험한 환경에서의 에이전트 훈련에 유용할 수 있다.