Quo Vadis, World Modeling?

arXiv:2608.02713 · 2026-08-05 공개 · arXiv · PDF

co-evolution world-modeling training-optimization execution-outcomes skill-proxies reward-proxies in-context-guidance dynamic-feedback

Abstract

Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to parallelize. World modeling offers a natural intermediate proxy that allows agents to query lower-cost, more controllable feedback before committing to real actions. Classical world models instantiate this proxy primarily through future physical-state prediction, a formulation useful yet narrow for agents that require actionable feedback beyond raw state transitions. In this work, we conceptualize Agent-Centric Interactive World Proxies, shifting the fundamental paradigm from physical state transitions to agent-usable information transitions, such as execution outcomes, retrieved experiences or skills, and verification signals, broadening the scope of world modeling to provide versatile feedback for continually improving agents. To systematically map this design space, we organize world proxies into six functional forms based on their feedback modalities: dynamics, spatial, execution, memory/experience, skill, and reward/verification proxies, which together characterize the primary ways world modeling serves agent improvement. We further analyze how these proxies empower agents across three progressive levels: L.1 Inference-Time Guidance, where proxy outputs enrich in-context information for superior decisions; L.2 Training-Time Optimization, where proxy outputs yield rewards, critiques, or synthetic rollouts for policy learning; and L.3 Agent-Proxy Co-Evolution, where real-environment evidence continuously updates both the proxy and the agent for co-evolution. Ultimately, this work recasts world modeling into an agent-centric paradigm, establishing a roadmap for building world proxies that empower agents to plan better, learn faster, and evolve continually.

한국어 요약

한 줄 요약

세계 모델링을 에이전트 중심의 정보 전환 프록시로 확장하여 지속적 개선을 지원한다.

핵심 기여도

핵심 아이디어

기존 세계 모델링은 주로 미래 물리 상태 예측에 초점을 맞추어 왔으나, 이는 에이전트가 행동 결정에 필요한 정보를 제공하는 데 한계가 있다. 본 연구는 에이전트가 사용 가능한 정보 전환(예: 실행 결과, 경험, 검증)을 기반으로 세계 프록시를 구성함으로써, 세계 모델링의 범위를 확장한다. 이는 에이전트가 더 유용한 피드백을 받아 지속적으로 개선할 수 있도록 한다. 특히, 정보 전환에 기반한 프록시는 에이전트의 실행 시 판단, 학습 시 최적화, 그리고 프록시와의 공진화를 가능하게 한다는 점에서 혁신적이다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 세계 모델링의 핵심 개념을 에이전트 중심의 정보 전환으로 재정의함으로써, 에이전트가 더 유연하고 지속적으로 학습할 수 있는 기반을 제공한다. 특히, 실행 결과, 경험, 검증 신호 등을 활용한 프록시는 실제 환경과의 상호작용을 줄이면서도 성능을 유지할 수 있는 가능성을 제시한다. 그러나 구체적인 실험 수치나 비교 대상이 명시되지 않아, 실제 성능 개선 여부는 추가 연구가 필요하다.

실용적 활용

본 프레임워크는 로봇, 자율 주행, 게임 AI 등 에이전트가 실제 환경과 상호작용을 최소화하면서도 지속적으로 학습해야 하는 분야에 적용 가능하다. 특히, 비용이 높거나 위험한 환경에서의 에이전트 훈련에 유용할 수 있다.