World Editing: Intervening on Executable Worlds at Increasing Depth
Max Ku, Nok-Kan Law, Yu-Chien Tang, Shih-Ying Yeh, Ping Nie, Andy Zheng, Tat Hei Lai, Fei-Yueh Chen, Nikko Yu, Wei-Chieh Sun, Suzy Huang, Chiao-Wei Hsu, Chih-Chuan Huang, Chak-Wing Mak, Ho Yin Sam Ng, Edisy Kin Wai Chan, Min-Hung Chen, Ho Kei Cheng
arXiv:2610.02331 · 2026-10-06 공개 · arXiv · PDF
coding-agents visual-consistency minecraft executable-worlds world-editing game-modding igmworld igmbench
Abstract
Interactive world models are increasingly capable of generating environments and acting within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on an existing world while preserving properties that should remain unchanged, and introduce intervention depth as an axis describing how strongly an edit couples world entities, dynamics, and systems. We instantiate this capability through industry-grade game modding and introduce IGMWorld, together with IGMBench, a benchmark of 110 tasks and over 1.1K executable state and behavioral criteria across Minecraft and Terraria. The tasks span property, entity, dynamics, and system interventions and are evaluated through deterministic executability, behavioral, preservation, and visual checks. Frontier coding agents already exhibit substantial world-editing capability: the strongest configuration solves 78.2% of tasks under a strict task-level criterion, while criterion-level performance reaches 94.8%. Reliability generally decreases with intervention depth, and this pattern persists even among tasks with similar numbers of evaluation criteria. Most failed edits still build and load successfully, suggesting that the main difficulty is making the edited world behave as requested. Visual consistency remains a separate weakness, with all evaluated configurations below 50% joint visual pass rate. These results show that world editing is a distinct capability from world generation and interaction, and that executable games provide a practical testbed for studying it.
한국어 요약
한 줄 요약
실행 가능한 게임 세계에 대한 의도적 편집 능력을 평가하는 IGMWorld와 IGMBench를 제안하며, 편집 깊이가 성능에 미치는 영향을 분석한다.
핵심 기여도
- **World Editing**을 실행 가능한 세계에 대한 의도적 개입으로 정의하고, **Intervention Depth** 개념을 도입하여 편집의 복잡도를 정량화.
- **IGMWorld**와 **IGMBench**를 제안: Minecraft와 Terraria에서 110개의 편집 작업과 1.1K 이상의 실행 가능한 평가 기준 포함.
- 최고 성능 설정에서 **78.2%의 엄격한 작업 수준 성능**, **94.8%의 기준 수준 성능** 달성.
- **시각 일관성**(Joint Visual Pass Rate)은 모든 평가 설정에서 50% 미만으로 나타남.
핵심 아이디어
기존 연구는 주로 세계 생성(World Generation)과 상호작용(World Interaction)에 집중했으나, 본 연구는 **기존 실행 가능한 세계에 대한 의도적 편집**(World Editing)을 새로운 연구 영역으로 제시한다. 이는 단순히 시각적 수정을 넘어, **엔티티, 역학, 시스템 간의 결합**(Coupling)을 고려한 개입을 포함한다.
**Intervention Depth**는 편집이 세계 내 요소들을 얼마나 강하게 연결하는지를 기준으로 정의되며, 이는 편집의 복잡도와 신뢰도 사이의 관계를 분석하는 핵심 지표이다. 연구는 **게임 모드**(Game Modding)를 통해 이 개념을 구현하고, **실행 가능성**(Executability), **행동 정확도**(Behavioral Correctness), **시각 일관성**(Visual Consistency)을 평가하는 세 단계 평가 프레임워크를 제안한다.
기술적 접근법
- **IGMWorld**: 실행 가능한 게임 세계에 대한 편집 작업을 수행하고, **실행 가능성**, **행동 정확도**, **시각 일관성**을 평가하는 환경.
- **IGMBench**: Minecraft와 Terraria에서 110개의 편집 작업과 1.1K 이상의 실행 가능한 평가 기준을 포함.
- **세 단계 평가**:
1. **실행 가능성**: 수정된 세계가 빌드되고 로드되는지 확인.
2. **편집 정확도**: 요청된 개입이 실행 세계에서 올바르게 반영되었는지, **상태 검사**(State Checks), **행동 검사**(Behavioral Checks), **회귀 검사**(Regression Checks)로 평가.
3. **시각 일관성**: **TPIPS**(Text-Conditioned Perceptual Image Patch Similarity)를 사용하여 텍스트 조건에 따른 **의미 일관성**(Semantic Consistency)과 **스타일 일관성**(Style Consistency)을 평가.
주요 결과
- **78.2%**의 엄격한 작업 수준 성능, **94.8%**의 기준 수준 성능 달성.
- **Minecraft와 Terraria**에서 평가한 110개 작업 중, **실행 가능성**은 대부분 충족되나, **행동 정확도**가 낮은 경우가 많음.
- **시각 일관성**은 모든 평가 설정에서 **50% 미만**의 통과율을 기록.
- **Intervention Depth**가 높을수록 **신뢰도가 감소**하며, 이는 평가 기준 수와 무관하게 지속됨.
의의 및 한계
- **World Editing**은 기존의 세계 생성 및 상호작용과 구분되는 **독립적인 능력**으로, 실행 가능한 게임 세계는 이를 연구하기 위한 실용적인 테스트베드이다.
- **실행 가능성**은 상대적으로 높으나, **의도된 행동의 실현**과 **시각 일관성**은 여전히 주요한 도전 과제이다.
- **TPIPS**를 사용한 시각 평가가 효과적이지만, **주관적 해석**과 **다양한 자산 간 비교의 한계**가 존재.
실용적 활용
- **게임 개발**, **AI 기반 자동화 툴**, **교육용 시뮬레이션** 등에서 실행 가능한 세계의 **의도적 편집**이 필요한 상황에 적용 가능.
- **IGMWorld와 IGMBench**는 AI 에이전트의 **세계 이해 및 수정 능력**을 평가하는 표준화된 테스트베드로 활용될 수 있다.