ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu, Zheng Zhou, Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo
arXiv:2607.19191 · 2026-07-22 공개 · arXiv · PDF
video-generation long-horizon action-conditioned world-model ode-distillation vae-decoder streaming-inference low-bit-inference
Abstract
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduce LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift. Raw keyboard actions provide a unified control interface for scene roaming and third-person character interaction, while reference-character memory provides persistent appearance cues for identity consistency during third-person rollouts. For deployment, we co-design a streaming inference stack with a lightweight VAE decoder, efficient attention, memory-aware scheduling, and low-bit DiT inference. Across optimized low-bit configurations, ABot-World-0 streams 720P video at up to 16 FPS on a single NVIDIA RTX 5090 desktop GPU, with 1.2s action-to-first-frame latency and approximately 19GiB peak VRAM. Experiments on WorldRoamBench and extended interactive rollouts demonstrate competitive controllability and coherent long-horizon world evolution.
한국어 요약
한 줄 요약
ABot-World-0는 단일 데스크톱 GPU에서 실시간으로 장기적이고 일관된 인터랙티브 월드를 생성하는 행동 조건부 비디오 월드 모델이다.
핵심 기여도
- **WorldExplorer**를 통해 멀티소스 데이터(AAA 게임, 시뮬레이션, 인터넷 영상)를 통합하고 훈련 피드백에 따라 데이터 수집을 동적으로 조정.
- **LongForcing** 알고리즘을 도입하여 장기 자가롤아웃과 확장된 시간 범위의 교사 모델을 정렬, 분포 이동과 오토회귀 드리프트를 완화.
- **720P 해상도에서 16 FPS**, **1.2초의 액션-첫프레임 지연**, **최대 19GiB VRAM 소비**로 실시간 렌더링 가능.
- **WorldRoamBench**에서 경쟁력 있는 성능을 보이며 장기적 월드 진화와 제어 가능성 증명.
핵심 아이디어
ABot-World-0는 단순한 비디오 생성을 넘어, 사용자의 행동에 따라 지속적으로 진화하는 월드를 생성하는 것을 목표로 한다. 이는 단일 모델이 장기적 일관성, 제어 가능성, 실시간 성능을 모두 달성하는 체계적인 접근을 요구한다. 핵심 아이디어는 **LongForcing**를 통해 장기 자가롤아웃과 확장된 시간 범위의 교사 모델을 정렬하는 것이다. 이는 단기 훈련 상태와 장기 추론 상태 간의 불일치를 줄이고, 시각적 오류 누적을 억제한다. 또한, **Reference-Character Memory**를 도입하여 장기적인 3인칭 롤아웃에서 일관된 외형 정보를 제공함으로써 인물의 정체성을 유지한다.
기술적 접근법
- **WorldExplorer**: 멀티모달 게임 및 시뮬레이션 트레이젝토리 수집, 14개의 결정론적 품질 검사, VLM 기반 평가, 동작 및 텍스트 주석 처리.
- **Bidirectional Teacher → Causal Student**: 교사 강요(Teacher Forcing)와 ODE 디스틸레이션을 통해 이중 방향 교사 모델을 인과적 학습자 모델로 점진적으로 변환.
- **LongForcing**: 장기 자가롤아웃과 확장된 시간 범위의 교사 모델을 정렬하여 분포 이동을 완화.
- **배포 스택**: 가벼운 VAE 디코더, 효율적인 어텐션, 메모리 인식 스케줄링, 저비트 DiT 추론을 결합하여 720P 16 FPS, 1.2초 지연, 19GiB VRAM 소비 달성.
주요 결과
- **WorldRoamBench**에서 경쟁력 있는 성능: 행동 추종, 경로 추종, 시각 품질, 물리 메커니즘, 메모리 유지 등에서 높은 점수.
- **60초 롤아웃 평가**에서 LongForcing은 Causal-Forcing 기반 기준 대비 시각적 오류 누적을 감소시켜 장기적 시각 안정성 향상.
- **일일 롤아웃**에서 시각적 품질, 활성 역학, 장면 일관성을 유지하며 콜랩스 없이 지속적 생성 가능.
- **단일 RTX 5090 GPU**에서 720P 16 FPS, 1.2초 지연, 19GiB VRAM 소비로 실시간 렌더링 가능.
의의 및 한계
ABot-World-0는 단일 모델로 장기적 월드 진화, 제어 가능성, 실시간 성능을 모두 달성함으로써 인터랙티브 월드 모델링의 새로운 기준을 제시한다. 특히, **LongForcing**을 통해 장기 자가롤아웃의 안정성을 향상시키는 점이 학술적·실용적 가치를 높인다. 그러나, **고해상도(예: 1080P 이상) 실시간 렌더링**은 여전히 어려운 점이며, **복잡한 세마틱 이벤트나 조건**에 대한 처리는 추가 연구가 필요하다. 또한, **장기적 월드 진화의 일관성**은 여전히 제한적일 수 있으며, **더 복잡한 물리적 상호작용**을 지원하기 위한 확장이 필요하다.
실용적 활용
ABot-World-0는 게임 시뮬레이션, 인터랙티브 콘텐츠 생성, 에이전트 학습, 임바디드 AI 연구 등에 활용 가능하다. 특히, **단일 데스크톱 GPU에서 실시간 월드 생성**이 가능하므로, **소규모 연구실이나 개발자**에게도 접근성이 높은 도구로 활용될 수 있다. **Reference-Character Memory**와 **LongForcing**를 통해 장기적 인물 일관성과 월드 안정성을 유지할 수 있어, **게임 개발 및 VR/AR 환경 구축**에 적합하다.