Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

Ying Yang, Guiyu Zhang, Lianghua Huang, Chang Nie, Chenyang Si, Haofan Wang, Shaoshuai Shi, Li Jiang

arXiv:2610.02521 · 2026-10-05 공개 · arXiv · PDF

world-models multimodal-llm long-video-generation memory-management spatial-memory generation-stability spatial-clustering reliability-aware-filtering

Abstract

Long-video generation and world models have shown strong potential for interactive entertainment and embodied simulation by predicting future observations conditioned on user actions and historical memory. However, as memory sequences grow longer and their structures become increasingly complex, managing long-range spatial context becomes increasingly challenging, calling for a more intelligent and systematic memory-management strategy. Building on the advancing spatial reasoning capabilities of multimodal large language models (MLLMs) and the broader vision of unified models, we propose Spatial Memory Intelligence (SMI), the first framework to systematically employ an understanding model for spatial-memory management in long-video world models. SMI introduces four coordinated atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering. Extensive experiments across multiple baselines, benchmarks, and world-model backbones demonstrate the effectiveness and generalizability of SMI, achieving comprehensive improvements in memory sparsity, generation stability, and spatial consistency.

한국어 요약

한 줄 요약

SMI는 장기 비디오 월드 모델에서 공간 기억 관리를 위한 첫 번째 통합 프레임워크로, 4가지 원자적 연산을 통해 기억 효율성과 생성 안정성을 향상시킨다.

핵심 기여도

핵심 아이디어

기존 월드 모델은 장기 기억 관리에서 공간 일관성과 생성 안정성 문제를 겪는다. SMI는 MLLM을 활용해 공간 기억을 체계적으로 관리함으로써 이 문제를 해결한다. 핵심 아이디어는 공간 기억을 4가지 원자적 연산으로 분해하여, 각 단계에서 공간 및 의미적 판단을 수행하는 것이다. 예를 들어, spatial clustering은 공간적으로 인접한 기억을 클러스터로 그룹화하고, within-cluster sparsification은 클러스터 내 중복된 관측을 제거하여 저장 비용을 줄인다. Action-aware retrieval은 현재 행동과 최근 맥락을 결합해 가장 유용한 기억을 검색하며, reliability-aware filtering은 시각 드리프트가 있는 새로운 관측을 메모리에 저장하지 않도록 한다. 이러한 연산은 MLLM이 학습된 의미적 및 공간적 판단을 기반으로 수행된다.

기술적 접근법

주요 결과

의의 및 한계

SMI는 장기 비디오 월드 모델에서 공간 기억 관리를 체계적으로 수행할 수 있는 첫 번째 통합 프레임워크로, MLLM의 공간 및 의미적 추론 능력을 활용한 새로운 접근법을 제시한다. 특히, 4가지 원자적 연산을 통해 기억 효율성과 생성 안정성을 동시에 개선한 점이 학술적 및 실용적 가치를 높인다. 그러나 MLLM의 계산 비용이 높아 실시간 적용에는 한계가 있을 수 있으며, 특정 장면에서 의미적 판단 오류가 발생할 가능성도 있다. 또한, 현재 실험은 특정 벤치마크에 제한되어 있어 더 다양한 환경에서의 일반화 능력을 검증해야 한다.

실용적 활용

SMI는 인터랙티브 엔터테인먼트, 가상 현실, 자율 시스템 등에서 장기 비디오 시뮬레이션을 구현하는 데 유용하다. 특히, 대규모 장면에서의 공간 일관성과 생성 안정성이 중요한 애플리케이션에 적합하며, MLLM 기반의 기억 관리는 복잡한 환경에서도 유연한 시뮬레이션을 가능하게 한다.