Adaptive Latent Capacity for World Models

Idan Achituve, Lior Dikstein, Idit Diamant, Arnon Netzer, Hai Victor Habi

arXiv:2609.32921 · 2026-10-07 공개 · arXiv · PDF

world-models latent-representation predictive-modeling visual-control jepa adaptive-latent mixsigreg recursive-planning

Abstract

We introduce Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation. To encourage this ordering, ALeWM learns a sequence-conditioned distribution over prefix lengths and trains the predictor to estimate the full next embedding from a sampled input prefix. As standard anti-collapse objectives encourage variation across latent coordinates and do not organize them by predictive importance, we also introduce MixSIGReg. MixSIGReg regularizes the masked embeddings against a prior-weighted mixture with Gaussian active prefixes and zeros in the remaining coordinates. As a result, the ALeWM objective encourages early coordinates to retain information useful for prediction and recursive planning. Our analysis shows that the mixture distribution used by MixSIGReg assigns higher variance to earlier coordinate blocks and lower variance to later ones. In addition, we show that, under specified assumptions, prediction error is minimized by placing the information most useful for prediction in earlier blocks. Empirically, we study the behavior of ALeWM in a controlled dynamical system with known state variables and in goal-conditioned visual control. We show that ALeWM consistently achieves higher mean success rates than tuned fixed-width LeWM, with lower planning capacity on average.

한국어 요약

한 줄 요약

ALeWM은 JEPA 기반 월드 모델에서 예측 정보를 짧은 접두사에 집중시키는 Adaptive Latent Capacity 기법을 제안한다.

핵심 기여도

핵심 아이디어

기존 월드 모델은 예측 정보를 전체 잠재 좌표에 동일하게 분포시키지만, ALeWM은 이 정보를 짧은 접두사에 집중시켜 계획 효율성을 높인다. 이는 특히, 재귀적 예측과 목표 비교를 통해 동작을 생성하는 월드 모델에서 유리하다. MixSIGReg는 SIGReg의 등방성 가우시안 가정을 수정하여, 앞쪽 좌표에 더 높은 분산을 부여함으로써 예측에 중요한 정보가 앞쪽에 배치되도록 유도한다. 이는 수학적으로 예측 오차를 최소화하기 위해 정보를 앞쪽 좌표에 배치하는 것이 최적임을 보여준다.

기술적 접근법

주요 결과

의의 및 한계

ALeWM은 월드 모델에서 잠재 표현의 유연한 활용을 가능하게 하며, 다양한 태스크에 따라 접두사 길이를 자동으로 조절함으로써 계획 효율성을 높인다. MixSIGReg는 기존 SIGReg의 한계를 극복하고, 예측에 중요한 정보가 앞쪽 좌표에 집중되도록 유도함으로써 성능 향상에 기여한다. 그러나, ALeWM은 학습 초기에 capacity network가 비효율적으로 작동할 수 있으며, 이는 학습 과정에서 비자연적인 분포 이동을 유발할 수 있다. 또한, 특정 환경에서는 고정된 접두사 길이가 더 효과적일 수 있으므로, 모든 상황에서 adaptive capacity가 최적이라는 보장은 없다.

실용적 활용

ALeWM은 로봇 제어, 시뮬레이션 기반 학습, 비전 기반 탐색 등에서 유용하게 활용될 수 있다. 특히, 다양한 태스크나 환경에 따라 자동으로 계획 용량을 조절할 수 있어, 실시간 성능 최적화가 필요한 시스템에 적합하다.