PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram Đorđević, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen

arXiv:2608.27345 · 2026-08-28 공개 · arXiv · PDF

video-generation world-models model-training stochastic-sampling distributional-criteria language-prompts initial-noise-sampling probabilistic-alignment

Abstract

Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.

한국어 요약

한 줄 요약

PAWBench는 영상 생성 모델이 확률적으로 정렬된 세계 모델링에 얼마나 가까운지를 평가하는 벤치마크로, 50개 시나리오에서 기존 11개 시스템이 모두 일관된 분포를 재현하지 못함을 밝힘.

핵심 기여도

핵심 아이디어

기존 영상 생성 모델은 단일 영상의 타당성만 평가받으며, 동일한 초기 관찰과 행동에 대해 가능한 미래의 분포를 정확히 모델링하는지 여부는 무시되어 왔다. 이는 **확률적 정렬**(probabilistic alignment)이라는 개념으로 정의되는데, 이는 동일한 조건에서 생성된 영상이 **물리적으로 가능한 미래의 전체 범위**(valid behaviors)와 **각각의 발생 확률**(probability mass)을 반영해야 한다는 것을 의미한다.

이를 위해 연구팀은 **PAWBench**를 제안하며, 50개 시나리오를 8개의 물리 메커니즘 그룹으로 분류하고, **PAW-Calibration**과 **PAW-Coverage**이라는 두 평가 스위트를 구성했다. PAW-Calibration은 분포가 수학적으로 명시된 시나리오(예: 동전 던지기, 바퀴 회전)를 대상으로 하며, PAW-Coverage는 가능한 결과는 열거 가능하지만 확률은 명시되지 않은 시나리오(예: 볼링공 굴리기, 병 뒤집기)를 대상으로 한다.

기술적 접근법

주요 결과

의의 및 한계

PAWBench는 영상 생성 모델이 단순히 타당한 영상만 생성하는 것을 넘어, **확률적 분포를 정확히 모델링하는 능력을 평가**할 수 있는 첫 번째 시도로, 세계 모델링의 학술적 기준을 정립하는 데 기여한다. 그러나 PAWBench는 **물리적 메커니즘의 범위가 제한적**이며, **모든 가능한 행동 분포를 포괄하지는 못**한다. 또한, **모델의 학습 분포 자체를 재구성하는 방법론은 여전히 미흡**하며, **언어나 노이즈 조절은 분포를 일관되게 변화시키지 못**한다는 한계가 드러났다.

실용적 활용

PAWBench는 **로봇 제어, 자율 주행, 시뮬레이션 기반 학습** 등에서 미래 행동의 확률적 분포를 정확히 예측해야 하는 분야에 적용 가능하다. 예를 들어, **자율 주행 시스템**이 동일한 상황에서 다양한 운전 방식을 고려할 수 있도록 도울 수 있으며, **강화 학습 환경**에서 보다 현실적인 세계 모델을 학습하는 데 활용될 수 있다.