PAWBench: 확률적으로 정렬된 세계 모델링에서 우리는 어디까지 왔는가?

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

August 27, 2026
저자: Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram Đorđević, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen
cs.AI

초록

최근 비디오 생성 모델은 점점 더 세계 모델(world model)로 간주되고 있다. 많은 물리적 과정은 둘 이상의 유효한 방식으로 전개될 수 있다. 따라서 세계 모델은 그럴듯한 궤적 하나만 재현해서는 안 되며, 동일한 초기 관측과 행동 하에서 가능한 행동들의 분포도 재현해야 한다. 우리는 이러한 분포 수준의 요구사항을 확률적 정합(probabilistic alignment)이라고 부른다. 그러나 기존 평가는 주로 개별 비디오의 타당성을 평가할 뿐, 반복 생성이 올바른 분포를 복원하는지 여부는 검증하지 않는다. 이는 핵심 질문을 제기한다: 현재 비디오 생성기는 확률적으로 정합된 세계 모델링에서 얼마나 떨어져 있는가? 이 질문에 답하기 위해, 우리는 확률적 정합을 세계 모델의 분포적 기준으로 정식화하고, 비디오 생성기를 세계 역학의 확률적 샘플러로 평가하기 위한 벤치마크인 PAWBench를 도입한다. 또한 반복적인 비디오 롤아웃을 가능한 물리적 행동들에 대한 경험적 분포로 변환하는 결과 수준 프로토콜인 PAWEval을 제안한다. 50개의 시나리오와 11개의 현재 시스템에 걸쳐, 유효한 행동의 범위를 복원하면서 기준 확률과 일관되게 일치하는 모델은 없었다. 이러한 격차를 확인한 후, 우리는 언어 프롬프트, 초기 노이즈 샘플링, 또는 모델 훈련이 모델의 예측 분포를 재형성할 수 있는지 시험한다. 우리는 이 연구가 확률적으로 정합된 세계 모델링으로 나아가기 위한 향후 노력의 기초가 될 수 있다고 믿는다.
English
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.
PDF731August 29, 2026