PAWBench:我们距离概率对齐的世界建模还有多远?
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
August 27, 2026
作者: Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram Đorđević, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen
cs.AI
摘要
近期视频生成模型日益被视作世界模型。许多物理过程可以以不止一种有效方式展开。因此,世界模型不仅应再现一条合理的轨迹,还应再现相同初始观测和动作下可能行为的分布。我们将这种分布层面的要求称为概率对齐。然而,现有评估大多关注单个视频的合理性,并未检验重复生成是否恢复了正确的分布。这引出了一个核心问题:现有视频生成器距离概率对齐的世界建模还有多远?为回答这一问题,我们将概率对齐形式化为世界模型的分布性准则,并引入PAWBench——一个将视频生成器作为世界动力学随机采样器进行评估的基准。我们进一步提出PAWEval,一种结果层面的评估方案,将重复视频展开转化为可能物理行为上的经验分布。在50个场景和11个现有系统中,没有模型能在恢复有效行为范围的同时始终与参考概率保持一致。在确认这一差距后,我们检验了语言提示、初始噪声采样或模型训练是否能重塑模型的预测分布。我们相信,我们的工作可以为未来迈向概率对齐世界建模的努力奠定基础。
English
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.