PAWBench:確率的整合世界モデリングへの到達度はどの程度か?
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
August 27, 2026
著者: Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram Đorđević, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen
cs.AI
要旨
近年のビデオ生成モデルは、ますます世界モデルとして捉えられるようになっている。多くの物理的プロセスは、複数の妥当な経路で進行し得る。したがって、世界モデルは、もっともらしい軌道だけでなく、同一の初期観測と行動の下で可能な行動の分布も再現すべきである。我々は、この分布レベルの要件を確率的アライメントと呼ぶ。しかしながら、既存の評価は主に個々のビデオの妥当性を評価しており、繰り返し生成が正しい分布を回復するかどうかを検証していない。これにより、中心的な疑問が生じる:現在のビデオ生成モデルは、確率的に整合した世界モデリングからどれほど離れているのか? これに答えるため、我々は確率的アライメントを世界モデルの分布的基準として形式化し、世界力学の確率的サンプラーとしてビデオ生成モデルを評価するベンチマークであるPAWBenchを導入する。さらに、繰り返しビデオロールアウトを可能な物理的行動に関する経験分布に変換する、結果レベルのプロトコルであるPAWEvalを導入する。50のシナリオと11の現在のシステムにわたって、妥当な行動の範囲を回復しつつ参照確率に一貫して一致するモデルは存在しなかった。このギャップを確認した上で、我々は言語プロンプト、初期ノイズサンプリング、またはモデル学習がモデルの予測分布を再形成できるかを検証する。我々は、本研究が確率的に整合した世界モデリングに向けた今後の取り組みの基盤となると考えている。
English
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.