ChatPaper.aiChatPaper

Apple-π: 動画を用いた思考のベンチマーク評価 ─ 法則に基づく物理的知能に向けて

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

July 17, 2026
著者: Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang, Shulin Tian, Ziang Cao, Weichen Fan, Ziqi Huang, Yuhao Dong, Hao Li, Zhaoxi Chen, Zhongang Cai, Lei Yang, Ziwei Liu
cs.AI

要旨

近年の動画生成モデルは、物理法則を内在化した新興の世界モデルとして広く注目されている。しかし既存のベンチマークの大半は、出力レベルでの物理的妥当性のみを評価しており、モデルが法則に基づいた確かな推論プロセスを経て出力に至ったかどうかを検証していない。本論文では、動画モデルの評価を物理法則に明示的に基づかせた初のベンチマーク「Apple-PI」を提案する。Apple-PIは三つの構成要素から成る。1) Orchard:古典力学における10の標準タスクをカバーする400本の動画データセット。交絡因子のない診断を可能とする単一法則タスクと、汎化性能を探る複合法則タスクを分離している。2) ベンチマークプロトコル:科学的推論に基づく「知覚」「定式化」「演繹」の三段階プロトコル。インフォグラフィックで注釈付けされた先頭フレームに対してフレーム連鎖プロンプティングを適用し、生成された動画をモデルの可視化された推論過程として扱う。3) 評価スイート:マルチモーダル大規模言語モデル(MLLM)による主観的スコアリングと、物理法則に基づく客観的指標を組み合わせたハイブリッド評価スイート。これにより、モデルが失敗したかどうかだけでなく、どの段階で失敗したかを段階別に診断できる。11のモデルを評価した結果、現状の動画モデルは信頼できる法則準拠型世界シミュレーターには程遠く、最良の動画モデルでもスコアは0.473にとどまった。段階別・柱別・原因別の分析により、知覚→定式化→演繹のボトルネック、複合法則間の状態遷移の弱さ、そして持続的なシミュレーション現実ギャップが明らかになった。これらの知見により、Apple-PIは将来の動画モデルを法則に基づく物理的知能を持つ世界モデルへと導く診断基盤として位置づけられる。
English
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We introduce Apple-PI, the first benchmark that anchors video-model evaluation explicitly in physical laws. Apple-PI comprises three components. 1) Orchard: a dataset of 400 videos covering ten canonical tasks in classical mechanics. It separates single-law tasks for confounder-free diagnosis from multi-law tasks for probing generalization. 2) Benchmark Protocol: a three-stage protocol based on scientific reasoning, including Perception, Formulation, and Deduction. It uses chain-of-frames prompting on infographic-annotated first frames, treating the generated video as the model's visible reasoning trace. 3) Evaluation Suite: a hybrid evaluation suite that combines MLLM-based subjective scoring with physics-law-grounded objective measures. This enables stage-resolved diagnosis of not only whether a model fails, but where it fails. Benchmarking 11 models shows that current video models remain far from reliable law-grounded world simulators, with the best video model scoring only 0.473. Our stage-, pillar-, and source-resolved analyses further expose a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap. These findings position Apple-PI as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.