Apple-π:面向基于物理定律的物理智能的视频思考基准测试
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
July 17, 2026
作者: Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang, Shulin Tian, Ziang Cao, Weichen Fan, Ziqi Huang, Yuhao Dong, Hao Li, Zhaoxi Chen, Zhongang Cai, Lei Yang, Ziwei Liu
cs.AI
摘要
现代视频生成模型正日益被视为具备内在物理定律理解能力的新兴世界模型。然而,现有基准主要仅在输出层面评估物理合理性,并未验证模型是否通过忠实且基于定律的推理过程得出结果。我们引入Apple-PI,这是首个明确以物理定律为锚点对视频模型进行评估的基准。Apple-PI包含三个组成部分:1) Orchard:一个包含400个视频的数据集,覆盖经典力学中的十项典型任务。该数据集将单定律任务(用于无混杂因素诊断)与多定律任务(用于探测泛化能力)分离。2) 基准测试协议:基于科学推理的三阶段协议,包括感知、公式化和演绎。它采用对信息图表标注的首帧进行帧链提示的方法,将生成的视频视为模型的可视化推理轨迹。3) 评估套件:一个混合评估套件,结合基于多模态大语言模型的主观评分与基于物理定律的客观度量。这使得不仅能诊断模型是否失败,还能分阶段定位其失败环节。对11个模型的基准测试表明,当前视频模型仍远非可靠的、基于定律的世界模拟器,最佳视频模型得分仅为0.473。我们的分阶段、分支柱及分来源分析进一步揭示了从感知到公式化再到演绎的瓶颈、多定律状态转移能力薄弱以及持续存在的模拟到现实差距。这些发现将Apple-PI定位为诊断基础,旨在引导未来视频模型向具备基于定律的物理智能的世界模型发展。
English
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We introduce Apple-PI, the first benchmark that anchors video-model evaluation explicitly in physical laws. Apple-PI comprises three components. 1) Orchard: a dataset of 400 videos covering ten canonical tasks in classical mechanics. It separates single-law tasks for confounder-free diagnosis from multi-law tasks for probing generalization. 2) Benchmark Protocol: a three-stage protocol based on scientific reasoning, including Perception, Formulation, and Deduction. It uses chain-of-frames prompting on infographic-annotated first frames, treating the generated video as the model's visible reasoning trace. 3) Evaluation Suite: a hybrid evaluation suite that combines MLLM-based subjective scoring with physics-law-grounded objective measures. This enables stage-resolved diagnosis of not only whether a model fails, but where it fails. Benchmarking 11 models shows that current video models remain far from reliable law-grounded world simulators, with the best video model scoring only 0.473. Our stage-, pillar-, and source-resolved analyses further expose a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap. These findings position Apple-PI as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.