ChatPaper.aiChatPaper

Apple-π:以影片進行思考基準測試,邁向基於法則的物理智能

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

July 17, 2026
作者: Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang, Shulin Tian, Ziang Cao, Weichen Fan, Ziqi Huang, Yuhao Dong, Hao Li, Zhaoxi Chen, Zhongang Cai, Lei Yang, Ziwei Liu
cs.AI

摘要

现代视频生成模型日益被视为内化了物理规律的新兴世界模型。然而,现有基准测试大多仅在输出层面评估物理合理性,而未验证模型是否通过忠实且符合规律的推理过程得出结果。我们提出Apple-PI,这是首个将视频模型评估明确锚定于物理定律的基准框架。Apple-PI包含三个组成部分:1)果园数据集(Orchard):涵盖经典力学中十项典型任务,包含400个视频片段。该数据集区分了用于排除混杂因素诊断的单一定律任务与用于探索泛化能力的多定律任务。2)基准协议(Benchmark Protocol):基于科学推理的三阶段协议,包括感知、形式化与演绎。该协议采用基于信息图表注释的首帧逐帧链式提示,将生成视频视为模型可视化的推理轨迹。3)评估套件(Evaluation Suite):混合评估套件,结合基于多模态大语言模型的主观评分与基于物理定律的客观度量。这实现了分阶段诊断——不仅能判断模型是否失败,更能定位其失败阶段。对11个模型的基准测试表明,当前视频模型远非可靠的规律性世界模拟器,最佳视频模型仅得0.473分。我们的分阶段、分维度、分来源分析进一步揭示了从感知到形式化再到演绎的瓶颈、多定律间状态传递薄弱以及持续存在的模拟到现实差距。这些发现使Apple-PI成为指导未来视频模型迈向具备规律性物理智能的世界模型的诊断基础。
English
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We introduce Apple-PI, the first benchmark that anchors video-model evaluation explicitly in physical laws. Apple-PI comprises three components. 1) Orchard: a dataset of 400 videos covering ten canonical tasks in classical mechanics. It separates single-law tasks for confounder-free diagnosis from multi-law tasks for probing generalization. 2) Benchmark Protocol: a three-stage protocol based on scientific reasoning, including Perception, Formulation, and Deduction. It uses chain-of-frames prompting on infographic-annotated first frames, treating the generated video as the model's visible reasoning trace. 3) Evaluation Suite: a hybrid evaluation suite that combines MLLM-based subjective scoring with physics-law-grounded objective measures. This enables stage-resolved diagnosis of not only whether a model fails, but where it fails. Benchmarking 11 models shows that current video models remain far from reliable law-grounded world simulators, with the best video model scoring only 0.473. Our stage-, pillar-, and source-resolved analyses further expose a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap. These findings position Apple-PI as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.