ChatPaper.aiChatPaper

Principia:视频模型的关系性物理测试

Principia: Relational Physics Tests for Video Models

September 3, 2026
作者: Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan, Anand Bhattad
cs.AI

摘要

评估视频模型中的物理推理能力具有挑战性,因为绝对运动测量依赖于帧率、物体尺度和相机标定,而这些因素在生成视频中往往存在歧义或不可获得。我们提出了一种不同的方法。当同一场景中的两个物体遵循相同的物理定律时,它们的运动必须满足可预测的关系,且这些关系独立于标定而成立。我们引入了Principia基准,通过配对物体之间的关系一致性来评估牛顿物理学。Principia涵盖八种现象——重力、恢复系数、摩擦力、转动惯量、抛体运动、动量、单摆和质量-弹簧振荡——涉及平移、旋转、碰撞和振荡动力学,并采用受控协议记录的真实场景。我们还引入了一种与标定无关的一致性评分,可直接在图像空间中量化物理违背程度。在来自六种最先进视频生成器的数千次生成结果中,尽管所有模型在VBench上的得分均约为0.8,但没有模型在Principia上的得分超过0.42。我们还评估了视觉语言模型检测关系性物理违背的能力,最佳模型仅达到67%的准确率,而大多数模型的性能接近随机水平。
English
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the same physical law, their motions must satisfy predictable relationships, and these relationships hold independent of calibration. We introduce Principia, a benchmark that evaluates Newtonian physics through relational consistency between paired objects. Principia spans eight phenomena - gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass-spring oscillation - across translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. We also introduce a calibration-independent consistency score that quantifies physical violation directly in image space. Across thousands of generations from six state-of-the-art video generators, no model exceeds 0.42 on Principia despite all scoring around 0.8 on VBench. Vision-language models are evaluated on their ability to detect relational physics violations, with the best model achieving only 67% accuracy and most performing near chance level.