ChatPaper.aiChatPaper

Principia: ビデオモデル向け関係性物理学テスト

Principia: Relational Physics Tests for Video Models

September 3, 2026
著者: Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan, Anand Bhattad
cs.AI

要旨

動画モデルにおける物理的推論の評価は、絶対的な運動測定がフレームレート、物体スケール、カメラキャリブレーションに依存し、それらが生成動画では曖昧または利用不可能であることが多いため困難である。本研究では、異なるアプローチを提案する。同一シーン内の2つの物体が同じ物理法則に従う場合、それらの運動は予測可能な関係を満たさなければならず、この関係はキャリブレーションとは独立に成立する。我々は、対となる物体間の関係的一貫性を通じてニュートン力学を評価するベンチマークPrincipiaを導入する。Principiaは、制御されたプロトコルで記録された実世界シーンを用いて、並進・回転・衝突・振動のダイナミクスにわたり、重力、反発、摩擦、回転慣性、投射運動、運動量、振り子、質量-ばね振動の8つの現象を網羅する。また、物理則違反を画像空間で直接定量化する、キャリブレーション非依存の整合性スコアも導入する。6つの最先端動画生成モデルによる数千の生成結果において、全モデルがVBenchで約0.8を記録しているにもかかわらず、Principiaで0.42を超えるモデルはなかった。視覚言語モデルについては、物体間の関係における物理則違反を検出する能力を評価したところ、最高性能のモデルでも精度は67%にとどまり、大半のモデルは偶然水準に近い性能であった。
English
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the same physical law, their motions must satisfy predictable relationships, and these relationships hold independent of calibration. We introduce Principia, a benchmark that evaluates Newtonian physics through relational consistency between paired objects. Principia spans eight phenomena - gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass-spring oscillation - across translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. We also introduce a calibration-independent consistency score that quantifies physical violation directly in image space. Across thousands of generations from six state-of-the-art video generators, no model exceeds 0.42 on Principia despite all scoring around 0.8 on VBench. Vision-language models are evaluated on their ability to detect relational physics violations, with the best model achieving only 67% accuracy and most performing near chance level.