Apple-π: 비디오 기반 추론 벤치마킹을 통한 법칙 기반 물리 지능으로의 접근
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
July 17, 2026
저자: Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang, Shulin Tian, Ziang Cao, Weichen Fan, Ziqi Huang, Yuhao Dong, Hao Li, Zhaoxi Chen, Zhongang Cai, Lei Yang, Ziwei Liu
cs.AI
초록
현대 비디오 생성 모델은 물리 법칙에 대한 내재화된 이해를 지닌 떠오르는 세계 모델로 점차 주목받고 있다. 그러나 기존 벤치마크는 대부분 출력 수준에서만 물리적 타당성을 평가할 뿐, 모델이 신뢰할 수 있고 법칙에 기반한 추론 과정을 통해 해당 결과에 도달하는지 검증하지 않는다. 본 논문에서는 비디오 모델 평가를 물리 법칙에 명시적으로 고정시키는 최초의 벤치마크인 Apple-PI를 소개한다. Apple-PI는 세 가지 구성 요소로 이루어진다. 1) 오차드(Orchard): 고전 역학의 10가지 표준 과제를 다루는 400개의 비디오로 구성된 데이터셋이다. 교란 요인이 없는 진단을 위한 단일 법칙 과제와 일반화 능력을 탐색하기 위한 다중 법칙 과제로 구분된다. 2) 벤치마크 프로토콜(Benchmark Protocol): 지각(Perception), 정식화(Formulation), 추론(Deduction)의 세 단계로 구성된 과학적 추론 기반 프로토콜이다. 인포그래픽이 주석 처리된 첫 프레임에 체인-오브-프레임 프롬프팅을 적용하며, 생성된 비디오를 모델의 가시적 추론 과정으로 간주한다. 3) 평가 스위트(Evaluation Suite): MLLM 기반 주관적 평가와 물리 법칙에 기반한 객관적 측정을 결합한 하이브리드 평가 스위트이다. 이를 통해 모델이 실패하는지 여부뿐만 아니라 어디서 실패하는지에 대한 단계별 진단이 가능하다. 11개 모델을 벤치마킹한 결과, 현재 비디오 모델은 신뢰할 수 있는 법칙 기반 세계 시뮬레이터와는 거리가 멀며, 최고 비디오 모델도 0.473점에 그쳤다. 단계별, 축별, 원천별 분석을 통해 지각-정식화-추론 병목 현상, 취약한 다중 법칙 상태 전이, 지속적인 시뮬레이션-현실 간극(sim-to-real gap)이 드러났다. 이러한 발견은 Apple-PI가 미래 비디오 모델을 법칙에 기반한 물리적 지능을 갖춘 세계 모델로 안내하기 위한 진단 기반 역할을 할 수 있음을 보여준다.
English
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We introduce Apple-PI, the first benchmark that anchors video-model evaluation explicitly in physical laws. Apple-PI comprises three components. 1) Orchard: a dataset of 400 videos covering ten canonical tasks in classical mechanics. It separates single-law tasks for confounder-free diagnosis from multi-law tasks for probing generalization. 2) Benchmark Protocol: a three-stage protocol based on scientific reasoning, including Perception, Formulation, and Deduction. It uses chain-of-frames prompting on infographic-annotated first frames, treating the generated video as the model's visible reasoning trace. 3) Evaluation Suite: a hybrid evaluation suite that combines MLLM-based subjective scoring with physics-law-grounded objective measures. This enables stage-resolved diagnosis of not only whether a model fails, but where it fails. Benchmarking 11 models shows that current video models remain far from reliable law-grounded world simulators, with the best video model scoring only 0.473. Our stage-, pillar-, and source-resolved analyses further expose a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap. These findings position Apple-PI as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.