VeriPhy: 세계 모델 평가 및 개선을 위한 에이전트적 물리 추론
VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement
September 2, 2026
저자: Wenzhuo Xu, Yuchen Zhu, Chongjian Ge, Xuan Shen, Jing Shi, Jason Kuen, Yongxin Chen, Molei Tao, Christopher McComb, Noelia Grande Gutiérrez, Jiuxiang Gu
cs.AI
초록
생성된 비디오의 시각적 유창성은 물리적 신뢰성을 보장하지 않으며, 스칼라 품질 점수 하나만으로는 클립이 위반하는 물리적 의무나 실패 지점을 나타낼 수 없다. 이에 우리는 VeriPhy를 제시한다. VeriPhy는 감사 가능한 물리 검증 시스템으로, 텍스트 전용 플래너가 어떤 프레임도 관측하기 전에 프롬프트를 유형화된 물리적 의무 및 정적으로 검증된 실행 계획으로 컴파일한다. 실행 중에는 관측 결과가 동결된 저수준 전문가들(예: 분할 및 추적, 계수, 추적 결과 트랙에 대한 열한 가지 유형화된 물리 측정, 깊이, OCR, 오디오 이벤트 감지)에 대한 선언된 호출만 게이팅하고 그 범위를 제한한다. 각 동작은 출처(provenance)를 지니는 증거 레코드를 반환하며, 그 페이로드는 사용 가능한 경우 유형화된 측정값 또는 명시적으로 태그된 학습 상태 중 하나이다. 유형화된 리졸버와 고정된 합성은 사용 가능한 레코드를 완전한 출처 정보와 함께 세 값 상태(지지됨, 반박됨, 알 수 없음; 이는 각각 그럴듯함, 그럴듯하지 않음, 판단 유보로 표면화됨)로 매핑하므로, 모든 판정은 그 판정을 산출한 증거까지 추적될 수 있다.
평가의 기준은 프롬프트 참조, 공간, 시간에서 실제 생성 실패의 위치를 특정하는 인간 주석 결함 기록들로 이루어진 1,500개 클립 코퍼스에 둔다. 304개의 그러한 기록을 담은 149개 클립 핵심 세트에서 VeriPhy는 228개를 포착하는 반면, 동일한 클립과 동일한 주장이 주어진 공개된 질문 분해 평가기는 164개를 포착한다. 재현율만으로는 이 방식이 동일한 백본을 모놀리식하게 프롬프팅하는 방식과 구별되지 않으며, 그 방식은 222개에 도달한다. 둘을 실제로 구별하는 것은 각 판정이 자신의 증거 레코드와 출처를 유지한다는 점이며, 이로써 추적 기록은 판정 단위로 감사 가능해지고, 비평가 판정이 생성 단계로 다시 환류될 수 있는 인터페이스로 사용될 수 있다.
English
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is incapable of indicating the obligation a clip violates or the moment it fails. We present VeriPhy, an auditable physical-verification system in which a text-only planner compiles the prompt into typed physical obligations and a statically validated execution plan before any frame is observed. During execution, observations gate and scope only declared calls to frozen low-level experts (e.g., segmentation and tracking, counting, eleven typed physical measurements over the resulting tracks, depth, OCR, and audio-event detection). Each action returns a provenance-carrying evidence record whose payload, when usable, is either a typed measurement or an explicitly tagged learned state. Typed resolvers and fixed composition map usable records to a three-valued state (supported, contradicted, or unknown, surfaced as plausible, implausible, or abstain) with full provenance, so that every verdict is traceable to the evidence that produced it. We anchor evaluation in a 1,500-clip corpus of human-annotated flaw records that localize real generation failures in prompt reference, space, and time. On a 149-clip core carrying 304 such records, VeriPhy accounts for 228, against 164 for a published question-decomposition evaluator given the same clips and the same claims. Recall alone does not separate it from prompting the same backbone monolithically, which reaches 222; what separates them is that each decision retains its evidence record and provenance, making the traces auditable one verdict at a time and usable as the interface through which a critic verdict could be written back into generation.