WorldExam: 겉보기 외형에서 내재적 반응성까지 세계 모델 벤치마킹
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
August 3, 2026
저자: Yuxue Yang, Shuyao Shang, Jiahe Wang, Zitong Zhou, Liang Tan, Junhan Zeng, Ruizhi Li, Junyan Li, Yu Liu, Xiao Yang, Yong Li, Jun Zhu, Hongsheng Li, Tieniu Tan, Lue Fan, Zhaoxiang Zhang
cs.AI
초록
제어 가능한 비디오 생성 모델은 점점 더 세계 모델로서 개발되고 있다. 이에 따라 이러한 역할에서의 평가는 생성된 비디오의 표면적 외관을 넘어, 묘사된 세계의 내재적 반응성, 즉 장면 상태로부터 세계가 어떻게 반응해야 하는지를 추론하고 입력에 명시적으로 기술되지 않은 그럴듯한 결과를 생성하는 능력으로 확장된다. 그러나 기존 벤치마크는 주로 요청된 행동과 상호작용 결과가 실제로 구현되는지 확인하여 시각적 품질이나 명시적 지시 충족 여부를 평가할 뿐, 내재적 반응성은 충분히 검토하지 않는다. 우리는 시각적 품질, 제어 준수, 공간 일관성, 세계 반응성의 네 가지 수준에 걸친 계층적 진단 벤치마크인 WorldExam을 제안한다. 이는 총 8개의 전용 과제에 걸쳐 1,474개의 사례로 구성되며, 카메라 기반, 행동 기반, 언어 기반 모델 패러다임의 통합 평가를 지원한다. 세계 반응성 수준은 입력에 명시적으로 지정된 것 이상의 장면 조건 반응과 목표 지향 행동을 평가한다. 20개의 대표 모델에 대한 평가는 명확한 성능 분화를 드러낸다. 카메라 기반 모델은 카메라 제어에 뛰어나지만 동적 상호작용을 지원하는 인터페이스가 없으며, 행동 기반 모델은 피사체를 더 정밀하게 제어하지만 세계를 비반응적으로 남겨두는 경우가 많고, 언어 기반 모델은 상호작용에서 더 나은 성능을 보이지만 복잡한 제어를 덜 충실히 따른다. 광범위한 과제 커버리지와 일관된 강력한 성능을 겸비한 모델은 없으며, 이는 높은 시각적 품질과 명시적 지시 충족이 내재적 반응성을 보장하지 않음을 보여준다.
English
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.