ChatPaper.aiChatPaper

WorldExam:从表观外观到内在反应性的世界模型基准测试

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

August 3, 2026
作者: Yuxue Yang, Shuyao Shang, Jiahe Wang, Zitong Zhou, Liang Tan, Junhan Zeng, Ruizhi Li, Junyan Li, Yu Liu, Xiao Yang, Yong Li, Jun Zhu, Hongsheng Li, Tieniu Tan, Lue Fan, Zhaoxiang Zhang
cs.AI

摘要

可控视频生成模型正越来越多地被开发为世界模型。因此,在这一角色下对其进行评估,需要超越生成视频的表观外观,深入考察其所描绘世界的固有反应性:即从场景状态推断世界应如何反应,并生成输入中未明确描述的合理后果的能力。然而,现有基准测试主要评估视觉质量或显式指令执行情况(通过检查所请求的动作和交互结果是否实现),对固有反应性的考察不足。我们提出WorldExam,一个分层诊断基准测试,涵盖四个层级:视觉质量、控制遵从性、空间一致性和世界反应性。该基准包含八个专门任务中的1,474个案例,支持对相机驱动、动作驱动和语言驱动模型范式的统一评估。世界反应性层级评估超出输入显式规定的场景条件反应和目标导向行为。对20个代表性模型的评估揭示了明显的能力分化。相机驱动模型擅长相机控制,但其接口不支持动态交互;动作驱动模型对主体的控制更精确,但往往使世界缺乏响应;语言驱动模型在交互方面表现更好,但对复杂控制的遵循度较低。没有任何模型能同时兼顾广泛的任务覆盖和持续强劲的表现,这表明高视觉质量和显式指令执行并不能保证固有反应性。
English
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.