WorldExam:從表面外觀到內在反應性的世界模型基準測試
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
August 3, 2026
作者: Yuxue Yang, Shuyao Shang, Jiahe Wang, Zitong Zhou, Liang Tan, Junhan Zeng, Ruizhi Li, Junyan Li, Yu Liu, Xiao Yang, Yong Li, Jun Zhu, Hongsheng Li, Tieniu Tan, Lue Fan, Zhaoxiang Zhang
cs.AI
摘要
可控影片生成模型日益被開發為世界模型。因此,評估它們在此角色中的表現,需超越生成影片的表面外觀,延伸到其所描繪世界的內在反應性:即從場景狀態推斷世界應如何反應,並生成輸入中未明確描述的合理後果的能力。然而,現有基準評測主要透過檢查請求的動作和互動結果是否實現,來評估視覺品質或明確指令的達成度,導致內在反應性未被充分檢驗。我們提出 WorldExam,一個涵蓋四個層級的分層診斷基準評測:視覺品質、控制遵循度、空間一致性與世界反應性。它包含八項專門任務、共 1,474 個案例,並支持對攝影機驅動、動作驅動及語言驅動三種模型範式的統一評估。世界反應性層級評估超越輸入中明確指定內容的場景條件反應與目標導向行為。對 20 個具代表性模型的評估揭示了明顯的能力分化。攝影機驅動模型擅長攝影機控制,但其介面不支援動態互動;動作驅動模型能更精確地控制主體,但往往使世界缺乏反應;語言驅動模型在互動方面表現較佳,但對複雜控制的遵循度較低。沒有任何模型能同時兼顧廣泛的任務覆蓋與一致的強勁表現,顯示高視覺品質與明確指令達成並不能保證內在反應性。
English
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.