ChatPaper.aiChatPaper

WorldExam: 見かけの外観から本質的な反応性へのワールドモデルのベンチマーキング

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

August 3, 2026
著者: Yuxue Yang, Shuyao Shang, Jiahe Wang, Zitong Zhou, Liang Tan, Junhan Zeng, Ruizhi Li, Junyan Li, Yu Liu, Xiao Yang, Yong Li, Jun Zhu, Hongsheng Li, Tieniu Tan, Lue Fan, Zhaoxiang Zhang
cs.AI

要旨

制御可能なビデオ生成モデルは、世界モデルとしてますます開発が進められている。それに伴い、この役割における評価は、生成されたビデオの見た目の品質だけでなく、描かれた世界が本来備える反応性、すなわちシーンの状態から世界がどのように反応すべきかを推論し、入力に明示的に記述されていない妥当な結果を生成する能力にまで拡張されている。しかし既存のベンチマークは主に、要求された動作と相互作用の結果が実現されているかを確認することで、視覚品質や明示的な指示の達成度を評価しており、本質的反応性は十分に検証されていない。 本稿では、階層的診断ベンチマークであるWorldExamを導入する。これは視覚品質、制御追従、空間的一貫性、世界反応性の4つのレベルにわたり、8つの専用タスクで1,474ケースから構成され、カメラ駆動、行動駆動、言語駆動のモデルパラダイムの統一的評価をサポートする。世界反応性レベルでは、入力に明示的に指定された内容を超えた、シーン条件付きの反応と目標指向的な行動を評価する。 20の代表的モデルの評価により、明確な能力の分断が明らかになった。カメラ駆動モデルはカメラ制御に優れるが、そのインターフェースは動的な相互作用をサポートしない。行動駆動モデルは被写体の制御はより正確であるが、世界を無反応のままにすることが多い。言語駆動モデルは相互作用では優れた性能を示すものの、複雑な制御への追従は不正確である。幅広いタスク網羅性と一貫した高性能を兼ね備えたモデルは存在せず、高い視覚品質と明示的な指示の達成は本質的反応性を保証しないことが示されている。
English
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.