VGI-Bench: 動画生成モデルにおける視覚的知能の探求

VGI-Bench: Probing Visual Intelligence in Video Generation Models

August 26, 2026
著者: Xuan He, Cong Wei, Yuhao Cheng, Linrui Ma, Yuxuan Zhang, Zuojun Li, Yuhao Wen, Jize Jiang, Zeyi Liu, Yuren Hao, Songcheng Cai, Keming Wu, Penghui Du, Kai Zou, Rui Yang, Chenkai Sun, Ke Yang, Ping Nie, Kelsey R Allen, Chenglong Wang, Michel Galley, Jianfeng Gao, ChengXiang Zhai
cs.AI

要旨

近年の研究は、動画生成モデルが生成フレームを通じてゼロショット視覚推論の特定の形態を示し得ることを示唆している。しかし、信頼性の高い評価は依然として困難であり、ベンチマークは、現在の動画モデルの視覚的先行知識と整合する入力を採用し、もっともらしい最終状態だけでなく妥当な進化過程を要求し、挑戦的でありつつ部分的に達成可能な難易度に調整されるべきである。この目的のために、我々はVGI-benchを導入する。これは27タスクと810インスタンスを含み、タスク領域とスキルタグの2階層分類法により、動画生成モデルの視覚推論能力の詳細な評価を可能にする。我々の評価は、現在の生成システムが視覚に基づく推論タスクの一部を解くことができる一方で、信頼性には程遠く、最強モデルであるSeedance 2.0でさえ我々の評価基準の下では51.0%しか達成しないことを示している。さらに、我々の分析は、出力の失敗モード、入力条件への感度、合成データによる微調整からの性能転移の限界、および内部ノイズ除去の観点から、限定的な自己補正——後のステップは主に初期仮説を洗練するだけで、推論エラーを修正しない——を明らかにする。VGI-benchが次世代動画生成モデルの発展を促進する一助となることを期待する。ウェブサイト: https://hexuan21.github.io/VGI-Bench/
English
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible. To this end, we introduce VGI-bench, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generation models. Our evaluations show that current generative systems can solve a subset of visually grounded reasoning tasks, but remain far from reliable, with even the strongest model, Seedance 2.0, achieving only 51.0% under our evaluation criteria. Our analysis further explore the output failure modes, input condition sensitivity, performance transfer boundary from synthetic fine-tuning, and internal denoising perspective revealing limited self-correction, where later steps mainly refine early hypotheses rather than correct reasoning errors. We hope VGI-bench will help stimulate the development of next-generation video generation models. Website: https://hexuan21.github.io/VGI-Bench/
PDF1431August 28, 2026