SemComp-Bench: 動画生成におけるセマンティックタスク完了のベンチマーキング
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
August 18, 2026
著者: Keyu Tu, Zhuowei Chen, Mengqi Huang, Yuxin Wang, Jiahao Zhu, Zhendong Mao, Yongdong Zhang
cs.AI
要旨
我々は、成果指向のビデオ生成タスクであるSemantic Task Completion Video Generationを導入する。この定式化の下では、成功には意図した成果の達成と意味的グラウンディングの両方が必要である。意味的グラウンディングは、タスクに関連する高レベルの意味の観点から、参照画像と生成された成果の間の対応関係を特徴付ける。評価は生成された成果に焦点を当てており、中間タスクステップの完全な系列の提示や、参照画像との従来の外観一貫性を要求しない。体系的評価を支援するため、6つのドメインをカバーする評価データセットであるSemComp-Dataを構築する。各インスタンスは、参照画像、詳細な指示、簡潔な指示、および成果中心のビデオクリップから構成される。スケーラブルな4段階のキュレーションパイプラインは、生ビデオを標準化されたSemComp-Dataインスタンスに変換する。さらに、視覚言語モデル(VLM)を用いて構造化された二値質問に回答する評価プロトコルであるSemComp-Benchを導入する。SemComp-Benchは、成果達成と生成信頼性について、それぞれOAスコアとGRスコアを報告する。代表的なビデオ生成モデルを用いた実験は、意図した成果を達成しながら、参照画像におけるタスク関連の意味的グラウンディングを維持することが依然として困難であることを示している。
English
We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires both achievement of the intended outcome and semantic grounding. Semantic grounding characterizes the correspondence between the reference image and the generated outcome in terms of high-level semantics relevant to the task. Evaluation focuses on the generated outcome and requires neither the presentation of a complete sequence of intermediate task steps nor conventional appearance consistency with the reference image. To support systematic evaluation, we construct SemComp-Data, an evaluation dataset covering six domains. Each instance comprises a reference image, a detailed instruction, a brief instruction, and an outcome-centric video clip. A scalable four-stage curation pipeline converts raw videos into standardized SemComp-Data instances. We further introduce SemComp-Bench, an evaluation protocol that uses a vision-language model (VLM) to answer structured binary questions. SemComp-Bench reports the OA Score and the GR Score for Outcome Achievement and Generation Reliability, respectively. Experiments on representative video generation models show that achieving intended outcomes while maintaining task-relevant semantic grounding in reference images remains challenging.