SemComp-Bench:影片生成中語義任務完成的基準測試
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
August 18, 2026
作者: Keyu Tu, Zhuowei Chen, Mengqi Huang, Yuxin Wang, Jiahao Zhu, Zhendong Mao, Yongdong Zhang
cs.AI
摘要
我們引入「語義任務完成影片生成」這項以成果為導向的影片生成任務。在此架構下,成功需要同時達成預期結果與語義對應。語義對應描述了參考影像與生成結果之間,就任務相關的高層次語義而言的對應關係。評估重點在於生成的結果,既不需要呈現完整的逐步中間任務過程,也不要求與參考影像具備傳統的外觀一致性。為了支持系統性評估,我們建構了 SemComp-Data,一個涵蓋六個領域的評估資料集。每個實例包含一張參考影像、一份詳細指令、一份簡短指令,以及一段以成果為中心的影片片段。一個可擴展的四階段篩選管線將原始影片轉換為標準化的 SemComp-Data 實例。我們進一步引入 SemComp-Bench,這是一套評估協議,利用視覺語言模型(VLM)回答結構化的二元問題。SemComp-Bench 分別報告 OA 分數與 GR 分數,對應結果達成度與生成可靠性。在具代表性的影片生成模型上的實驗顯示,要在維持參考影像中與任務相關的語義對應的同時達成預期結果,仍然具有挑戰性。
English
We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires both achievement of the intended outcome and semantic grounding. Semantic grounding characterizes the correspondence between the reference image and the generated outcome in terms of high-level semantics relevant to the task. Evaluation focuses on the generated outcome and requires neither the presentation of a complete sequence of intermediate task steps nor conventional appearance consistency with the reference image. To support systematic evaluation, we construct SemComp-Data, an evaluation dataset covering six domains. Each instance comprises a reference image, a detailed instruction, a brief instruction, and an outcome-centric video clip. A scalable four-stage curation pipeline converts raw videos into standardized SemComp-Data instances. We further introduce SemComp-Bench, an evaluation protocol that uses a vision-language model (VLM) to answer structured binary questions. SemComp-Bench reports the OA Score and the GR Score for Outcome Achievement and Generation Reliability, respectively. Experiments on representative video generation models show that achieving intended outcomes while maintaining task-relevant semantic grounding in reference images remains challenging.