ChatPaper.aiChatPaper

SemComp-Bench: 비디오 생성에서의 의미론적 작업 완료 벤치마킹

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

August 18, 2026
저자: Keyu Tu, Zhuowei Chen, Mengqi Huang, Yuxin Wang, Jiahao Zhu, Zhendong Mao, Yongdong Zhang
cs.AI

초록

우리는 결과 중심의 비디오 생성 과업인 의미론적 과업 완료 비디오 생성(Semantic Task Completion Video Generation)을 소개한다. 이 과업 설정에서 성공을 위해서는 의도된 결과의 달성과 의미론적 근거 확보가 모두 요구된다. 의미론적 근거는 참조 이미지와 생성된 결과 간의 대응 관계를 과업과 관련된 고수준 의미론 측면에서 규명한다. 평가는 생성된 결과에 초점을 맞추며, 중간 과업 단계의 완전한 시퀀스 제시나 참조 이미지와의 통상적인 외관 일관성은 요구하지 않는다. 체계적인 평가를 지원하기 위해 우리는 여섯 개의 도메인을 포괄하는 평가 데이터셋인 SemComp-Data를 구축한다. 각 인스턴스는 참조 이미지, 상세 지시문, 간결 지시문, 그리고 결과 중심 비디오 클립으로 구성된다. 확장 가능한 4단계 큐레이션 파이프라인은 원시 비디오를 표준화된 SemComp-Data 인스턴스로 변환한다. 또한 우리는 비전-언어 모델(VLM)을 사용하여 구조화된 이분형 질문에 응답하는 평가 프로토콜인 SemComp-Bench를 도입한다. SemComp-Bench는 결과 달성(Outcome Achievement)과 생성 신뢰성(Generation Reliability)에 대해 각각 OA 점수와 GR 점수를 보고한다. 대표적인 비디오 생성 모델들에 대한 실험 결과, 의도된 결과를 달성하면서 참조 이미지에서 과업 관련 의미론적 근거를 유지하는 것은 여전히 어려운 과제로 남아 있음이 확인된다.
English
We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires both achievement of the intended outcome and semantic grounding. Semantic grounding characterizes the correspondence between the reference image and the generated outcome in terms of high-level semantics relevant to the task. Evaluation focuses on the generated outcome and requires neither the presentation of a complete sequence of intermediate task steps nor conventional appearance consistency with the reference image. To support systematic evaluation, we construct SemComp-Data, an evaluation dataset covering six domains. Each instance comprises a reference image, a detailed instruction, a brief instruction, and an outcome-centric video clip. A scalable four-stage curation pipeline converts raw videos into standardized SemComp-Data instances. We further introduce SemComp-Bench, an evaluation protocol that uses a vision-language model (VLM) to answer structured binary questions. SemComp-Bench reports the OA Score and the GR Score for Outcome Achievement and Generation Reliability, respectively. Experiments on representative video generation models show that achieving intended outcomes while maintaining task-relevant semantic grounding in reference images remains challenging.