ChatPaper.aiChatPaper

DramaChain Bench: 숏드라마 생성을 위한 엔드투엔드 벤치마크

DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation

September 1, 2026
저자: Haoyuan Shi, Mingtao Chen, Shuo Jiang, Ziyan Chen, Xuyi Sheng, Yiming Liu, Ying Zhang, Miao Wang, Jianxiang Lu, Fanyang Lu, Songyuanyi Lu, Xiele Wu, Zhichao Hu, Yuhong Liu, Richeng Xuan
cs.AI

초록

상용 숏드라마 제작은 대본, 스토리보드, 키프레임 이미지, 샷 단위 비디오, 완성된 숏드라마로 이어지는 다단계 체인을 따른다. 기존 대부분의 벤치마크는 사전에 작성된 입력을 사용해 비디오 생성 단계만 평가하며, 실제 업스트림 파이프라인 출력은 사용하지 않는다. 따라서 각 단계가 바로 직전 입력 프롬프트뿐 아니라 원본 대본의 의도를 준수하는지, 그리고 서로 다른 샷들이 여러 에피소드로 편집된 뒤에도 일관성을 유지하는지에 대한 두 가지 핵심 질문에는 답할 수 없다. 본 논문은 전체 제작 체인의 모든 단계를 평가하는 최초의 숏드라마 벤치마크인 DramaChain Bench를 제안한다. DramaChain Bench는 하나의 차원 체계를 공유하는 세 가지 사내 시스템을 기반으로 구축되었다. DramaChain Dimensions는 각 단계에서 구체화되는 다섯 가지 평가 축으로 구성되며, 63개의 말단 차원(leaf dimension)으로 세분된다. DramaChain Agent는 작업 흐름과 완성된 숏드라마 품질 모두에서 상용 숏드라마 플랫폼에 맞춰 보정되어, 모델 간 단계별 공정한 비교를 가능하게 한다. DramaChain Labeling System에서는 5,785개 항목 각각을 세 명의 전문 주석자가 독립적으로 채점하며, 모든 결함은 사전 정의된 결함 목록에서 선택되고 시공간적으로 위치가 특정된다. 이 과정은 17,488개의 유효 점수와 255,925개의 추적 가능한 귀속 기록을 생성한다. 인간 주석 데이터는 상위 단계(업스트림)의 결함이 파이프라인 전반으로 연쇄 전파됨을 확인하며, 최종 에피소드 품질이 비디오 생성만으로 결정되지 않음을 보여준다. 이후 DramaChain Agentic Judge는 모든 말단 차원을 자동으로 채점한다. 이 판정기는 항목별 체크리스트에 따라 판정하기 전에 여러 차례의 에이전트 기반 라운드(agentic round)를 거쳐 증거를 수집하며, 평균 PLCC 0.918로 모델 순위를 재현한다. 이는 주석 비용 없이도 새로운 모델을 수용할 수 있을 만큼 충분한 정확도이다.
English
Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real upstream pipeline outputs. This leaves two critical questions unanswerable: whether each stage adheres to the original script intent (rather than only its immediate input prompt), and whether disparate shots remain coherent after assembly into multi-episode releases. We present DramaChain Bench, the first short-drama benchmark that evaluates every stage of the complete production chain. It is built upon three in-house systems sharing one dimension system, DramaChain Dimensions: five evaluation axes instantiated at every stage, resolving into 63 leaf dimensions. DramaChain Agent is calibrated against commercial short-drama platforms in both workflow and finished short-drama quality, enabling stage-wise fair comparison across models. DramaChain Labeling System has each of the 5,785 items scored independently by three professional annotators, with all defects spatio-temporally localised and selected from a predefined defect list. This process produces 17,488 valid scores and 255,925 traceable attribution records. The human annotations confirm that upstream defects cascade across the pipeline, demonstrating that final episode quality is not governed by video generation alone. DramaChain Agentic Judge then scores every leaf dimension automatically, gathering evidence over multiple agentic rounds before judging against a per-item checklist; it reproduces the model ranking at a mean PLCC of 0.918, enough to admit new models at no annotation cost.