DramaChain Bench:短劇生成的端到端基準
DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
September 1, 2026
作者: Haoyuan Shi, Mingtao Chen, Shuo Jiang, Ziyan Chen, Xuyi Sheng, Yiming Liu, Ying Zhang, Miao Wang, Jianxiang Lu, Fanyang Lu, Songyuanyi Lu, Xiele Wu, Zhichao Hu, Yuhong Liu, Richeng Xuan
cs.AI
摘要
商業短劇製作遵循多階段鏈條:劇本、分鏡、關鍵幀圖像、鏡頭級視頻,以及最終成片短劇。多數現有基準僅使用預先撰寫的輸入而非真實上游管線輸出,來評估視頻生成階段。這使得兩個關鍵問題無法回答:每個階段是否遵從原始劇本意圖(而非僅遵從其直接輸入提示),以及不同鏡頭在組裝為多集發布後是否保持連貫性。我們提出 DramaChain Bench,首個評估完整製作鏈條中每一階段的短劇基準。它建構於三個共享同一維度系統的內部系統之上,即 DramaChain Dimensions:五個評估軸在每個階段實例化,解析為 63 個葉節點維度。DramaChain Agent 在工作流程與成片短劇品質兩方面均以商業短劇平台為校準基準,從而實現跨模型的階段級公平比較。DramaChain Labeling System 讓三位專業標注員獨立為全部 5,785 個項目評分,所有缺陷均進行時空定位,並從預先定義的缺陷清單中選取。此過程產生 17,488 個有效分數與 255,925 條可追溯的歸因記錄。人工標注證實上游缺陷會沿管線級聯傳播,表明最終劇集品質並非僅由視頻生成單獨決定。DramaChain Agentic Judge 隨後自動為每個葉節點維度評分,在對照逐項目檢查清單進行判定之前,透過多輪智能體回合收集證據;其以平均 PLCC 0.918 重現模型排序,足以在零標注成本下接納新模型。
English
Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real upstream pipeline outputs. This leaves two critical questions unanswerable: whether each stage adheres to the original script intent (rather than only its immediate input prompt), and whether disparate shots remain coherent after assembly into multi-episode releases. We present DramaChain Bench, the first short-drama benchmark that evaluates every stage of the complete production chain. It is built upon three in-house systems sharing one dimension system, DramaChain Dimensions: five evaluation axes instantiated at every stage, resolving into 63 leaf dimensions. DramaChain Agent is calibrated against commercial short-drama platforms in both workflow and finished short-drama quality, enabling stage-wise fair comparison across models. DramaChain Labeling System has each of the 5,785 items scored independently by three professional annotators, with all defects spatio-temporally localised and selected from a predefined defect list. This process produces 17,488 valid scores and 255,925 traceable attribution records. The human annotations confirm that upstream defects cascade across the pipeline, demonstrating that final episode quality is not governed by video generation alone. DramaChain Agentic Judge then scores every leaf dimension automatically, gathering evidence over multiple agentic rounds before judging against a per-item checklist; it reproduces the model ranking at a mean PLCC of 0.918, enough to admit new models at no annotation cost.