DramaChain Bench:面向短剧生成的端到端基准测试
DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
September 1, 2026
作者: Haoyuan Shi, Mingtao Chen, Shuo Jiang, Ziyan Chen, Xuyi Sheng, Yiming Liu, Ying Zhang, Miao Wang, Jianxiang Lu, Fanyang Lu, Songyuanyi Lu, Xiele Wu, Zhichao Hu, Yuhong Liu, Richeng Xuan
cs.AI
摘要
商业短剧制作遵循多阶段链条:剧本、分镜、关键帧图像、镜头级视频,直至成片短剧。现有大多数基准仅使用预先编写的输入而非真实上游流水线输出,对视频生成单一阶段进行评估。这导致两个关键问题无法回答:每个阶段是否遵循原始剧本意图(而非仅遵循其直接输入提示),以及不同镜头在组装为多集作品后能否保持连贯性。我们提出DramaChain Bench,这是首个对完整制作链条中每一阶段进行评估的短剧基准。它基于三个共享同一维度系统的内部系统构建,即DramaChain Dimensions:在每一阶段实例化的五个评估轴,细化为63个叶子维度。DramaChain Agent在流程和成片短剧质量两方面均以商业短剧平台为校准基准,从而实现跨模型的阶段级公平比较。DramaChain Labeling System对全部5,785个项目中的每一项均由三名专业标注员独立评分,所有缺陷均进行时空定位,并从预定义的缺陷列表中选取。该流程产生17,488个有效评分和255,925条可追溯的归因记录。人工标注确认上游缺陷沿流水线级联传导,表明最终剧集质量并非仅由视频生成阶段决定。DramaChain Agentic Judge随后自动对每个叶子维度评分,在多个智能体回合中收集证据,再依据逐项检查清单进行评判;其复现模型排序的平均PLCC达0.918,足以在零标注成本下接纳新模型。
English
Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real upstream pipeline outputs. This leaves two critical questions unanswerable: whether each stage adheres to the original script intent (rather than only its immediate input prompt), and whether disparate shots remain coherent after assembly into multi-episode releases. We present DramaChain Bench, the first short-drama benchmark that evaluates every stage of the complete production chain. It is built upon three in-house systems sharing one dimension system, DramaChain Dimensions: five evaluation axes instantiated at every stage, resolving into 63 leaf dimensions. DramaChain Agent is calibrated against commercial short-drama platforms in both workflow and finished short-drama quality, enabling stage-wise fair comparison across models. DramaChain Labeling System has each of the 5,785 items scored independently by three professional annotators, with all defects spatio-temporally localised and selected from a predefined defect list. This process produces 17,488 valid scores and 255,925 traceable attribution records. The human annotations confirm that upstream defects cascade across the pipeline, demonstrating that final episode quality is not governed by video generation alone. DramaChain Agentic Judge then scores every leaf dimension automatically, gathering evidence over multiple agentic rounds before judging against a per-item checklist; it reproduces the model ranking at a mean PLCC of 0.918, enough to admit new models at no annotation cost.