ChatPaper.aiChatPaper

DramaChain Bench:ショートドラマ生成のためのエンドツーエンドベンチマーク

DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation

September 1, 2026
著者: Haoyuan Shi, Mingtao Chen, Shuo Jiang, Ziyan Chen, Xuyi Sheng, Yiming Liu, Ying Zhang, Miao Wang, Jianxiang Lu, Fanyang Lu, Songyuanyi Lu, Xiele Wu, Zhichao Hu, Yuhong Liu, Richeng Xuan
cs.AI

要旨

商業ショートドラマの制作は、脚本、絵コンテ、キーフレーム画像、ショット単位のビデオ、完成したショートドラマへと至る多段階のチェーンで構成される。既存のベンチマークのほとんどは、実際の上流パイプラインの出力ではなく、事前に作られた入力を用いて、ビデオ生成段階のみを評価している。そのため、各段階が(直前の入力プロンプトのみではなく)元の脚本の意図に忠実であるかどうか、また、異なるショットが複数エピソードの作品へと組み立てられた後も一貫性を保つかどうか、という二つの重要な問いは未解決のままである。本稿では、制作チェーン全体のすべての段階を評価する初のショートドラマ・ベンチマークであるDramaChain Benchを提案する。DramaChain Benchは、単一の次元体系を共有する三つの社内システム上に構築される。その次元体系であるDramaChain Dimensionsは、各段階で具体化される五つの評価軸から成り、最終的には63のリーフ次元へと分解される。DramaChain Agentは、ワークフローと完成したショートドラマの品質の両方について、商業ショートドラマのプラットフォームに較正されており、モデル間での段階別の公平な比較を可能にする。DramaChain Labeling Systemでは、全5,785項目の各項目が三人の専門アノテーターによって独立にスコアリングされ、すべての欠陥は時空間的に局所化されるとともに、事前定義された欠陥リストから選択される。このプロセスにより、17,488件の有効スコアと255,925件の追跡可能な帰属記録が生成される。人手によるアノテーションは、上流工程の欠陥がパイプライン全体へ連鎖することを裏付けており、最終的なエピソード品質がビデオ生成のみによって決まるのではないことを示している。その後、DramaChain Agentic Judgeは、すべてのリーフ次元を自動的にスコアリングする。その際、項目別チェックリストに基づいて判定する前に、複数回のエージェント・ラウンドにわたって証拠を収集する。これにより、モデルランキングは平均PLCC 0.918で再現され、アノテーションコストをかけることなく新規モデルを受け入れることが可能になる。
English
Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real upstream pipeline outputs. This leaves two critical questions unanswerable: whether each stage adheres to the original script intent (rather than only its immediate input prompt), and whether disparate shots remain coherent after assembly into multi-episode releases. We present DramaChain Bench, the first short-drama benchmark that evaluates every stage of the complete production chain. It is built upon three in-house systems sharing one dimension system, DramaChain Dimensions: five evaluation axes instantiated at every stage, resolving into 63 leaf dimensions. DramaChain Agent is calibrated against commercial short-drama platforms in both workflow and finished short-drama quality, enabling stage-wise fair comparison across models. DramaChain Labeling System has each of the 5,785 items scored independently by three professional annotators, with all defects spatio-temporally localised and selected from a predefined defect list. This process produces 17,488 valid scores and 255,925 traceable attribution records. The human annotations confirm that upstream defects cascade across the pipeline, demonstrating that final episode quality is not governed by video generation alone. DramaChain Agentic Judge then scores every leaf dimension automatically, gathering evidence over multiple agentic rounds before judging against a per-item checklist; it reproduces the model ranking at a mean PLCC of 0.918, enough to admit new models at no annotation cost.