欠落した時間的リンク:スクリプト駆動型音声・映像生成のための時間的コンテキストルーティング
The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation
September 2, 2026
著者: Yichen Liu, Quanwei Zhang, Haozhe Wang, Donghao Zhou, Xiaojie Li, Yang Shi, Jiaming Liu, Ruihua Huang, Yingtian Zou, Daquan Zhou
cs.AI
要旨
音声と映像の統合生成モデルは、映像品質と音声・映像同期の点で大きな進歩を遂げてきた。しかし、ショット遷移が発生するタイミングや台詞が発話されるタイミングを制御する余地は依然として限られている。この制約は、脚本駆動型のコンテンツ制作への応用を妨げるものであり、そこではタイミングの誤りが物語の一貫性や視聴体験を損ない得る。現在の統合生成モデルは、映像と音声の表現を共有の時間軸上で整列させるが、構造化プロンプトに指定されたショットと台詞の正確なタイミングは、プロンプトのテキスト表現にのみ符号化され、映像・音声のいずれのモダリティの時間座標とも整列されないままである。その結果、映像と音声は互いに同期したままであっても、その両方が脚本のタイムラインに従わない可能性がある。この不一致が、時間的整列を映像と音声の枠を超え、構造化脚本まで拡張する動機を与える。そこで我々は、Temporal Context Routing(TCR)を導入する。これは、脚本のタイミングを映像・音声生成の共有時間軸上に写像し、各プロンプトのガイダンスを両モダリティの対応する位置へルーティングする。200本のテスト用脚本を用いたベースラインとの比較では、TCRはショット境界MAEを96%削減し、1.11秒から0.042秒へ改善するとともに、Dialogue Acc@0.5秒を28.3%から84.1%へ引き上げる。TCRはこれらの改善を、ベースラインと同等の映像品質と音声・映像同期を維持したまま達成する。さらにユーザー調査では、参加者は評価された5つの全次元においてTCRを選好することが示された。
English
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization. However, they still provide limited control over when shot transitions occur and dialogue is spoken. This limitation constrains their application in script-driven content creation, where timing errors can undermine narrative coherence and the viewing experience. Current joint generators align video and audio representations on a shared temporal axis, yet the precise timing of shots and dialogue specified in a structured prompt is encoded only in the prompt's text representation and remains unaligned with the temporal coordinates of either modality. Consequently, video and audio may remain synchronized with each other while both fail to follow the script timeline. This mismatch motivates us to extend temporal alignment beyond video and audio to include the structured script. We therefore introduce Temporal Context Routing (TCR), which maps the script timing onto the shared temporal axis of video and audio generation and routes each prompt's guidance to the corresponding positions in both modalities. Compared with the baseline on 200 test scripts, TCR reduces Shot Boundary MAE by 96%, from 1.11 s to 0.042 s, and raises Dialogue Acc@0.5 s from 28.3% to 84.1%. TCR achieves these improvements while maintaining visual quality and audio-visual synchronization comparable to those of the baselines. A user study further shows that participants prefer TCR on all five evaluated dimensions.