ChatPaper.aiChatPaper

缺失的时间环节:面向脚本驱动音视频生成的时间上下文路由

The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation

September 2, 2026
作者: Yichen Liu, Quanwei Zhang, Haozhe Wang, Donghao Zhou, Xiaojie Li, Yang Shi, Jiaming Liu, Ruihua Huang, Yingtian Zou, Daquan Zhou
cs.AI

摘要

联合音视频生成模型在视觉质量与音视频同步方面已取得显著进展,然而在面对镜头切换时机和对白开口时刻时,这些模型仍只能提供有限的控制能力。这一局限性制约了其在脚本驱动内容创作中的应用,因为在该场景下,时序错误会破坏叙事连贯性并影响观看体验。当前的联合生成模型将视频与音频表征对齐到共享的时间轴上,但结构化提示所指定的镜头与对白的精确时序,仅编码在提示的文本表征中,并未与任一模态的时序坐标对齐。因此,视频与音频彼此间可能保持同步,但两者却都无法遵循脚本时间线。这一不匹配促使我们将时间对齐的范围从视频与音频之间拓展至结构化脚本。为此,我们提出时间上下文路由(Temporal Context Routing, TCR),它将脚本时序映射到视频与音频生成的共享时间轴上,并将每条提示的引导路由到两种模态中的对应位置。在200条测试脚本上与基线方法相比,TCR将镜头边界平均绝对误差从1.11秒降低至0.042秒,降幅达96%;同时将对白准确率(0.5秒内命中)从28.3%提升至84.1%。TCR在实现上述提升的同时,保持了与基线相当的视觉质量和音视频同步水平。用户研究进一步表明,参与者在全部五项评估维度上均更偏好TCR生成的結果。
English
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization. However, they still provide limited control over when shot transitions occur and dialogue is spoken. This limitation constrains their application in script-driven content creation, where timing errors can undermine narrative coherence and the viewing experience. Current joint generators align video and audio representations on a shared temporal axis, yet the precise timing of shots and dialogue specified in a structured prompt is encoded only in the prompt's text representation and remains unaligned with the temporal coordinates of either modality. Consequently, video and audio may remain synchronized with each other while both fail to follow the script timeline. This mismatch motivates us to extend temporal alignment beyond video and audio to include the structured script. We therefore introduce Temporal Context Routing (TCR), which maps the script timing onto the shared temporal axis of video and audio generation and routes each prompt's guidance to the corresponding positions in both modalities. Compared with the baseline on 200 test scripts, TCR reduces Shot Boundary MAE by 96%, from 1.11 s to 0.042 s, and raises Dialogue Acc@0.5 s from 28.3% to 84.1%. TCR achieves these improvements while maintaining visual quality and audio-visual synchronization comparable to those of the baselines. A user study further shows that participants prefer TCR on all five evaluated dimensions.