ChatPaper.aiChatPaper

누락된 시간적 연결고리: 스크립트 기반 오디오-비디오 생성을 위한 시간적 맥락 라우팅

The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation

September 2, 2026
저자: Yichen Liu, Quanwei Zhang, Haozhe Wang, Donghao Zhou, Xiaojie Li, Yang Shi, Jiaming Liu, Ruihua Huang, Yingtian Zou, Daquan Zhou
cs.AI

초록

오디오-비디오 공동 생성 모델은 시각 품질과 오디오-비디오 동기화 측면에서 상당한 진전을 이루었다. 그러나 이러한 모델들은 샷 전환이 발생하는 시점과 대사가 발화되는 시점에 대한 제어는 여전히 제한적이다. 이러한 제약은 시간 오류가 서사적 일관성과 시청 경험을 저해할 수 있는 대본 기반 콘텐츠 생성에의 적용을 제한한다. 현재의 공동 생성기들은 비디오와 오디오 표현을 공유된 시간 축 위에 정렬하지만, 구조화된 프롬프트에 명시된 샷과 대사의 정밀한 타이밍은 프롬프트의 텍스트 표현에만 인코딩될 뿐 두 모달리티 중 어느 것의 시간 좌표와도 정렬되지 않은 상태로 남는다. 그 결과 비디오와 오디오는 서로 동기화된 상태를 유지할 수 있지만, 둘 다 대본 타임라인을 따르지 못할 수 있다. 이러한 불일치는 시간적 정렬을 비디오와 오디오를 넘어 구조화된 대본까지 포함하도록 확장해야 할 동기를 제공한다. 이에 우리는 TCR(Temporal Context Routing)을 제안한다. TCR은 대본 타이밍을 비디오 및 오디오 생성의 공유 시간 축에 매핑하고, 각 프롬프트의 안내 정보를 두 모달리티의 해당 위치로 라우팅한다. 200개의 테스트 대본에 대해 기준 모델과 비교했을 때, TCR은 샷 경계 MAE를 1.11초에서 0.042초로 96% 감소시키고 대사 Acc@0.5s를 28.3%에서 84.1%로 향상시킨다. TCR은 이러한 개선을 이루면서도 기준 모델과 유사한 수준의 시각 품질과 오디오-비디오 동기화를 유지한다. 사용자 연구는 또한 참가자들이 평가된 다섯 가지 차원 모두에서 TCR을 선호함을 보여준다.
English
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization. However, they still provide limited control over when shot transitions occur and dialogue is spoken. This limitation constrains their application in script-driven content creation, where timing errors can undermine narrative coherence and the viewing experience. Current joint generators align video and audio representations on a shared temporal axis, yet the precise timing of shots and dialogue specified in a structured prompt is encoded only in the prompt's text representation and remains unaligned with the temporal coordinates of either modality. Consequently, video and audio may remain synchronized with each other while both fail to follow the script timeline. This mismatch motivates us to extend temporal alignment beyond video and audio to include the structured script. We therefore introduce Temporal Context Routing (TCR), which maps the script timing onto the shared temporal axis of video and audio generation and routes each prompt's guidance to the corresponding positions in both modalities. Compared with the baseline on 200 test scripts, TCR reduces Shot Boundary MAE by 96%, from 1.11 s to 0.042 s, and raises Dialogue Acc@0.5 s from 28.3% to 84.1%. TCR achieves these improvements while maintaining visual quality and audio-visual synchronization comparable to those of the baselines. A user study further shows that participants prefer TCR on all five evaluated dimensions.