The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation

Yichen Liu, Quanwei Zhang, Haozhe Wang, Donghao Zhou, Xiaojie Li, Yang Shi, Jiaming Liu, Ruihua Huang, Yingtian Zou, Daquan Zhou

arXiv:2609.02367 · 2026-09-04 공개 · arXiv · PDF

multimodal-generation audio-video-generation temporal-alignment shot-boundary dialogue-timing prompt-guidance temporal-context-routing script-driven

Abstract

Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization. However, they still provide limited control over when shot transitions occur and dialogue is spoken. This limitation constrains their application in script-driven content creation, where timing errors can undermine narrative coherence and the viewing experience. Current joint generators align video and audio representations on a shared temporal axis, yet the precise timing of shots and dialogue specified in a structured prompt is encoded only in the prompt's text representation and remains unaligned with the temporal coordinates of either modality. Consequently, video and audio may remain synchronized with each other while both fail to follow the script timeline. This mismatch motivates us to extend temporal alignment beyond video and audio to include the structured script. We therefore introduce Temporal Context Routing (TCR), which maps the script timing onto the shared temporal axis of video and audio generation and routes each prompt's guidance to the corresponding positions in both modalities. Compared with the baseline on 200 test scripts, TCR reduces Shot Boundary MAE by 96%, from 1.11 s to 0.042 s, and raises Dialogue [email protected] s from 28.3% to 84.1%. TCR achieves these improvements while maintaining visual quality and audio-visual synchronization comparable to those of the baselines. A user study further shows that participants prefer TCR on all five evaluated dimensions.

한국어 요약

한 줄 요약

TCR은 스크립트 기반 동영상 생성에서 비디오와 오디오의 타이밍을 스크립트와 정확히 맞추는 기술로, Shot Boundary MAE 96% 감소, Dialogue [email protected] 84.1% 달성.

핵심 기여도

핵심 아이디어

기존의 비디오-오디오 생성 모델은 공유 시간 축에서 동작하지만, 스크립트에 명시된 씬 전환과 대사 타이밍은 텍스트 조건에만 포함되어 있어, 실제 생성물과 스크립트 사이의 시간 차이가 발생했다. 이 문제를 해결하기 위해 TCR은 각 텍스트 프롬프트의 시간 정보를 **공유 시간 축**에 매핑하고, 해당 시간 위치에 맞춰 **비디오-텍스트 및 오디오-텍스트 크로스 어텐션 로짓에 라우팅 점수를 추가**하는 방식을 도입했다. 이는 프롬프트별로 독립적으로 시간 제어가 가능하도록 하며, 기존 텍스트, 쿼리, 키 표현을 수정하지 않아 시각 품질과 오디오-비디오 동기화를 유지할 수 있다.

기술적 접근법

주요 결과

의의 및 한계

TCR은 스크립트 기반 생성에서 **정확한 시간 제어**를 가능하게 하며, 기존 시각 품질과 오디오-비디오 동기화를 유지하면서도 스크립트와의 일치도를 높인다. 이는 드라마, 광고 등 구조화된 콘텐츠 제작에 큰 기여를 할 수 있다. 그러나 TCR은 **정밀한 어노테이션 데이터**가 필요하며, **코어스 타이밍 정보**를 사용하면 성능이 크게 저하되는 한계가 있다. 또한, 현재는 **0.1초 단위의 격자 기반 타이밍**만 지원하며, 더 세밀한 제어는 추가 연구가 필요하다.

실용적 활용

TCR은 **쇼트폼 드라마, 광고, 인터랙티브 콘텐츠 제작** 등 스크립트 기반 동영상 생성이 필요한 산업에 적용 가능하다. 특히, 대사와 씬 전환의 정확한 타이밍이 요구되는 **드라마 제작, 캐릭터 애니메이션, 게임 내 콘텐츠 생성** 분야에서 유용하게 사용될 수 있다.