LeapTalk: 토킹 헤드 생성에서의 지연 시간-품질 트레이드오프 극복
LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation
July 29, 2026
저자: Rongxiang Zhang, Songhua Liu
cs.AI
초록
장기간 및 실시간 토킹헤드 생성은 지연 시간-품질 트레이드오프로 인해 여전히 어려운 문제로 남아 있다. 비효율적인 다단계 확산은 스트리밍 생성을 불가능하게 하는 반면, 실시간 자기회귀적 접근법은 오류 누적과 정체성 드리프트를 겪는다. 이러한 단점을 해결하기 위해, 우리는 단일 순방향 단계만으로 안정적인 실시간 토킹헤드 생성을 달성하고 임의로 긴 비디오로 확장 가능한 새로운 프레임워크인 LeapTalk를 제안한다. 우리 접근법의 핵심은 단일 단계 브리지 증류 기법에 있다. 한편으로, 기존의 노이즈-데이터 패러다임에서 벗어나 브라운 브리지에 기반한 데이터-데이터 수송 정식화를 도입한다. 지속적인 참조에 의해 고정된 이 전략은 정체성 드리프트를 효과적으로 완화하고 장기적인 시간적 안정성을 향상시킨다. 다른 한편으로, 사전 훈련된 확산 교사 모델에서 학생 브리지 모델로의 원활한 지식 전달을 가능하게 하기 위해, 두 모델 간의 기능적 차이를 연결하는 SNR 정렬 시간 변환 Φ(τ)을 갖춘 이종 증류 프레임워크를 탐구한다. 또한, 극단적인 단계 축소 상황에서도 정밀한 입술 동기화를 유지하기 위해 오디오 기반 분류기-프리 가이던스 메커니즘을 제안한다. 광범위한 실험을 통해 우리의 방법이 단 1단계만으로 최대 200 FPS에서 고충실도 및 시간적으로 일관된 비디오 생성을 달성하며, 효율성과 안정성 모두에서 기존 접근법을 크게 능가함을 입증한다. 프로젝트 페이지: https://zhangrongxiang.github.io/leaptalk-page/
English
Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient multi-step diffusion prohibits streaming generation, whereas real-time autoregressive approaches suffer from error accumulation and identity drift. To address this drawback, we propose LeapTalk, a novel framework that achieves stable and real-time talking-head generation with a single forward step, scaling to arbitrarily long videos. At the heart of our approach lies a single-step bridge distillation scheme. On the one hand, departing from the conventional noise-to-data paradigm, we introduce a data-to-data transport formulation based on a Brownian bridge. Anchored by a persistent reference, this strategy effectively mitigates identity drift and enhances long-term temporal stability. On the other hand, to enable smooth knowledge transfer from a pre-trained diffusion teacher to the student bridge model, we explore a heterogeneous distillation framework with an SNR-aligned time transformation Φ(τ), which bridges the functional discrepancy between the two models. Moreover, we propose an audio-driven classifier-free guidance mechanism to maintain fine-grained lip synchronization under extreme step reduction. Extensive experiments demonstrate that our method achieves high-fidelity and temporally consistent video generation with only 1 step at up to 200 FPS, significantly outperforming existing approaches in both efficiency and stability. Project Page: https://zhangrongxiang.github.io/leaptalk-page/