LeapTalk: トーキングヘッド生成におけるレイテンシと品質のトレードオフを打破する
LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation
July 29, 2026
著者: Rongxiang Zhang, Songhua Liu
cs.AI
要旨
長時間かつリアルタイムなトーキングヘッド生成は、遅延と品質のトレードオフにより依然として困難である。すなわち、非効率的な多段階拡散はストリーミング生成を妨げる一方、リアルタイムの自己回帰的手法は誤差蓄積とアイデンティティのドリフトに悩まされる。この課題に対処するため、我々はLeapTalkを提案する。これは、単一のフォワードステップで安定かつリアルタイムなトーキングヘッド生成を実現し、任意の長さの動画へ拡張可能な新しいフレームワークである。我々のアプローチの中核は単一ステップのブリッジ蒸留スキームにある。一方で、従来のノイズからデータへのパラダイムから離れ、ブラウン橋に基づくデータからデータへの輸送定式化を導入する。永続的な参照に固定されることで、この戦略はアイデンティティのドリフトを効果的に軽減し、長期的な時間的安定性を向上させる。他方で、事前学習済みの拡散教師モデルから生徒ブリッジモデルへの円滑な知識転移を可能にするため、SNR整合型時間変換Φ(τ)を備えた異種間蒸留フレームワークを探究する。これは2つのモデル間の機能的乖離を埋める。さらに、極端なステップ数削減下でもきめ細かな口唇同期を維持するため、音声駆動の分類器フリーガイダンス機構を提案する。広範な実験により、我々の手法はわずか1ステップで最大200FPSの高忠実かつ時間的に一貫した動画生成を実現し、効率と安定性の両方において既存手法を大幅に上回ることを示す。プロジェクトページ: https://zhangrongxiang.github.io/leaptalk-page/
English
Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient multi-step diffusion prohibits streaming generation, whereas real-time autoregressive approaches suffer from error accumulation and identity drift. To address this drawback, we propose LeapTalk, a novel framework that achieves stable and real-time talking-head generation with a single forward step, scaling to arbitrarily long videos. At the heart of our approach lies a single-step bridge distillation scheme. On the one hand, departing from the conventional noise-to-data paradigm, we introduce a data-to-data transport formulation based on a Brownian bridge. Anchored by a persistent reference, this strategy effectively mitigates identity drift and enhances long-term temporal stability. On the other hand, to enable smooth knowledge transfer from a pre-trained diffusion teacher to the student bridge model, we explore a heterogeneous distillation framework with an SNR-aligned time transformation Φ(τ), which bridges the functional discrepancy between the two models. Moreover, we propose an audio-driven classifier-free guidance mechanism to maintain fine-grained lip synchronization under extreme step reduction. Extensive experiments demonstrate that our method achieves high-fidelity and temporally consistent video generation with only 1 step at up to 200 FPS, significantly outperforming existing approaches in both efficiency and stability. Project Page: https://zhangrongxiang.github.io/leaptalk-page/