RhymeFlow: 基于异步去噪流调度的免训练视频生成加速方法

摘要

基于扩散变换器（DiTs）的视频生成模型在视频合成中取得了显著性能，但由于3D注意力的二次复杂度，其推理延迟和计算成本较高。现有加速方法主要通过稀疏注意力和KV缓存等技术降低单个去噪步骤内的计算复杂度，但它们严格遵循标准扩散流程的固有约束：目标视频序列中的每一帧都必须在所有扩散时间步中经历完整的密集去噪过程。我们观察到，由于相邻帧之间内容与运动的对应关系，当锚定具有关键语义转换的关键帧时，其他帧的中间状态往往遵循更可预测的轨迹，这表明这种均匀、密集的去噪过程对于自然视频数据而言本质上是冗余的。为此，我们提出RhymeFlow——一种无训练框架，用于解耦不同帧的去噪轨迹。具体而言，我们首先识别出一组稀疏的关键帧，它们主导着潜在语义演化。随后，仅对这些关键帧进行密集、逐步的去噪以保持结构完整性，而非关键帧则逐步跳过去噪步骤以最小化计算成本。由于非关键帧跳过的中间状态破坏了关键帧去噪步骤中的时间一致性，导致视觉质量下降，我们进一步引入潜在轨迹投影模块，使关键帧能够与完整且时序一致的序列表示进行交互。在现有基于DiT的视频生成模型上的大量实验表明，我们的方法在推理速度和视觉质量上均优于现有基线。

English

Video generation models based on Diffusion Transformers (DiTs) have achieved remarkable performance in video synthesis, yet they suffer from high inference latency and computational costs due to the quadratic complexity of 3D attention. Existing acceleration methods primarily reduce computational complexity within each individual denoising steps through techniques such as sparse attention and KV-caching. However, they rigidly adhere to the inherent constraint of the standard diffusion pipeline: every frame in the target video sequence must be subjected to a complete, dense denoising process across all diffusion timesteps. We observe that due to the corresponding contents and motions among adjacent frames, when keyframes with critical semantic transitions are anchored, the intermediate states of others often follow more predictable trajectories, which indicates that such uniform, dense denoising process is inherently redundant for natural video data. To this end, we introduce RhymeFlow, a training-free framework that decouples the denoising trajectories of different frames. Specifically, we first identify a sparse set of pivotal key frames that dominate the latent semantic evolution. Then, only these keyframes undergo dense, step-by-step denoising to ensure structural integrity, while non-keyframes progressively skip denoising steps to minimize computational cost. Since skipped intermediate states of non-keyframes break the temporal coherence in keyframe denoising steps, leading to visual degradation, we further introduce a latent trajectory projection module, which enables keyframes to interact with a complete and temporally consistent sequence representation. Extensive experiments on current DiT-based video generation models demonstrate our method outperforms existing baselines with higher inference speed and better visual quality.