ChatPaper.aiChatPaper

FlashRender:基于相机控制视频MeanFlow的少步生成式渲染

FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow

September 3, 2026
作者: Byeongjun Park, Byung-Hoon Kim, Hyungjin Chung
cs.AI

摘要

我们提出了 FlashRender,一种少步生成式渲染框架,能够在数秒内沿目标相机轨迹对源视频进行重渲染。我们将依赖于采样步数的相机控制识别为现有多步生成式渲染模型中离散化误差的突出表现,并表明解决这一不一致性可大幅降低去噪轨迹的曲率,从而有利于后续的步蒸馏。为此,我们引入了表示转换与对齐(Representation Transformation and Alignment, RETA),该方法将源视频的隐藏表示与来自冻结的视觉几何模型的目标视频特征进行对齐。这直接将几何变换编码到源视频流中,从而实现了跨采样步数一致的相机控制。随后,我们在由 RETA 导致的低曲率去噪轨迹上,使用 MeanFlow 目标函数对模型进行微调,使模型能够更有效地处理离散化误差。最后,我们应用在策略流映射蒸馏,以修正固定少步采样下的自生成轨迹误差。大量实验表明,RETA、MeanFlow 与在策略流映射蒸馏在少步生成式渲染中发挥着互补作用。三者共同使我们的方法在视频质量和几何一致性上与多步基线方法相当,而采样成本仅为原来的 1/25,同时实现更优的相机可控性,即便面对分布外的目标相机轨迹亦是如此。
English
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target camera trajectory in seconds. We identify sampling-step-dependent camera control as a prominent manifestation of discretization error in existing multi-step generative rendering models and show that resolving this inconsistency substantially lowers denoising trajectory curvature, facilitating subsequent step distillation. To this end, we introduce Representation Transformation and Alignment (RETA), which aligns hidden source-video representations with target-video features from a frozen visual geometry model. This directly encodes the geometric transformation within the source-video stream, enabling sampling-step-consistent camera control. We then fine-tune the model with the MeanFlow objective on the lower-curvature denoising trajectory induced by RETA, allowing the model to more effectively address discretization error. Finally, we apply on-policy flow map distillation to correct self-rollout errors under fixed few-step sampling. Extensive experiments show that RETA, MeanFlow, and on-policy flow map distillation play complementary roles in few-step generative rendering. Together, they enable our approach to match multi-step baselines in video quality and geometric consistency at 25x lower sampling cost while achieving superior camera controllability, even under out-of-distribution target camera trajectories.