ChatPaper.aiChatPaper

FlashRender:カメラ制御ビデオMeanFlowによる数ステップ生成的レンダリング

FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow

September 3, 2026
著者: Byeongjun Park, Byung-Hoon Kim, Hyungjin Chung
cs.AI

要旨

我々は、ソースビデオをターゲットカメラ軌道に沿って数秒で再撮影する、少数ステップ生成レンダリングフレームワークである FlashRender を提案する。既存の多段階生成レンダリングモデルにおける離散化誤差の顕著な現れとしてサンプリングステップ数に依存するカメラ制御を特定し、この不整合を解消することでデノイジング軌道の曲率が大幅に低下し、その後のステップ蒸留が容易になることを示す。この目的のために、凍結された視覚幾何モデルからのターゲットビデオの特徴量と、ソースビデオの隠れ表現を整列させる表現変換と整列(RETA)を導入する。これにより、ソースビデオのストリーム内に幾何学的変換が直接エンコードされ、サンプリングステップ間で一貫したカメラ制御が可能になる。次に、RETA によって誘導される低曲率のデノイジング軌道上で MeanFlow 目的関数を用いてモデルをファインチューニングし、離散化誤差により効果的に対処できるようにする。最後に、固定された少数ステップサンプリング下での自己ロールアウト誤差を修正するため、オン方策フローマップ蒸留を適用する。広範な実験により、RETA、MeanFlow、オン方策フローマップ蒸留が、少数ステップ生成レンダリングにおいて補完的な役割を果たすことが示される。これらを組み合わせることで、提案手法は、分布外のターゲットカメラ軌道下でも、サンプリングコストを25分の1に抑えつつ、映像品質と幾何学的整合性において多段階ベースラインに匹敵し、優れたカメラ制御性を達成する。
English
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target camera trajectory in seconds. We identify sampling-step-dependent camera control as a prominent manifestation of discretization error in existing multi-step generative rendering models and show that resolving this inconsistency substantially lowers denoising trajectory curvature, facilitating subsequent step distillation. To this end, we introduce Representation Transformation and Alignment (RETA), which aligns hidden source-video representations with target-video features from a frozen visual geometry model. This directly encodes the geometric transformation within the source-video stream, enabling sampling-step-consistent camera control. We then fine-tune the model with the MeanFlow objective on the lower-curvature denoising trajectory induced by RETA, allowing the model to more effectively address discretization error. Finally, we apply on-policy flow map distillation to correct self-rollout errors under fixed few-step sampling. Extensive experiments show that RETA, MeanFlow, and on-policy flow map distillation play complementary roles in few-step generative rendering. Together, they enable our approach to match multi-step baselines in video quality and geometric consistency at 25x lower sampling cost while achieving superior camera controllability, even under out-of-distribution target camera trajectories.