ChatPaper.aiChatPaper

CineMobile:設備端圖像到視頻擴散的電影級攝影機運動生成

CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation

July 4, 2026
作者: Xuyao Huang, Zelai Deng, Xu Wang, Xizhong Xiao, Zhijie Deng
cs.AI

摘要

行動裝置上日益增長的圖像轉影片生成需求,越來越聚焦於電影級動態效果,如子彈時間、推軌縮放、慢動作等。儘管擴散變換器(DiTs)在影片生成方面表現出色,但其龐大的參數量與多步驟迭代去噪過程導致巨大的運算開銷,使得在行動裝置上實現高效生成極具挑戰。我們提出CineMobile以填補此鴻溝。具體而言,CineMobile採用三重優化策略:(1)利用蒸餾引導剪枝方法,推導出一個精簡而高效的模型,同時保留電影效果所需的核心影片生成能力;(2)透過擴散蒸餾與強化學習的結合,將壓縮後的模型優化為4步生成器;(3)採用混合訓練後量化策略,將模型大小壓縮至1 GB以下。實驗結果顯示,相較於採用Wan 2.1架構的教師模型,CineMobile在生成速度上實現40倍加速,同時保持可比的視覺品質。具體而言,CineMobile在NVIDIA H200 GPU上生成49幀480p影片時,每步去噪延遲為0.6秒;在聯發科天璣8400 Ultimate 5G平台上,則為20秒,峰值記憶體使用量為1.8 GB,證明了其在行動裝置圖像轉影片生成中的實際應用可行性。
English
The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dolly zoom, slow motion, etc. While Diffusion Transformers (DiTs) exhibit strong performance in video generation, their large parameter sizes and multi-step iterative denoising processes lead to substantial computational overhead, making efficient generation on mobile devices challenging. We propose CineMobile to bridge the gap. In particular, CineMobile adopts a three-fold optimization strategy: (1) leveraging a distillation-guided pruning approach to derive a compact yet efficient model that retains the essential video generation capabilities required for cinematic effects; (2) optimizing the compressed model into a 4-step generator via a combination of diffusion distillation and reinforcement learning; (3) employing a hybrid post-training quantization strategy to compress the model footprint to under 1 GB. Experimental results show that compared to the teacher model with the Wan 2.1 architecture, CineMobile achieves a 40x speedup in generation while maintaining comparable visual quality. Specifically, CineMobile generates 49-frame 480p videos with a per-step denoising latency of 0.6s on an NVIDIA H200 GPU and 20s on the MediaTek Dimensity 8400 Ultimate 5G platform, with a peak memory usage of 1.8 GB, demonstrating its practical applicability for mobile-based image-to-video creation.