CineMobile: 기기 내 이미지-비디오 확산을 통한 영화적 카메라 움직임 생성
CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation
July 4, 2026
저자: Xuyao Huang, Zelai Deng, Xu Wang, Xizhong Xiao, Zhijie Deng
cs.AI
초록
모바일 기기에서 이미지-비디오 생성에 대한 수요가 증가함에 따라, bullet time, dolly zoom, 슬로우 모션 등과 같은 영화적 모션 효과에 대한 관심이 높아지고 있다. Diffusion Transformers(DiTs)는 비디오 생성에서 뛰어난 성능을 보이지만, 큰 파라미터 크기와 다단계 반복적 잡음 제거 과정으로 인해 상당한 계산 오버헤드가 발생하여 모바일 기기에서의 효율적인 생성이 어렵다. 본 논문에서는 이러한 격차를 해소하기 위해 CineMobile을 제안한다. 구체적으로, CineMobile은 세 가지 최적화 전략을 채택한다: (1) 증류 기반 가지치기(distillation-guided pruning) 방식을 활용하여 영화적 효과에 필수적인 비디오 생성 능력을 유지하는 컴팩트하면서도 효율적인 모델을 도출하고, (2) 확산 증류(diffusion distillation)와 강화 학습의 조합을 통해 압축된 모델을 4단계 생성기로 최적화하며, (3) 하이브리드 사후 훈련 양자화(post-training quantization) 전략을 적용하여 모델 크기를 1GB 미만으로 압축한다. 실험 결과, Wan 2.1 아키텍처를 기반으로 한 교사 모델과 비교하여 CineMobile은 유사한 시각적 품질을 유지하면서 생성 속도를 40배 향상시켰다. 특히, CineMobile은 NVIDIA H200 GPU에서 단계당 잡음 제거 지연 시간이 0.6초, MediaTek Dimensity 8400 Ultimate 5G 플랫폼에서 20초로 49프레임 480p 비디오를 생성하며, 최대 메모리 사용량은 1.8GB로 모바일 기반 이미지-비디오 생성에 대한 실용적 적용 가능성을 입증한다.
English
The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dolly zoom, slow motion, etc. While Diffusion Transformers (DiTs) exhibit strong performance in video generation, their large parameter sizes and multi-step iterative denoising processes lead to substantial computational overhead, making efficient generation on mobile devices challenging. We propose CineMobile to bridge the gap. In particular, CineMobile adopts a three-fold optimization strategy: (1) leveraging a distillation-guided pruning approach to derive a compact yet efficient model that retains the essential video generation capabilities required for cinematic effects; (2) optimizing the compressed model into a 4-step generator via a combination of diffusion distillation and reinforcement learning; (3) employing a hybrid post-training quantization strategy to compress the model footprint to under 1 GB. Experimental results show that compared to the teacher model with the Wan 2.1 architecture, CineMobile achieves a 40x speedup in generation while maintaining comparable visual quality. Specifically, CineMobile generates 49-frame 480p videos with a per-step denoising latency of 0.6s on an NVIDIA H200 GPU and 20s on the MediaTek Dimensity 8400 Ultimate 5G platform, with a peak memory usage of 1.8 GB, demonstrating its practical applicability for mobile-based image-to-video creation.