키메라: 하이브리드 비주얼 확산 트랜스포머의 설계 및 친칠라 스케일링
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
July 30, 2026
저자: Chongjian Ge, Hanwen Jiang, Tianyu Wang, Jiuxiang Gu, Yiran Xu, Ziwen Chen, Shaoteng Liu, Jing Shi, Yicong Hong, Zefan Cai, Hailin Jin, Hao Tan
cs.AI
초록
시각 생성은 점점 더 고해상도 이미지, 긴 비디오, 멀티모달 컨텍스트를 요구하며, 이에 따라 전체 어텐션의 2차 비용은 감당하기 어려워진다. 우리는 원리 기반 확장 방식을 갖춘 하이브리드 시각 확산 백본인 Chimera를 소개한다. Chimera는 위치 임베딩 없이 텍스트, 이미지, 비디오 토큰을 하나의 래스터 순서 스트림으로 처리한다. 이는 O(N) 복잡도의 장기 컨텍스트 상태 추적을 위한 Kimi Delta Attention(KDA), 직접적인 전역 상호작용을 위한 인터리브된 멀티헤드 잠재 어텐션(MLA), 국소적 시공간 컨텍스트를 위한 모달리티 인식 단기 합성곱을 결합한다. 희소 전문가 혼합(MoE) 계층은 활성화 연산량을 제어하면서 용량을 확장한다. 이 이종 아키텍처를 확장하기 위해, 우리는 각 텐서의 기능적 팬인(fan-in)과 모델 깊이에 따라 폭과 깊이에 걸쳐 하이퍼파라미터를 전이하는 모듈별 기법인 HeteroP를 도입한다. HeteroP는 활성화 모델 크기, 학습 토큰 수, 이미지-비디오 데이터 비율에 대한 Chinchilla 스타일의 계산 최적 법칙을 피팅하는 데 사용되는 일관되게 튜닝된 모델 패밀리를 생성한다. 이러한 법칙에 따라 우리는 20억 개의 활성화 파라미터를 가진 110억 파라미터 규모의 Chimera를 학습한다. 실험은 세 가지 결과를 보여준다. 첫째, 사전 학습 확산 손실로 측정했을 때, 밀집 백본은 대응되는 전체 어텐션 Wan-2.1 2B 기준 모델보다 계산 효율이 1.7배 높으며, 전체 시스템은 7.3배에 도달한다. 둘째, 길이별 미세 조정 없이 Chimera는 5초 학습 클립에서 30초 비디오로 제로샷 외삽을 수행하며, 마지막 5초에서 단 6.5%의 FID 저하만을 보인다. 셋째, 피팅된 법칙은 계산 최적 이미지 사전 학습이 계산량을 활성화 모델 크기와 학습 토큰 수 사이에 거의 균등하게 분배하는 반면, 비디오 사전 학습은 더 높은 예산에서 모델 크기를 다소 선호함을 보여준다. 이러한 결과는 효율적인 장기 컨텍스트 확산 아키텍처를 설계하고 확장하기 위한 기반을 마련한다.
English
Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera, a hybrid visual diffusion backbone with a principled scaling recipe. Chimera processes text, image, and video tokens in one raster-ordered stream without positional embeddings. It combines Kimi Delta Attention (KDA) for long-context state tracking with O(N) complexity, interleaved Multi-head Latent Attention (MLA) for direct global interaction, and modality-aware short convolutions for local spatiotemporal context. Sparse Mixture-of-Experts (MoE) layers expand capacity while controlling activated compute. To scale this heterogeneous architecture, we introduce HeteroP, a module-wise scheme that transfers hyperparameters across width and depth according to each tensor's functional fan-in and model depth. HeteroP yields a consistently tuned family used to fit Chinchilla-style compute-optimal laws for activated model size, training-token count, and image-video data ratio. Guided by these laws, we train an 11B-parameter Chimera with 2B activated parameters. Experiments show three results. First, measured by pretraining diffusion loss, the dense backbone is 1.7x as compute-efficient as a matched full-attention Wan-2.1 2B baseline, while the complete system reaches 7.3x. Second, without length-specific fine-tuning, Chimera extrapolates zero-shot from 5-second training clips to 30-second videos, with only 6.5% FID degradation in the last five seconds. Third, the fitted laws show that compute-optimal image pretraining divides compute nearly evenly between activated model size and training-token count, whereas video pretraining modestly favors model size at higher budgets. These results establish a foundation for designing and scaling efficient long-context diffusion architectures.