ChatPaper.aiChatPaper

Mage-Flow: 이미지 생성 및 편집을 위한 효율적인 네이티브 해상도 파운데이션 모델

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

July 21, 2026
저자: Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, Zhaoyang Jia, Yifei Shen, Xun Guo, Yuxuan Luo, Jiahao Li, Wenxuan Xie, Fanyi Pu, Xiaoyi Zhang, Kaichen Zhang, Zongyu Guo, Tianci Bi, Dongnan Gui, Zhening Liu, Zimo Wen, Zihan Zheng, Senqiao Yang, Xiao Li, Jinglu Wang, Bin Li, Yan Lu
cs.AI

초록

대규모 시각 생성기는 점점 더 강력해지고 있지만, 훈련, 미세 조정 및 배포에 많은 비용이 소요됩니다. 본 논문에서는 효율적인 텍스트-이미지 생성 및 명령 기반 이미지 편집을 위한 콤팩트한 4B 규모 생성 스택인 Mage-Flow를 소개합니다. 이 스택은 두 가지 공동 설계 구성 요소, 즉 경량의 고충실도 잠재 토크나이저인 Mage-VAE와 정류 흐름 매칭(rectified flow matching)으로 훈련된 Native-Resolution Multimodal Diffusion Transformer로 구축됩니다. Mage-VAE는 앵커-잠재 정규화(anchor-latent regularization)를 사용한 단일 단계 확산 방식 인코딩 및 디코딩을 활용하여 강력한 공용 VAE의 재구성 품질을 유지하면서 토큰화 비용을 10배 이상 절감합니다. 기본 해상도 패킹 및 스택 수준 CUDA 커널 퓨전과 함께, 이 스택은 유연한 해상도 훈련을 지원하고 종단 간 훈련 처리량을 약 2.5배 향상시킵니다. 이러한 기반 위에 생성 및 편집 모두를 위한 Base, RL 정렬 및 Turbo 변형 모델군을 완전히 개발합니다. Diffusion-NFT는 프롬프트 따르기, 텍스트 렌더링, 미적 품질 및 편집 정확도를 개선하며, 적대적 지각 유도(adversarial perceptual guidance)를 통한 소수 단계 증류는 저지연 추론을 위한 4단계 Turbo 모델을 생성합니다. 콤팩트한 규모에도 불구하고, Mage-Flow와 Mage-Flow-Edit는 표준 생성 및 편집 벤치마크에서 경쟁력 있는 성능을 달성합니다. 더 중요하게는, Turbo 변형은 대화형 사용에 실용적인 고해상도 생성 및 편집을 가능하게 합니다. 단일 NVIDIA A100 GPU에서 1024² 해상도로, Mage-Flow-Turbo는 0.59초 안에 이미지를 생성하고, Mage-Flow-Edit-Turbo는 1.02초 안에 이미지를 편집하며, 작은 메모리 사용량을 유지합니다. 이러한 결과는 세심한 토크나이저-백본-시스템 공동 설계를 통해 효율적인 4B 모델군 내에서 강력한 고해상도 생성 및 편집을 제공할 수 있음을 보여줍니다.
English
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding with anchor-latent regularization, preserving the reconstruction quality of strong public VAEs while reducing tokenization cost by more than an order of magnitude. Together with native-resolution packing and stack-level CUDA kernel fusion, the stack supports flexible-resolution training and improves end-to-end training throughput by about 2.5times. Built on this foundation, we develop a complete model family with Base, RL-aligned, and Turbo variants for both generation and editing. Diffusion-NFT improves prompt following, text rendering, aesthetic quality, and editing fidelity, while few-step distillation with adversarial perceptual guidance produces 4-step Turbo models for low-latency inference. Despite its compact scale, Mage-Flow and Mage-Flow-Edit achieves competitive performance across standard generation and editing benchmarks. More importantly, the Turbo variants make high-resolution generation and editing practical for interactive use: at 1024^2 resolution on a single NVIDIA A100 GPU, Mage-Flow-Turbo generates an image in 0.59s, and Mage-Flow-Edit-Turbo edits an image in 1.02s, while maintaining a small memory footprint. These results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.