Mage-Flow:一種高效原生解析度的基礎模型,用於影像生成與編輯
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
July 21, 2026
作者: Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, Zhaoyang Jia, Yifei Shen, Xun Guo, Yuxuan Luo, Jiahao Li, Wenxuan Xie, Fanyi Pu, Xiaoyi Zhang, Kaichen Zhang, Zongyu Guo, Tianci Bi, Dongnan Gui, Zhening Liu, Zimo Wen, Zihan Zheng, Senqiao Yang, Xiao Li, Jinglu Wang, Bin Li, Yan Lu
cs.AI
摘要
大規模視覺生成模型的能力日益增強,但其訓練、微調與部署成本高昂。我們提出Mage-Flow,一個緊湊的4B規模生成堆疊,用於高效文字到影像生成與基於指令的影像編輯。該堆疊由兩個協同設計的元件構成:Mage-VAE(輕量高保真潛在分詞器),以及使用整流流匹配訓練的原生解析度多模態擴散變壓器。Mage-VAE採用一步擴散式編碼/解碼與錨點潛在正則化,在維持強大公共VAE重建品質的同時,將分詞成本降低一個數量級以上。結合原生解析度打包與堆疊級CUDA核心融合,該堆疊支援彈性解析度訓練,並將端到端訓練吞吐量提升約2.5倍。在此基礎上,我們開發了完整的模型系列,包含用於生成與編輯的Base、RL對齊與Turbo變體。擴散NFT改善了提示跟隨、文字渲染、美學品質與編輯保真度,而結合對抗性感知引導的少步蒸餾則產生了適用於低延遲推論的4步Turbo模型。儘管規模緊湊,Mage-Flow與Mage-Flow-Edit在標準生成與編輯基準測試中仍展現出競爭力。更重要的是,Turbo變體使高解析度生成與編輯可用於互動場景:在單張NVIDIA A100 GPU上,以1024²解析度,Mage-Flow-Turbo可在0.59秒內生成影像,Mage-Flow-Edit-Turbo可在1.02秒內編輯影像,同時維持較小的記憶體佔用。這些結果表明,精心的分詞器-骨幹網路-系統協同設計,能在高效的4B模型家族中實現強大的高解析度生成與編輯。
English
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding with anchor-latent regularization, preserving the reconstruction quality of strong public VAEs while reducing tokenization cost by more than an order of magnitude. Together with native-resolution packing and stack-level CUDA kernel fusion, the stack supports flexible-resolution training and improves end-to-end training throughput by about 2.5times. Built on this foundation, we develop a complete model family with Base, RL-aligned, and Turbo variants for both generation and editing. Diffusion-NFT improves prompt following, text rendering, aesthetic quality, and editing fidelity, while few-step distillation with adversarial perceptual guidance produces 4-step Turbo models for low-latency inference. Despite its compact scale, Mage-Flow and Mage-Flow-Edit achieves competitive performance across standard generation and editing benchmarks. More importantly, the Turbo variants make high-resolution generation and editing practical for interactive use: at 1024^2 resolution on a single NVIDIA A100 GPU, Mage-Flow-Turbo generates an image in 0.59s, and Mage-Flow-Edit-Turbo edits an image in 1.02s, while maintaining a small memory footprint. These results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.