Mage-Flow: 効率的なネイティブ解像度の画像生成・編集向け基盤モデル
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
July 21, 2026
著者: Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, Zhaoyang Jia, Yifei Shen, Xun Guo, Yuxuan Luo, Jiahao Li, Wenxuan Xie, Fanyi Pu, Xiaoyi Zhang, Kaichen Zhang, Zongyu Guo, Tianci Bi, Dongnan Gui, Zhening Liu, Zimo Wen, Zihan Zheng, Senqiao Yang, Xiao Li, Jinglu Wang, Bin Li, Yan Lu
cs.AI
要旨
大規模な視覚生成モデルは高性能化が進んでいるが、訓練、微調整、デプロイに多大なコストを要する。本稿では、高効率なテキスト画像生成と指示ベースの画像編集を実現するコンパクトな4B規模生成スタック「Mage-Flow」を紹介する。このスタックは、軽量で高忠実度な潜在トークナイザ「Mage-VAE」と、整流流れマッチングで訓練されたネイティブ解像度対応マルチモーダル拡散Transformerという、共設計された2つのコンポーネントで構成される。Mage-VAEは、ワンステップの拡散型符号化・復号化とアンカー潜在正則化を採用し、既存の強力な公開VAEの再構成品質を維持しつつ、トークン化コストを一桁以上削減する。さらに、ネイティブ解像度パッキングとスタックレベルのCUDAカーネル融合により、柔軟な解像度での訓練を可能にし、エンドツーエンドの訓練スループットを約2.5倍向上させる。この基盤の上に、生成と編集の両方に対応したBase、RL整合、Turboの各バリアントからなる完全なモデルファミリを開発する。Diffusion-NFTはプロンプト追従性、テキストレンダリング、美的品質、編集忠実度を向上させ、敵対的知覚ガイダンスによる少数ステップ蒸留により、低レイテンシ推論のための4ステップTurboモデルを生成する。コンパクトな規模にもかかわらず、Mage-FlowとMage-Flow-Editは標準的な生成・編集ベンチマークで競争力のある性能を達成する。さらに重要なことに、Turboバリアントは高解像度の生成と編集をインタラクティブな用途に実用的なものにする。すなわち、単一のNVIDIA A100 GPU上で1024^2解像度での動作時、Mage-Flow-Turboは0.59秒で画像を生成し、Mage-Flow-Edit-Turboは1.02秒で画像を編集し、かつメモリフットプリントも小さい。これらの結果は、トークナイザ、バックボーン、システムの注意深い共設計により、効率的な4Bモデルファミリ内で強力な高解像度生成と編集を実現できることを示している。
English
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding with anchor-latent regularization, preserving the reconstruction quality of strong public VAEs while reducing tokenization cost by more than an order of magnitude. Together with native-resolution packing and stack-level CUDA kernel fusion, the stack supports flexible-resolution training and improves end-to-end training throughput by about 2.5times. Built on this foundation, we develop a complete model family with Base, RL-aligned, and Turbo variants for both generation and editing. Diffusion-NFT improves prompt following, text rendering, aesthetic quality, and editing fidelity, while few-step distillation with adversarial perceptual guidance produces 4-step Turbo models for low-latency inference. Despite its compact scale, Mage-Flow and Mage-Flow-Edit achieves competitive performance across standard generation and editing benchmarks. More importantly, the Turbo variants make high-resolution generation and editing practical for interactive use: at 1024^2 resolution on a single NVIDIA A100 GPU, Mage-Flow-Turbo generates an image in 0.59s, and Mage-Flow-Edit-Turbo edits an image in 1.02s, while maintaining a small memory footprint. These results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.