ChatPaper.aiChatPaper

Mage-VL: 効率的なコーデックネイティブストリーミングマルチモーダル基盤モデル

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

July 27, 2026
著者: Senqiao Yang, Kaichen Zhang, Zhaoyang Jia, Jinghao Guo, Yifei Shen, Xinjie Zhang, Xiaoyi Zhang, Haoqing Wang, Xiao Li, Peng Zhang, Xiang An, Yin Xie, Zhening Liu, Xun Guo, Jiahao Li, Shicheng Zheng, Jinglu Wang, Zongyu Guo, Wenxuan Xie, Zihan Zheng, Yuxuan Luo, Bin Li, Yan Lu
cs.AI

要旨

標準的な視覚言語モデル(VLM)はモラベックのパラドックスに悩まされている。すなわち、複雑なオフライン視覚推論には優れる一方、単純なストリーミング知覚タスクでは苦戦し、非効率的に処理する。本稿では、リアルタイムマルチモーダル理解と相互作用のための、効率的なコーデックネイティブストリーミング基盤モデルMage-VLを提案する。その中核として、カスタムトークナイザーであるMage-ViTは、均一なフレームサンプリングを置き換え、スパースなアンカー(I)フレームと予測(P)フレーム間の動きベクトルと残差エネルギーを用いて、動的でエントロピーの高い領域を選択的に符号化する。16×16のパッチレベルで動作し、これにより時空間コンテキストを保持しながら、視覚トークン消費量を75%以上削減する。約5億6000万枚のラベルなし画像と1億フレームのラベルなし映像でゼロから訓練されたMage-ViTは、数十億の画像テキストペアで訓練された旗艦エンコーダに匹敵、または凌駕する。さらに、AI4AIデータパイプラインとして、マルチモーメントキャプションのためのプロンプト-コード共同最適化と、訓練レシピを導くAI駆動性能診断を確立する。また、生体に着想を得た二重システムアーキテクチャ(軽量なシステム1イベントゲートと因果的システム2デコーダ)により、Mage-VLはプロアクティブなストリーミング知覚を実現する。広範な評価により、Mage-VL-4Bは静的タスクでQwen3-VL-4Bと同等でありながら、映像理解および2D/3D空間推論で大きな向上を示し、最大3.5倍の実時間推論高速化を達成し、15BのPhi-4-reasoning-visionベースラインを総合的に凌駕する。モデル成果物に加え、事前学習データ効率、可変解像度スケーリング、コーデックシステム高速化、VideoQA SFT冗長性、動作-空間相乗効果、AI4AIデータパイプライン、マルチモーダルRLのためのZero-Vision SFTという7つの重要な実証的知見を提供する。
English
Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor (I) and predicted (P) frames. Operating at a 16 x 16 patch level, this reduces visual token consumption by over 75% while preserving spatiotemporal context. Trained from scratch on approximately 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or outperforms flagship encoders trained on billions of image-text pairs. We establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes. Furthermore, through a bio-inspired dual-system architecture - a lightweight System 1 event gate and a causal System 2 decoder - Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5x wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline. Beyond model artifacts, we deliver seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, VideoQA SFT redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL.