Mage-VL: 一种高效的编解码器原生流式多模态基础模型
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
July 27, 2026
作者: Senqiao Yang, Kaichen Zhang, Zhaoyang Jia, Jinghao Guo, Yifei Shen, Xinjie Zhang, Xiaoyi Zhang, Haoqing Wang, Xiao Li, Peng Zhang, Xiang An, Yin Xie, Zhening Liu, Xun Guo, Jiahao Li, Shicheng Zheng, Jinglu Wang, Zongyu Guo, Wenxuan Xie, Zihan Zheng, Yuxuan Luo, Bin Li, Yan Lu
cs.AI
摘要
标准视觉-语言模型(VLM)存在莫拉维克悖论:它们擅长复杂的离线视觉推理,却在简单的流式感知任务中表现不佳且处理效率低下。我们提出Mage-VL,一种高效的编解码原生流式基础模型,用于实时多模态理解与交互。其核心在于我们定制的分词器Mage-ViT,通过利用运动向量和残差能量选择性编码动态且熵值丰富的区域(稀疏锚定帧I与预测帧P),取代了均匀帧采样。在16×16块级操作下,视觉令牌消耗减少超过75%,同时保留了时空上下文。Mage-ViT基于约5.6亿未标注图像和1亿未标注视频帧从头训练,性能可媲美或超越在数十亿图像-文本对上训练的主流编码器。我们构建了涵盖提示-代码联合优化的多模态字幕生成AI4AI数据流水线,并引入基于AI的诊断指导训练策略。此外,通过仿生双系统架构——轻量级系统1事件门控与因果系统2解码器——Mage-VL实现了主动流式感知。大量评估表明,Mage-VL-4B在静态任务上与Qwen3-VL-4B持平,在视频理解及2D/3D空间推理上取得显著提升,推理速度最高提升3.5倍,全面超越150亿参数的Phi-4-reasoning-vision基线。除模型成果外,我们贡献了七项关键实证发现,涉及预训练数据效率、变分辨率缩放、编解码系统加速、VideoQA SFT冗余、运动-空间协同、AI4AI数据流水线,以及用于多模态强化学习的零视觉监督微调。
English
Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor (I) and predicted (P) frames. Operating at a 16 x 16 patch level, this reduces visual token consumption by over 75% while preserving spatiotemporal context. Trained from scratch on approximately 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or outperforms flagship encoders trained on billions of image-text pairs. We establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes. Furthermore, through a bio-inspired dual-system architecture - a lightweight System 1 event gate and a causal System 2 decoder - Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5x wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline. Beyond model artifacts, we deliver seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, VideoQA SFT redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL.