Mage-VL: 효율적인 코덱-네이티브 스트리밍 멀티모달 기초 모델
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
July 27, 2026
저자: Senqiao Yang, Kaichen Zhang, Zhaoyang Jia, Jinghao Guo, Yifei Shen, Xinjie Zhang, Xiaoyi Zhang, Haoqing Wang, Xiao Li, Peng Zhang, Xiang An, Yin Xie, Zhening Liu, Xun Guo, Jiahao Li, Shicheng Zheng, Jinglu Wang, Zongyu Guo, Wenxuan Xie, Zihan Zheng, Yuxuan Luo, Bin Li, Yan Lu
cs.AI
초록
표준 시각-언어 모델(VLM)은 모라베크의 역설(Moravec's paradox)을 겪는다. 즉 복잡한 오프라인 시각 추론에는 뛰어나지만 단순한 스트리밍 지각 작업에는 어려움을 겪고 비효율적으로 처리한다. 우리는 실시간 다중 모달 이해와 상호작용을 위한 효율적인 코덱 기반 스트리밍 기초 모델인 Mage-VL을 제시한다. 핵심적으로, 우리의 맞춤형 토크나이저인 Mage-ViT는 균일한 프레임 샘플링을 대체하여 희소 앵커(I) 프레임과 예측(P) 프레임 간의 움직임 벡터와 잔차 에너지를 사용하여 동적이고 엔트로피가 높은 영역을 선택적으로 인코딩한다. 16x16 패치 수준에서 작동하여 시공간적 맥락을 유지하면서 시각 토큰 소비를 75% 이상 줄인다. 약 5억 6천만 개의 레이블이 없는 이미지와 1억 개의 레이블이 없는 비디오 프레임으로 처음부터 훈련된 Mage-ViT는 수십억 개의 이미지-텍스트 쌍으로 훈련된 플래그십 인코더와 동등하거나 더 나은 성능을 보인다. 우리는 다중 모달 캡셔닝을 위한 프롬프트-코드 공동 최적화와 훈련 레시피를 안내하는 AI 기반 성능 진단을 포함하는 AI4AI 데이터 파이프라인을 구축한다. 또한, 생체 모방 이중 시스템 아키텍처(경량 System 1 이벤트 게이트와 인과적 System 2 디코더)를 통해 Mage-VL은 능동적인 스트리밍 지각을 가능하게 한다. 광범위한 평가 결과, Mage-VL-4B는 정적 작업에서 Qwen3-VL-4B와 동등한 성능을 보이면서 비디오 이해와 2D/3D 공간 추론에서 강력한 향상을 달성하고 최대 3.5배의 벽시계 추론 속도 향상을 보여주며, 15B Phi-4-reasoning-vision 베이스라인을 전반적으로 능가한다. 모델 산출물 외에도, 사전 훈련 데이터 효율성, 가변 해상도 스케일링, 코덱 시스템 가속, VideoQA SFT 중복성, 움직임-공간 시너지, AI4AI 데이터 파이프라인, 다중 모달 RL을 위한 Zero-Vision SFT를 포괄하는 7가지 주요 실증적 발견을 제시한다.
English
Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energy across sparse anchor (I) and predicted (P) frames. Operating at a 16 x 16 patch level, this reduces visual token consumption by over 75% while preserving spatiotemporal context. Trained from scratch on approximately 560M unlabeled images and 100M unlabeled video frames, Mage-ViT matches or outperforms flagship encoders trained on billions of image-text pairs. We establish AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes. Furthermore, through a bio-inspired dual-system architecture - a lightweight System 1 event gate and a causal System 2 decoder - Mage-VL enables proactive streaming perception. Extensive evaluations show that Mage-VL-4B matches Qwen3-VL-4B on static tasks while achieving strong gains in video understanding and 2D/3D spatial reasoning, with up to a 3.5x wall-clock inference speedup, and comprehensively surpasses the 15B Phi-4-reasoning-vision baseline. Beyond model artifacts, we deliver seven key empirical findings covering pre-training data efficiency, variable-resolution scaling, codec system acceleration, VideoQA SFT redundancy, motion-spatial synergy, AI4AI data pipelines, and Zero-Vision SFT for multimodal RL.