VideoChat3:完全開放的多模態大語言模型,實現高效且通用的視頻理解
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
July 16, 2026
作者: Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang
cs.AI
摘要
近期在视频理解方面的进展已涵蓋動作、長視頻及串流互動,推動此領域朝向實際應用發展。然而,儘管有此進展,當前開源模型仍存在若干局限。它們往往難以在不同類型的視頻間泛化,僅在特定領域有效;高運算需求進一步限制了其效率與可擴展性。此外,多數模型僅部分開放,訓練代碼、策略或數據集等關鍵元件無法取得,阻礙了可重現性並減緩社群驅動的開發進程。為解決這些問題,我們提出 VideoChat3——一個完全開放、高效且通用的以視頻為中心的 MLLM。VideoChat3 透過兩種互補設計推進視頻理解。在效率方面,我們引入了膨脹3D視覺 Transformer(I3D-ViT)與自適應幀解析度串流視頻感知,實現高效的時空表徵,並降低訓練與推理過程中處理視頻輸入的成本。在效能方面,我們開發了一個可擴展的視頻數據合成流程,策劃了三個多樣化、高品質的訓練數據集:VideoChat3-Academic2M、VideoChat3-LV116K 與 VideoChat3-OL617K,涵蓋通用、長格式與串流視頻場景,從而提升模型跨領域的泛化能力。透過整合這些設計,VideoChat3 實現了廣泛泛化與運算效率的罕見平衡。在通用、長格式與串流基準測試上的實驗表明,VideoChat3 以僅 4B 參數與更高效率,超越了參數量相等或更大的既有開源模型。
English
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse video types, making them effective only in specific domains. High computational demands further restrict their efficiency and scalability. Moreover, most models are only partially open, with key components such as training code, strategy, or datasets unavailable, which hinders reproducibility and slows community-driven development. To address these issues, we introduce VideoChat3, a fully open, efficient, and generalist video-centric MLLM. VideoChat3 advances video understanding through two complementary designs. For efficiency, we introduce Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for Streaming Video Perception, which enables efficient spatiotemporal representation and reduces the cost of processing video inputs during training and inference. For effectiveness, we develop a scalable video data synthesis pipeline that curates three diverse, high-quality training datasets: VideoChat3-Academic2M, VideoChat3-LV116K, and VideoChat3-OL617K, covering general, long-form, and streaming video scenarios, improving the model's generalization across domains. By integrating these designs, VideoChat3 achieves a rare balance of broad generalization and computational efficiency. Experiments across general, long-form, and streaming benchmarks demonstrate that VideoChat3 surpasses prior open-source models with equal or larger parameter counts with only 4B parameters and higher efficiency.