VideoChat3:面向高效通用视频理解的全开放视频多模态大语言模型

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

July 16, 2026
作者: Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang
cs.AI

摘要

视频理解领域的最新进展涵盖了运动、长视频和流式交互,推动该领域走向实际应用。尽管取得了这些进步,当前的开源模型仍存在若干局限。它们往往难以泛化到多种类型的视频,仅在特定领域有效;高计算需求进一步限制了其效率与可扩展性。此外,大多数模型仅部分开放,关键组件如训练代码、策略或数据集不可用,阻碍了可复现性和社区驱动的开发。为解决这些问题,我们提出VideoChat3——一个完全开放、高效且通用的视频中心多模态大语言模型。VideoChat3通过两个互补设计推进视频理解。在效率方面,我们引入膨胀3D视觉Transformer(I3D-ViT)和面向流式视频感知的自适应帧分辨率,实现了高效的时空表示,并降低了训练和推理过程中处理视频输入的成本。在效果方面,我们开发了可扩展的视频数据合成流程,构建了三个多样化、高质量的训练数据集:VideoChat3-Academic2M、VideoChat3-LV116K和VideoChat3-OL617K,覆盖通用、长视频和流式视频场景,提升了模型跨领域的泛化能力。通过整合这些设计,VideoChat3实现了广泛泛化与计算效率的罕见平衡。在通用、长视频和流式基准测试上的实验表明,VideoChat3仅凭40亿参数和更高效率,超越了同等或更大参数量的先前开源模型。
English
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse video types, making them effective only in specific domains. High computational demands further restrict their efficiency and scalability. Moreover, most models are only partially open, with key components such as training code, strategy, or datasets unavailable, which hinders reproducibility and slows community-driven development. To address these issues, we introduce VideoChat3, a fully open, efficient, and generalist video-centric MLLM. VideoChat3 advances video understanding through two complementary designs. For efficiency, we introduce Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for Streaming Video Perception, which enables efficient spatiotemporal representation and reduces the cost of processing video inputs during training and inference. For effectiveness, we develop a scalable video data synthesis pipeline that curates three diverse, high-quality training datasets: VideoChat3-Academic2M, VideoChat3-LV116K, and VideoChat3-OL617K, covering general, long-form, and streaming video scenarios, improving the model's generalization across domains. By integrating these designs, VideoChat3 achieves a rare balance of broad generalization and computational efficiency. Experiments across general, long-form, and streaming benchmarks demonstrate that VideoChat3 surpasses prior open-source models with equal or larger parameter counts with only 4B parameters and higher efficiency.
PDF1081July 18, 2026