VideoChat3:効率的かつ汎用的な動画理解を実現する完全オープンな動画マルチモーダル大規模言語モデル

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

July 16, 2026
著者: Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang
cs.AI

要旨

近年の動画理解の進歩は、動作、長時間動画、ストリーミングインタラクションに及び、この分野を実世界のアプリケーションへと押し進めています。この進歩にもかかわらず、現在のオープンソースモデルにはいくつかの点で限界があります。多様な動画タイプにわたって汎化することが難しく、特定の領域でのみ有効であることが多いです。高い計算負荷が効率性とスケーラビリティをさらに制限しています。さらに、ほとんどのモデルは部分的にしかオープンではなく、トレーニングコード、戦略、データセットなどの主要コンポーネントが利用できないため、再現性が損なわれ、コミュニティ主導の開発が遅れています。 これらの問題に対処するため、完全にオープンで効率的かつ汎用的な動画中心のMLLMであるVideoChat3を導入します。VideoChat3は、2つの相補的な設計を通じて動画理解を進化させます。効率性のために、膨張型3Dビジョントランスフォーマー(I3D-ViT)とストリーミング動画認識のための適応フレーム解像度を導入し、効率的な時空間表現を可能にし、トレーニングおよび推論中の動画入力処理のコストを削減します。効果性のために、スケーラブルな動画データ合成パイプラインを開発し、汎用、長時間形式、ストリーミング動画シナリオをカバーする3つの多様で高品質なトレーニングデータセット(VideoChat3-Academic2M、VideoChat3-LV116K、VideoChat3-OL617K)をキュレーションし、モデルのドメイン間での汎化を向上させます。これらの設計を統合することで、VideoChat3は幅広い汎化と計算効率の稀なバランスを達成します。汎用、長時間形式、ストリーミングの各ベンチマークにわたる実験は、VideoChat3が同等またはそれ以上のパラメータ数を持つ従来のオープンソースモデルを、わずか4Bパラメータでより高い効率性をもって凌駕することを示しています。
English
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse video types, making them effective only in specific domains. High computational demands further restrict their efficiency and scalability. Moreover, most models are only partially open, with key components such as training code, strategy, or datasets unavailable, which hinders reproducibility and slows community-driven development. To address these issues, we introduce VideoChat3, a fully open, efficient, and generalist video-centric MLLM. VideoChat3 advances video understanding through two complementary designs. For efficiency, we introduce Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for Streaming Video Perception, which enables efficient spatiotemporal representation and reduces the cost of processing video inputs during training and inference. For effectiveness, we develop a scalable video data synthesis pipeline that curates three diverse, high-quality training datasets: VideoChat3-Academic2M, VideoChat3-LV116K, and VideoChat3-OL617K, covering general, long-form, and streaming video scenarios, improving the model's generalization across domains. By integrating these designs, VideoChat3 achieves a rare balance of broad generalization and computational efficiency. Experiments across general, long-form, and streaming benchmarks demonstrate that VideoChat3 surpasses prior open-source models with equal or larger parameter counts with only 4B parameters and higher efficiency.
PDF1081July 18, 2026