VideoChat3: 효율적이고 범용적인 비디오 이해를 위한 완전 개방형 비디오 MLLM

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

July 16, 2026
저자: Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang
cs.AI

초록

최근 비디오 이해 분야의 발전은 움직임, 장편 비디오 및 실시간 스트리밍 상호작용을 포괄하며, 이 분야를 실제 응용으로 이끌고 있습니다. 이러한 진전에도 불구하고, 현재의 오픈소스 모델은 여러 측면에서 제한적입니다. 이들은 종종 다양한 비디오 유형에 걸쳐 일반화하는 데 어려움을 겪어 특정 도메인에서만 효과적입니다. 높은 계산 요구 사항은 효율성과 확장성을 더욱 제한합니다. 게다가 대부분의 모델은 부분적으로만 공개되어 있으며, 학습 코드, 전략 또는 데이터셋과 같은 핵심 구성 요소가 제공되지 않아 재현성을 저해하고 커뮤니티 주도 개발을 늦춥니다. 이러한 문제를 해결하기 위해, 우리는 완전히 공개되고 효율적이며 일반적인 비디오 중심 MLLM인 VideoChat3를 소개합니다. VideoChat3는 두 가지 보완적인 설계를 통해 비디오 이해를 발전시킵니다. 효율성 측면에서는 팽창형 3D 비전 트랜스포머(I3D-ViT)와 적응형 프레임 해상도를 위한 스트리밍 비디오 인식을 도입하여 효율적인 시공간 표현을 가능하게 하고 학습 및 추론 중 비디오 입력 처리 비용을 절감합니다. 효과성 측면에서는 확장 가능한 비디오 데이터 합성 파이프라인을 개발하여 세 가지 다양하고 고품질의 학습 데이터셋(VideoChat3-Academic2M, VideoChat3-LV116K, VideoChat3-OL617K)을 선별합니다. 이 데이터셋은 일반, 장편, 스트리밍 비디오 시나리오를 포괄하여 모델의 도메인 간 일반화 능력을 향상시킵니다. 이러한 설계를 통합함으로써, VideoChat3는 광범위한 일반화와 계산 효율성 사이의 드문 균형을 달성합니다. 일반, 장편, 스트리밍 벤치마크에 걸친 실험 결과, VideoChat3는 4B 파라미터와 더 높은 효율성으로 동일하거나 더 많은 파라미터를 가진 기존 오픈소스 모델을 능가하는 것으로 나타났습니다.
English
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse video types, making them effective only in specific domains. High computational demands further restrict their efficiency and scalability. Moreover, most models are only partially open, with key components such as training code, strategy, or datasets unavailable, which hinders reproducibility and slows community-driven development. To address these issues, we introduce VideoChat3, a fully open, efficient, and generalist video-centric MLLM. VideoChat3 advances video understanding through two complementary designs. For efficiency, we introduce Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for Streaming Video Perception, which enables efficient spatiotemporal representation and reduces the cost of processing video inputs during training and inference. For effectiveness, we develop a scalable video data synthesis pipeline that curates three diverse, high-quality training datasets: VideoChat3-Academic2M, VideoChat3-LV116K, and VideoChat3-OL617K, covering general, long-form, and streaming video scenarios, improving the model's generalization across domains. By integrating these designs, VideoChat3 achieves a rare balance of broad generalization and computational efficiency. Experiments across general, long-form, and streaming benchmarks demonstrate that VideoChat3 surpasses prior open-source models with equal or larger parameter counts with only 4B parameters and higher efficiency.
PDF1081July 18, 2026