VGI-Bench: 비디오 생성 모델의 시각 지능 탐구

VGI-Bench: Probing Visual Intelligence in Video Generation Models

August 26, 2026
저자: Xuan He, Cong Wei, Yuhao Cheng, Linrui Ma, Yuxuan Zhang, Zuojun Li, Yuhao Wen, Jize Jiang, Zeyi Liu, Yuren Hao, Songcheng Cai, Keming Wu, Penghui Du, Kai Zou, Rui Yang, Chenkai Sun, Ke Yang, Ping Nie, Kelsey R Allen, Chenglong Wang, Michel Galley, Jianfeng Gao, ChengXiang Zhai
cs.AI

초록

최근 연구들은 비디오 생성 모델이 생성된 프레임을 통해 특정 형태의 제로샷 시각적 추론을 보일 수 있음을 시사한다. 그러나 신뢰할 수 있는 평가는 여전히 어려운 과제이다. 벤치마크는 현재 비디오 모델의 시각적 사전 지식과 정렬된 입력을 채택해야 하며, 그럴듯한 최종 상태만이 아닌 타당한 진화 과정을 요구하고, 도전적이면서도 일부 실현 가능한 수준으로 과제 난이도를 조정해야 한다. 이를 위해, 우리는 비디오 생성 모델의 시각적 추론 능력에 대한 세분화된 평가를 위해 과제 도메인과 스킬 태그의 2단계 분류 체계로 구성된 27개 과제와 810개 인스턴스를 포함하는 VGI-bench를 소개한다. 우리의 평가는 현재 생성 시스템이 시각적으로 기반한 추론 과제의 일부를 해결할 수 있지만, 신뢰성과는 거리가 멀며, 가장 강력한 모델인 Seedance 2.0조차 우리의 평가 기준에서 51.0%에 그친다는 것을 보여준다. 우리의 분석은 출력 실패 모드, 입력 조건 민감도, 합성 파인튜닝으로부터의 성능 전이 경계, 그리고 제한된 자기 교정을 드러내는 내부 노이즈 제거 관점을 추가로 탐구한다. 즉, 후반 단계는 주로 초기 가설을 정제할 뿐 추론 오류를 교정하지 않는다. 우리는 VGI-bench가 차세대 비디오 생성 모델의 개발을 촉진하는 데 기여하기를 기대한다. 웹사이트: https://hexuan21.github.io/VGI-Bench/
English
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible. To this end, we introduce VGI-bench, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generation models. Our evaluations show that current generative systems can solve a subset of visually grounded reasoning tasks, but remain far from reliable, with even the strongest model, Seedance 2.0, achieving only 51.0% under our evaluation criteria. Our analysis further explore the output failure modes, input condition sensitivity, performance transfer boundary from synthetic fine-tuning, and internal denoising perspective revealing limited self-correction, where later steps mainly refine early hypotheses rather than correct reasoning errors. We hope VGI-bench will help stimulate the development of next-generation video generation models. Website: https://hexuan21.github.io/VGI-Bench/
PDF1431August 28, 2026