ChatPaper.aiChatPaper

RefCaptioner: 다중 참조 이미지 기반 비디오 캡셔닝

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

July 30, 2026
저자: Tengfei Liu, Yang Shi, Yuran Wang, Xiaohan Zhang, Yuqing Wen, Yuqi Tang, Qixun Wang, Zhuoran Zhang, Xuanyu Zhu, Weihong Lin, Xinlei Yu, Yujie Wei, Xinwei Long, Fengxiang Wang, Xinlong Chen, Yue Ding, Jialu Chen, Haotian Wang, Yuanxing Zhang
cs.AI

초록

기존 비디오 캡셔닝 모델은 비디오 콘텐츠의 자연스러운 설명을 생성하지만, 국부적인 시각 요소를 다중 참조 이미지에 명시적으로 그라운딩하지 못한다. 우리는 구 수준의 참조 그라운딩을 갖춘 사실적인 비디오 설명을 요구하는 새로운 태스크인 다중 참조 이미지 기반 비디오 캡셔닝(multi-reference image-grounded video captioning)을 소개하고, 이 태스크를 위한 2단계 사후 학습 프레임워크인 RefCaptioner를 제안한다. RefCaptioner는 혼합 데이터 SFT와 계층적 커버리지 할인 GRPO(Hierarchical Coverage-Discounted GRPO)를 결합하여, 일반적인 비디오 캡셔닝 능력을 유지하면서 참조 선택, 구 수준 바인딩, 디스트랙터 배제, 그리고 참조 간 일관성을 동시에 개선한다. 학습을 지원하기 위해, 우리는 20,000개의 비디오와 171,354개의 참조 이미지를 포함하는 코퍼스를 구축한다. 또한 우리는 실제 세계 비디오와 AI 생성 비디오 모두에서 캡션 사실성과 다중 참조 그라운딩을 평가하기 위한 벤치마크인 MRVBench를 소개한다. 실험 결과, RefCaptioner는 표준 비디오 캡셔닝 벤치마크에서 경쟁력을 유지하면서 오픈소스 모델 중 최고의 전반적 성능을 달성한다. 인간 평가는 또한 RefCaptioner의 캡션이 주석자에 의해 선호되며, 오픈소스 및 상용 비디오 생성기 모두에서 소스에 보다 충실한 비디오 재구성을 가능하게 함을 확인한다.
English
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing 20,000 videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.