RefCaptioner:多参考图像引导的视频字幕生成
RefCaptioner: Multi-Reference Image-Grounded Video Captioning
July 30, 2026
作者: Tengfei Liu, Yang Shi, Yuran Wang, Xiaohan Zhang, Yuqing Wen, Yuqi Tang, Qixun Wang, Zhuoran Zhang, Xuanyu Zhu, Weihong Lin, Xinlei Yu, Yujie Wei, Xinwei Long, Fengxiang Wang, Xinlong Chen, Yue Ding, Jialu Chen, Haotian Wang, Yuanxing Zhang
cs.AI
摘要
现有的视频字幕生成模型能够生成对视频内容的自然描述,但无法将局部视觉元素显式关联到多张参考图像上。我们提出了多参考图像引导的视频字幕生成这一新任务,该任务要求生成具有短语级参考定位的事实性视频描述,并为此设计了RefCaptioner——一个两阶段后训练框架。RefCaptioner将混合数据SFT与分层覆盖率折扣GRPO相结合,在保持通用视频字幕生成能力的同时,联合提升参考选择、短语级绑定、干扰项拒绝和跨参考一致性。为支持训练,我们构建了一个包含20,000个视频和171,354张参考图像的语料库。我们还提出了MRVBench基准,用于在真实世界视频和AI生成视频上评估字幕的事实准确性和多参考定位能力。实验表明,RefCaptioner在开源模型中取得了最佳的整体性能,同时在标准视频字幕生成基准上保持竞争力。人工评估进一步证实,其生成的字幕更受标注人员青睐,并且在使用开源和专有视频生成器时,能够实现更忠于来源的视频重构。
English
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing 20,000 videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.