RefCaptioner:多參考影像基礎的影片字幕生成
RefCaptioner: Multi-Reference Image-Grounded Video Captioning
July 30, 2026
作者: Tengfei Liu, Yang Shi, Yuran Wang, Xiaohan Zhang, Yuqing Wen, Yuqi Tang, Qixun Wang, Zhuoran Zhang, Xuanyu Zhu, Weihong Lin, Xinlei Yu, Yujie Wei, Xinwei Long, Fengxiang Wang, Xinlong Chen, Yue Ding, Jialu Chen, Haotian Wang, Yuanxing Zhang
cs.AI
摘要
現有的影片字幕生成模型雖能生成自然的影片內容描述,但無法將局部視覺元素明確地對齊至多張參考影像。我們提出多參考影像對齊的影片字幕生成任務,這是一項新任務,要求生成具事實性的影片描述,並在詞組層級上對齊參考影像。為此,我們提出 RefCaptioner,一個適用於此任務的兩階段後訓練框架。RefCaptioner 結合混合資料的監督式微調與分層覆蓋折扣 GRPO,在保留一般影片字幕生成能力的同時,共同改善參考影像選擇、詞組層級綁定、干擾項排除與跨參考一致性。為支援訓練,我們建構了一個包含 20,000 部影片與 171,354 張參考影像的語料庫。我們進一步提出 MRVBench,這是一個用於評估描述事實性與多參考對齊能力的基準,涵蓋真實世界影片與 AI 生成影片。實驗結果顯示,RefCaptioner 在開源模型中達到最佳整體表現,同時在標準影片字幕生成基準上仍具競爭力。人工評估進一步證實,其生成的描述更受標註者偏好,且能透過開源與專有的影片生成器達成更忠於來源的影片重建。
English
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing 20,000 videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.