RefCaptioner: 複数参照画像に基づくビデオキャプショニング
RefCaptioner: Multi-Reference Image-Grounded Video Captioning
July 30, 2026
著者: Tengfei Liu, Yang Shi, Yuran Wang, Xiaohan Zhang, Yuqing Wen, Yuqi Tang, Qixun Wang, Zhuoran Zhang, Xuanyu Zhu, Weihong Lin, Xinlei Yu, Yujie Wei, Xinwei Long, Fengxiang Wang, Xinlong Chen, Yue Ding, Jialu Chen, Haotian Wang, Yuanxing Zhang
cs.AI
要旨
既存のビデオキャプショニングモデルは、ビデオコンテンツの自然な記述を生成できるが、局所的な視覚要素を複数の参照画像に明示的に接地することはできない。本稿では、フレーズレベルの参照接地を伴う事実に基づくビデオ記述を必要とする新しいタスクであるマルチ参照画像接地ビデオキャプショニングを導入し、このタスクのための2段階ポストトレーニングフレームワークであるRefCaptionerを提案する。RefCaptionerは、混合データSFTと階層的カバレッジ割引GRPOを組み合わせることで、参照選択、フレーズレベルのバインディング、ディストラクタ棄却、および参照間一貫性を同時に改善しつつ、一般的なビデオキャプショニング能力を維持する。トレーニングを支援するため、20,000本のビデオと171,354枚の参照画像を含むコーパスを構築する。さらに、実世界およびAI生成ビデオの両方において、キャプションの事実性とマルチ参照接地を評価するためのベンチマークであるMRVBenchを導入する。実験により、RefCaptionerは標準的なビデオキャプショニングベンチマークにおいて競争力を維持しつつ、オープンソースモデルの中で最良の総合性能を達成することが示された。さらに、人間による評価により、そのキャプションがアノテータに選好され、オープンソースおよびプロプライエタリの両方のビデオ生成モデルにおいて、よりソースに忠実なビデオ再構成を可能にすることが確認された。
English
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing 20,000 videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.