ChatPaper.aiChatPaper

FIRM-Video:先檢查再評分——可靠的文本到視頻獎勵建模

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

August 22, 2026
作者: Peiyuan Zhang, Xiangyu Zhao, Hongbo Liu, Xiaoxing Hu, Mingxin Liu, Shuran Ma, Yunhang Shen, Jian Hu, Haihan Gao, Haoyu Cao, Xue Yang
cs.AI

摘要

可靠的獎勵模型對於文字轉影片的評估與對齊至關重要。然而,評估準確性與推論效率之間的取捨,對訓練監督的品質提出了極高的要求。現有方法常依賴具固定評分標準或開放式推理的整體性評估器,導致檢查不完整、論證不可靠以及歸因混淆等問題。我們提出FIRM-Video,這是一個基於「先檢查後評分」原則的統一檢查清單驅動資料建構框架:建構各維度專屬的檢查清單,根據時序視覺證據逐一驗證標準,並僅彙總已驗證的決策。在指令遵循方面,FIRM-Video將提示詞分解為加權的原子化需求;在世界一致性方面,它基於可見的實體與動作,建構經提示詞校準且針對特定目標的檢查項目;在感知品質方面,它套用通用的視覺缺陷分類體系。已驗證的標準與分數進一步轉化為自然語言分析,以進行端到端的獎勵建模。隨後,我們從29,348部影片建構了包含88,044個維度專屬實例的FIRM-Video-90K,並推出了包含250部影片、共750筆逐點人工標註的FIRM-Video-Bench。基於Qwen3-VL的FIRM-Video-8B在FIRM-Video-Bench上取得最佳的整體MAE,同時在三個影片生成器的Best-of-8採樣中,持續獲得最高的VBench總分、品質分數與語意分數。
English
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.