FIRM-Video:先检查后评分,实现可靠的文本到视频奖励建模
FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
August 22, 2026
作者: Peiyuan Zhang, Xiangyu Zhao, Hongbo Liu, Xiaoxing Hu, Mingxin Liu, Shuran Ma, Yunhang Shen, Jian Hu, Haihan Gao, Haoyu Cao, Xue Yang
cs.AI
摘要
可靠的奖励模型对于文生视频评估和对齐至关重要。然而,评估准确性与推理效率之间的权衡对训练监督的质量提出了很高要求。现有方法通常依赖具有固定评分标准或开放式推理的整体评判器,导致检查不完整、理由不忠实和归因纠缠。我们提出 FIRM-Video,一种基于“先检查后评分”原则的统一清单驱动数据构建框架:构建维度特定清单,依据时序视觉证据验证每条标准,并仅聚合已验证的判定。在指令遵循方面,FIRM-Video 将提示分解为加权原子要求;在世界一致性方面,它构建基于可见实体和动作的、经提示校准的目标特定检查;在感知质量方面,它应用通用的视觉缺陷分类体系。经过验证的标准与分数进一步转化为自然语言分析,用于端到端奖励建模。随后,我们构建了 FIRM-Video-90K,包含来自 29,348 个视频的 88,044 个维度特定实例,并推出 FIRM-Video-Bench,涵盖 250 个视频上的 750 条逐点人工标注。基于 Qwen3-VL 的 FIRM-Video-8B 在 FIRM-Video-Bench 上取得最佳整体 MAE,同时在三个视频生成器的 Best-of-8 采样中始终取得最高的 VBench 总评分、质量评分和语义评分。
English
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.