ChatPaper.aiChatPaper

FIRM-Video: 信頼性の高いテキスト・トゥ・ビデオ報酬モデリングのためのスコアリング前チェック

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

August 22, 2026
著者: Peiyuan Zhang, Xiangyu Zhao, Hongbo Liu, Xiaoxing Hu, Mingxin Liu, Shuran Ma, Yunhang Shen, Jian Hu, Haihan Gao, Haoyu Cao, Xue Yang
cs.AI

要旨

信頼性の高い報酬モデルは、テキストから動画への評価とアライメントに不可欠です。しかし、評価精度と推論効率の間のトレードオフは、トレーニング時の教師信号の質に高い要求を課します。既存の手法は、固定されたルーブリックを用いる全体的判定器や自由形式の推論に依存することが多く、結果として、検査の不完全さ、不正確な根拠付け、帰属の混在を招いています。我々は、check-before-score原則に基づく統一的なチェックリスト駆動データ構築フレームワーク、FIRM-Videoを提案します。これは、次元別チェックリストを構築し、各基準を時間的視覚エビデンスに照らして検証し、検証済みの判定のみを集約するものです。命令追従においては、FIRM-Videoはプロンプトを重み付きの原子的要件に分解します。世界一貫性においては、可視のエンティティと行動に基づく、プロンプト調整済みの対象特異的チェックを構築します。知覚品質においては、視覚的欠陥の一般的な分類法を適用します。検証済みの基準とスコアは、さらに自然言語による分析へと変換され、エンドツーエンドの報酬モデリングに利用されます。続いて、我々は29,348本の動画から88,044件の次元別インスタンスを備えるFIRM-Video-90Kを構築し、250本の動画にわたる750件のポイント単位の人手アノテーションを備えるFIRM-Video-Benchを導入します。Qwen3-VLベースのFIRM-Video-8Bは、FIRM-Video-Benchにおいて最良の総合MAEを達成し、3つの動画生成器におけるBest-of-8サンプリングでは、VBenchのTotal、Quality、Semanticスコアを一貫して最高とします。
English
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.