ChatPaper.aiChatPaper

FIRM-Video: 신뢰할 수 있는 텍스트-비디오 보상 모델링을 위해 점수를 매기기 전에 확인하라

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

August 22, 2026
저자: Peiyuan Zhang, Xiangyu Zhao, Hongbo Liu, Xiaoxing Hu, Mingxin Liu, Shuran Ma, Yunhang Shen, Jian Hu, Haihan Gao, Haoyu Cao, Xue Yang
cs.AI

초록

신뢰할 수 있는 보상 모델은 텍스트-투-비디오 평가와 정렬에 필수적이다. 그러나 평가 정확성과 추론 효율성 사이의 상충 관계는 훈련 감독의 품질에 높은 요구를 부과한다. 기존 접근법은 고정된 루브릭이나 개방형 추론을 사용하는 종합적 평가자에 의존하는 경우가 많아, 불완전한 검사, 신뢰할 수 없는 근거, 혼재된 귀인으로 이어진다. 우리는 점검-후-점수 원칙에 기반한 통합 체크리스트 기반 데이터 구축 프레임워크인 FIRM-Video를 제안한다: 차원별 체크리스트를 구성하고, 각 기준을 시간적 시각 증거와 대조하여 검증하며, 검증된 결정만 통합한다. 지시 따르기(Instruction Following)의 경우 FIRM-Video는 프롬프트를 가중치가 있는 원자적 요구사항으로 분해한다. 세계 일관성(World Coherence)의 경우 가시적 개체와 동작에 근거한, 프롬프트로 보정된 대상 특정 점검을 구성한다. 지각 품질(Perceptual Quality)의 경우 시각적 결함에 대한 일반적인 분류 체계를 적용한다. 검증된 기준과 점수는 또한 종단 간 보상 모델링을 위해 자연어 분석으로 변환된다. 이후 우리는 29,348개의 비디오에서 88,044개의 차원별 인스턴스를 갖춘 FIRM-Video-90K를 구축하고, 250개 비디오에 걸친 750개의 점별 인간 주석을 포함하는 FIRM-Video-Bench를 소개한다. Qwen3-VL 기반 FIRM-Video-8B는 FIRM-Video-Bench에서 전반적으로 최고의 MAE를 달성하며, 세 가지 비디오 생성기에서 Best-of-8 샘플링 시 VBench 총점, 품질 및 의미 점수에서 일관되게 가장 높은 점수를 제공한다.
English
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.