FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为提高文本到视频奖励模型的可靠性,提出FIRM-Video框架,通过检查清单驱动的方法验证各评估维度,确保决策基于实际视觉证据。
📝 Abstract
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.
Problem

Research questions and friction points this paper is trying to address.

text-to-video
reward modeling
evaluation accuracy
inference efficiency
training supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

check-before-score
dimension-specific checklists
temporal visual evidence
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30