WorldReward:基于相机条件的世界模型的奖励建模
WorldReward: Reward Modeling for Camera-Conditioned World Models
September 3, 2026
作者: Yibin Wang, Zehan Wang, Junshu Tang, Zhimin Li, Yujie Zhou, Jiazi Bu, Pengyang Ling, Feng Han, Zhixiong Zhang, Long Xing, Shengyuan Ding, Ziang Li, Cheng Jin, Yuhang Zang, Jiaqi Wang, Tianyu Pang
cs.AI
摘要
基于相机的世界模型生成交互式视频,其中指令动作应引发期望的场景变化,同时外观、几何和时间动态保持连贯。现有奖励方法分别评估这些需求:基于几何的奖励估计轨迹执行质量,但无法判断所执行动作的视觉质量;而基于图像的奖励衡量帧质量,却不捕捉动作执行或时间动态。我们提出,视觉语言模型(VLM)为关联动作与其视觉结果提供了共享推理空间。然而,针对完整动作序列评判整个长视频会形成冗长且嘈杂的上下文,其中短暂存在的局部动作证据可能被遗漏或稀释。我们提出WorldReward,一种基于VLM的成对偏好奖励模型,统一了基于相机的世界模型的动作一致性和视觉质量评估。WorldReward将成对视频分解为动作对齐的分块,将每个分块组织为结构化视觉证据,并通过投票将分块级决策聚合为独立的视频级动作偏好和视觉质量偏好。为训练该模型,我们使用由前沿VLM生成并经基于工具的主体审计和针对性人工审查精化的结构化判断,构建了大规模推理增强偏好数据集。我们进一步引入WorldReward-Bench,一个人工标注基准,用于衡量奖励模型在动作一致性、外观质量和运动质量三个维度上与人类偏好的一致性。WorldReward在所有三个维度上均取得最高一致性,分别超过GPT-5.5达3.42、1.45和3.56个百分点。当用于HY-WorldPlay 1.5的强化学习后训练时,它在短期到长期时间范围内持续改善动作执行和视觉质量。
English
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics remain coherent. Existing rewards assess these requirements separately: geometry-based rewards estimate trajectory execution but cannot judge the visual quality of the executed motion, whereas image-based rewards measure frame quality without capturing action execution or temporal dynamics. We posit that a vision-language model (VLM) offers a shared reasoning space for relating actions to their visual outcomes. However, judging a complete long video against its full action sequence creates a lengthy, noisy context in which short-lived local action evidence can be missed or diluted. We present WorldReward, a VLM-based pairwise preference reward model that unifies action-consistency and visual-quality evaluation for camera-conditioned world models. WorldReward decomposes paired videos into action-aligned chunks, organizes each chunk into structured visual evidence, and aggregates chunk-level decisions by voting into separate video-level action and visual-quality preferences. To train it, we construct a large-scale reasoning-augmented preference dataset using structured judgments generated by a frontier VLM and refined through tool-based agent auditing and targeted human review. We further introduce WorldReward-Bench, a human-annotated benchmark measuring reward-model agreement with human preferences across action consistency, appearance quality, and motion quality. WorldReward achieves the highest agreement on all three dimensions, exceeding GPT-5.5 by 3.42, 1.45, and 3.56 percentage points, respectively. When used for RL post-training of HY-WorldPlay 1.5, it consistently improves both action execution and visual quality across short- to long-term horizons.