ChatPaper.aiChatPaper

WorldReward: カメラ条件付きワールドモデルのための報酬モデリング

WorldReward: Reward Modeling for Camera-Conditioned World Models

September 3, 2026
著者: Yibin Wang, Zehan Wang, Junshu Tang, Zhimin Li, Yujie Zhou, Jiazi Bu, Pengyang Ling, Feng Han, Zhixiong Zhang, Long Xing, Shengyuan Ding, Ziang Li, Cheng Jin, Yuhang Zang, Jiaqi Wang, Tianyu Pang
cs.AI

要旨

カメラ条件付きワールドモデルは、指定されたアクションが期待されるシーン変化を引き起こす一方で、外観・幾何・時間的ダイナミクスの一貫性を保つインタラクティブ動画を生成する。既存の報酬はこれらの要件を別々に評価する。すなわち、幾何学的報酬は軌跡の実行を推定できるが、実行されたモーションの視覚品質を判断できず、画像ベースの報酬はフレーム品質を測定するものの、アクションの実行や時間的ダイナミクスを捉えられない。我々は、視覚言語モデル(VLM)がアクションをその視覚的結果に関連付けるための共有推論空間を提供すると考える。しかしながら、完全な長い動画をその全アクション系列と照合して判断することは、長大でノイズの多いコンテキストを生み出し、その中で短時間の局所的なアクション証拠が見落とされたり希釈されたりする可能性がある。我々は、カメラ条件付きワールドモデルに対するアクション一貫性と視覚品質評価を統合するVLMベースのペア選好報酬モデルであるWorldRewardを提案する。WorldRewardは、ペアの動画をアクション対応チャンクに分解し、各チャンクを構造化された視覚的証拠に整理し、チャンクレベルの決定を投票により集約して、動画レベルのアクション選好と視覚品質選好を別々に導出する。これを訓練するために、我々は、フロンティアVLMによって生成され、ツールベースのエージェント監査と対象を絞った人間によるレビューを通じて精緻化された構造化判定を用いて、大規模な推論拡張選好データセットを構築する。さらに我々は、アクション一貫性・外観品質・モーション品質の各側面における報酬モデルと人間の選好との一致度を測定する、人間アノテーション付きベンチマークであるWorldReward-Benchを導入する。WorldRewardは、これら3つの次元すべてにおいて最高の一致度を達成し、GPT-5.5をそれぞれ3.42、1.45、3.56パーセンテージポイント上回る。HY-WorldPlay 1.5のRLポストトレーニングに使用すると、短期から長期にわたるホライズンで、アクション実行と視覚品質の両方を一貫して改善する。
English
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics remain coherent. Existing rewards assess these requirements separately: geometry-based rewards estimate trajectory execution but cannot judge the visual quality of the executed motion, whereas image-based rewards measure frame quality without capturing action execution or temporal dynamics. We posit that a vision-language model (VLM) offers a shared reasoning space for relating actions to their visual outcomes. However, judging a complete long video against its full action sequence creates a lengthy, noisy context in which short-lived local action evidence can be missed or diluted. We present WorldReward, a VLM-based pairwise preference reward model that unifies action-consistency and visual-quality evaluation for camera-conditioned world models. WorldReward decomposes paired videos into action-aligned chunks, organizes each chunk into structured visual evidence, and aggregates chunk-level decisions by voting into separate video-level action and visual-quality preferences. To train it, we construct a large-scale reasoning-augmented preference dataset using structured judgments generated by a frontier VLM and refined through tool-based agent auditing and targeted human review. We further introduce WorldReward-Bench, a human-annotated benchmark measuring reward-model agreement with human preferences across action consistency, appearance quality, and motion quality. WorldReward achieves the highest agreement on all three dimensions, exceeding GPT-5.5 by 3.42, 1.45, and 3.56 percentage points, respectively. When used for RL post-training of HY-WorldPlay 1.5, it consistently improves both action execution and visual quality across short- to long-term horizons.