AgenticGen:面向廣告的獎勵引導代理式視訊生成
AgenticGen: Reward-Guided Agentic Video Generation for Advertising
August 31, 2026
作者: Xingyuan Bu, Chengru Song, Hao Zhou, Tao Zhou, Dong Li, Wei Li, Shilong Li, Hao Shi, Yongxin Guo, Donghao Zhou, Qiangpeng Yang, Shilei Wen
cs.AI
摘要
廣告影片生成不僅是一項影片合成任務,也是一個以產品為條件的推理問題,其成功與否取決於線上商業指標。近期的影片基礎模型能從多模態條件生成逼真的短片,但它們並未最佳化產品應如何轉化為有效廣告,或未來生成應如何根據線上商業回饋加以改進。為了閉合此迴路,我們提出 AgenticGen,一個獎勵引導的代理式框架,將廣告影片生成分解為兩個可訓練的推理階段:策略選擇與草稿生成,藉此揭露線上商業回饋可監督的優化目標。AgenticGen 從累積的線上回饋學習以績效為基礎的獎勵,以及一個與人類品質標準一致、具互補性的基於評分規準之獎勵,並以這些獎勵監督策略優化。DPO 首先使代理式策略朝向線上偏好移動,GRPO 則以過程與結果獎勵進一步精煉這兩個階段。離線實驗驗證了獎勵模型與逐次策略優化。TikTok 廣告系統中的線上 A/B 實驗顯示,經過 DPO 與 GRPO 後的 AgenticGen 相較於 SFT 基線,CTR 提升 2.72%、CVR 提升 2.63%,Advv 提升 9.61%。
English
Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose success is measured by online business metrics. Recent video foundation models can generate realistic clips from multimodal conditions, yet they do not optimize how a product should be transformed into an effective advertisement or how future generation should be improved from online business feedback. To close this loop, we propose AgenticGen, a reward-guided agentic framework that decomposes advertising video generation into two trainable reasoning stages, strategy selection and draft generation, thereby exposing optimization targets that online business feedback can supervise. AgenticGen learns a performance-based reward from accumulated online feedback and a complementary rubric-based reward aligned with human quality standards, then uses them to supervise policy optimization. DPO first moves the agentic policies toward online preferences, and GRPO further refines both stages with process and outcome rewards. Offline experiments validate the reward models and successive policy optimization. Online A/B experiments in the TikTok advertising system show that AgenticGen after DPO and GRPO improves CTR by 2.72%, CVR by 2.63%, and Advv by 9.61% over the SFT baseline.