AgenticGen: 광고를 위한 보상 기반 에이전틱 비디오 생성
AgenticGen: Reward-Guided Agentic Video Generation for Advertising
August 31, 2026
저자: Xingyuan Bu, Chengru Song, Hao Zhou, Tao Zhou, Dong Li, Wei Li, Shilong Li, Hao Shi, Yongxin Guo, Donghao Zhou, Qiangpeng Yang, Shilei Wen
cs.AI
초록
광고 비디오 생성은 비디오 합성 작업일 뿐만 아니라, 성공이 온라인 비즈니스 지표로 측정되는 제품 조건부 추론 문제이기도 하다. 최근의 비디오 파운데이션 모델은 멀티모달 조건으로부터 사실적인 클립을 생성할 수 있지만, 제품이 효과적인 광고로 어떻게 변환되어야 하는지나 향후 생성이 온라인 비즈니스 피드백으로부터 어떻게 개선되어야 하는지를 최적화하지는 않는다. 이 루프를 닫기 위해, 우리는 광고 비디오 생성을 두 가지 학습 가능한 추론 단계인 전략 선택과 초안 생성으로 분해함으로써 온라인 비즈니스 피드백이 감독할 수 있는 최적화 목표를 드러내는 보상 기반 에이전트 프레임워크인 AgenticGen을 제안한다. AgenticGen은 축적된 온라인 피드백으로부터 성과 기반 보상을 학습하고, 인간 품질 기준과 정렬된 보완적 루브릭 기반 보상을 학습한 다음, 이를 사용하여 정책 최적화를 감독한다. DPO는 먼저 에이전트 정책을 온라인 선호도 쪽으로 이동시키고, GRPO는 과정 및 결과 보상으로 두 단계를 추가로 정교화한다. 오프라인 실험은 보상 모델과 순차적 정책 최적화를 검증한다. TikTok 광고 시스템에서의 온라인 A/B 실험은 DPO와 GRPO를 거친 AgenticGen이 SFT 기준선 대비 CTR을 2.72%, CVR을 2.63%, Advv를 9.61% 향상시킨다는 것을 보여준다.
English
Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose success is measured by online business metrics. Recent video foundation models can generate realistic clips from multimodal conditions, yet they do not optimize how a product should be transformed into an effective advertisement or how future generation should be improved from online business feedback. To close this loop, we propose AgenticGen, a reward-guided agentic framework that decomposes advertising video generation into two trainable reasoning stages, strategy selection and draft generation, thereby exposing optimization targets that online business feedback can supervise. AgenticGen learns a performance-based reward from accumulated online feedback and a complementary rubric-based reward aligned with human quality standards, then uses them to supervise policy optimization. DPO first moves the agentic policies toward online preferences, and GRPO further refines both stages with process and outcome rewards. Offline experiments validate the reward models and successive policy optimization. Online A/B experiments in the TikTok advertising system show that AgenticGen after DPO and GRPO improves CTR by 2.72%, CVR by 2.63%, and Advv by 9.61% over the SFT baseline.