DREAM テクニカルレポート
DREAM Technical Report
August 13, 2026
著者: Bin Zhang, Bowen Zheng, Chao Yi, Chengyu Lai, Dian Chen, Dimin Wang, Gaoyang Guo, Jialin Zhu, Jian Wu, Jing Yu, Jiuning Lin, Lingqing Zhang, Lingyun Zheng, Mao Zhang, Mingming Pan, Ruiquan Lan, Shuai Zhong, Wen Chen, Wendong Zhang, Xiaodong Zhu, Xuan Chen, Xunke Xi, Yifan Lu, Yiheng Wang, Yue Zeng, Yujie Luo, Yuning Jiang, Zhe Hu, Zhibo Xiao, Zihong Huang, Binbin Cao, Bo Zheng, Danning Wang, Dixuan Wang, Ge Fan, Haixia Wu, Han Zhu, Hao Fang, Haoming Chen, Huiping Chu, Jian Wang, Jianjun Wu, Jiawei Wu, Jiaxin Yu, Jingwen Liu, Jinzhe Shan, Kai Meng, Kai Zhang, Keqin Xu, Kewei Zhu, Lang Tian, Leihui Chen, Li Chen, Licheng Xu, Lide Xiao, Ruitong Zhang, Shiyao Peng, Silu Zhou, Tao Wang, Wei Shi, Wenjun Yang, Xiang Chen, Xiang Gao, Xiao Ren, Xu Liu, Xuwen Wang, Yang Li, Yeqiu Yang, Yi Hu, Yichen Yuan, Yinnan Song, Yipeng Yu, Yuan Liu, Yunqi Gao, Zhiliang Huang, Zhujin Gao, Zongyuan Wu
cs.AI
要旨
産業用レコメンダーシステムは通常、カスケード型の検索(リトリーバル)、ランキング、再ランキングからなるパイプラインを採用する。効率的ではあるものの、これらのパイプラインはモジュール間で情報と目的を断片化し、固定的なルールに依存し、リアルタイムの意図に対する認識が限られており、閲覧、比較、購入の間のセッションレベルの変化が十分に対処されないままになっている。我々は、既存のパイプラインを置き換えることなく、その上位に認識対応型でオーケストレーション可能かつ監査可能なポリシーレイヤーを追加する自律的最適化制御アーキテクチャであるDREAM(エージェント的手法を用いたレコメンダーエンジン開発)を提案する。DREAMには2つの中核コンポーネントがある。第一に、3層インテントエンジンがオンデバイス信号を構造化されたL0/L1/L2インテント表現に融合する。そのエッジクラウドトリガーチェーンはレポート量を約8.7%に削減する。第二に、メタエンジンは、階層的なM1→M2→M3推論(インテント要約、Strategy Memoryに基づく戦略計画、パラメータ変換)のためのメタモデルを使用する。そして、安全ガードレール付きの統合出力経路を通じて、得られたパラメータをディスパッチする。報酬デュアルループは、戦略空間探索のためのオフラインシミュレーションと結果較正のためのオンラインフィードバックを組み合わせることで両コンポーネントを継続的に最適化し、生成、実行、評価、経験蓄積のサイクルを形成する。タオバオのホームページフィードでの大規模A/Bテストにより、再ランキング制御のみでIPVが2.06%、コアIPVが2.39%、GMVが0.88%改善されることが示された。制御をファインランキング(精密ランキング)まで拡張すると、これらの改善はそれぞれ2.71%、3.06%、1.31%に向上し、同時にPVも一貫して1%以上改善される。これらの改善は、パイプラインモデルの置き換えも、サービング安定性を損なうことも必要とせず、エージェント的メタ制御が産業レコメンデーションの実現可能なパラダイムであることを支持する。
English
Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.