DREAM 技術報告
DREAM Technical Report
August 13, 2026
作者: Bin Zhang, Bowen Zheng, Chao Yi, Chengyu Lai, Dian Chen, Dimin Wang, Gaoyang Guo, Jialin Zhu, Jian Wu, Jing Yu, Jiuning Lin, Lingqing Zhang, Lingyun Zheng, Mao Zhang, Mingming Pan, Ruiquan Lan, Shuai Zhong, Wen Chen, Wendong Zhang, Xiaodong Zhu, Xuan Chen, Xunke Xi, Yifan Lu, Yiheng Wang, Yue Zeng, Yujie Luo, Yuning Jiang, Zhe Hu, Zhibo Xiao, Zihong Huang, Binbin Cao, Bo Zheng, Danning Wang, Dixuan Wang, Ge Fan, Haixia Wu, Han Zhu, Hao Fang, Haoming Chen, Huiping Chu, Jian Wang, Jianjun Wu, Jiawei Wu, Jiaxin Yu, Jingwen Liu, Jinzhe Shan, Kai Meng, Kai Zhang, Keqin Xu, Kewei Zhu, Lang Tian, Leihui Chen, Li Chen, Licheng Xu, Lide Xiao, Ruitong Zhang, Shiyao Peng, Silu Zhou, Tao Wang, Wei Shi, Wenjun Yang, Xiang Chen, Xiang Gao, Xiao Ren, Xu Liu, Xuwen Wang, Yang Li, Yeqiu Yang, Yi Hu, Yichen Yuan, Yinnan Song, Yipeng Yu, Yuan Liu, Yunqi Gao, Zhiliang Huang, Zhujin Gao, Zongyuan Wu
cs.AI
摘要
工業推薦系統通常採用級聯式的召回、排序與重排管線。儘管這些管線高效,但它們使資訊與目標在各模組間碎片化,依賴僵化規則,且對即時意圖的感知有限,導致瀏覽、比較與購買之間的會話級轉換未獲充分處理。我們提出DREAM(Developing Recommender Engine with Agentic Methods,即以智能體方法開發推薦引擎),這是一種自主最佳化控制架構,在不取代現有管線的前提下,於其上增加一層具感知能力、可編排且可審計的策略層。DREAM包含兩大核心組件。第一,三層意圖引擎(Intent Engine)將設備端信號融合為結構化的L0/L1/L2意圖表徵;其邊雲觸發鏈將上報量降至約8.7%。第二,元引擎(Meta Engine)利用元模型進行M1→M2→M3的分層推理:意圖摘要、由策略記憶(Strategy Memory)輔助的策略規劃,以及參數轉譯。元引擎透過具安全護欄的統一出入口分發所產生的參數。獎勵雙迴路(Reward Dual Loop)將離線模擬與線上回饋相結合——前者用於策略空間探索,後者用於結果校準——持續最佳化上述兩大組件,形成生成、執行、評估與經驗累積的閉環。淘寶首頁資訊流上的大規模A/B測試表明,僅對重排進行控制即可使IPV提升2.06%、核心IPV提升2.39%、GMV提升0.88%;將控制範圍擴展至精排後,上述增益分別提升至2.71%、3.06%與1.31%,同時PV持續提升超過1%。上述增益既無需更換管線模型,亦不損害服務穩定性,支持智能體式元控制作為工業推薦系統的一種可行範式。
English
Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.