DREAM 기술 보고서
DREAM Technical Report
August 13, 2026
저자: Bin Zhang, Bowen Zheng, Chao Yi, Chengyu Lai, Dian Chen, Dimin Wang, Gaoyang Guo, Jialin Zhu, Jian Wu, Jing Yu, Jiuning Lin, Lingqing Zhang, Lingyun Zheng, Mao Zhang, Mingming Pan, Ruiquan Lan, Shuai Zhong, Wen Chen, Wendong Zhang, Xiaodong Zhu, Xuan Chen, Xunke Xi, Yifan Lu, Yiheng Wang, Yue Zeng, Yujie Luo, Yuning Jiang, Zhe Hu, Zhibo Xiao, Zihong Huang, Binbin Cao, Bo Zheng, Danning Wang, Dixuan Wang, Ge Fan, Haixia Wu, Han Zhu, Hao Fang, Haoming Chen, Huiping Chu, Jian Wang, Jianjun Wu, Jiawei Wu, Jiaxin Yu, Jingwen Liu, Jinzhe Shan, Kai Meng, Kai Zhang, Keqin Xu, Kewei Zhu, Lang Tian, Leihui Chen, Li Chen, Licheng Xu, Lide Xiao, Ruitong Zhang, Shiyao Peng, Silu Zhou, Tao Wang, Wei Shi, Wenjun Yang, Xiang Chen, Xiang Gao, Xiao Ren, Xu Liu, Xuwen Wang, Yang Li, Yeqiu Yang, Yi Hu, Yichen Yuan, Yinnan Song, Yipeng Yu, Yuan Liu, Yunqi Gao, Zhiliang Huang, Zhujin Gao, Zongyuan Wu
cs.AI
초록
산업용 추천 시스템은 일반적으로 계단식(cascaded) 검색, 랭킹, 재랭킹 파이프라인을 사용한다. 효율적이기는 하지만, 이러한 파이프라인은 모듈 간 정보와 목표를 단편화하고, 경직된 규칙에 의존하며, 실시간 의도에 대한 인식이 제한적이어서 탐색, 비교, 구매 사이의 세션 수준 전환을 충분히 처리하지 못한다. 본 논문은 기존 파이프라인을 대체하지 않으면서 그 위에 인식 기반(perception-aware), 조율 가능(orchestrable), 감사 가능(auditable)한 정책 계층을 추가하는 자율 최적화 제어 아키텍처인 DREAM(Developing Recommender Engine with Agentic Methods)을 제시한다. DREAM은 두 가지 핵심 구성 요소를 갖는다. 첫째, 3계층 의도 엔진(Intent Engine)은 기기 내 신호를 구조화된 L0/L1/L2 의도 표현으로 융합하며, 엣지-클라우드 트리거 체인을 통해 보고량을 약 8.7%로 감소시킨다. 둘째, 메타 엔진(Meta Engine)은 메타모델(MetaModel)을 사용하여 의도 요약, 전략 메모리(Strategy Memory)에 기반한 전략 계획, 파라미터 변환으로 이어지는 계층적 M1-M2-M3 추론을 수행한다. 이후 통합 출력 채널을 통해 안전 가드레일을 적용한 결과 파라미터를 전달한다. 이중 보상 루프(Reward Dual Loop)는 전략 공간 탐색을 위한 오프라인 시뮬레이션과 결과 보정을 위한 온라인 피드백을 결합하여 생성, 실행, 평가, 경험 축적의 순환을 형성함으로써 두 구성 요소를 지속적으로 최적화한다. 타오바오 홈페이지 피드에서의 대규모 A/B 테스트 결과, 재랭킹 제어만으로도 IPV가 2.06%, 핵심 IPV가 2.39%, GMV가 0.88% 개선되었다. 정밀 랭킹까지 제어 범위를 확장하면 이러한 개선은 각각 2.71%, 3.06%, 1.31%로 증가하며, PV는 일관되게 1% 이상 개선된다. 이러한 성과는 파이프라인 모델의 교체나 서빙 안정성의 저하 없이 달성되며, 에이전트형 메타 제어가 산업용 추천의 실행 가능한 패러다임임을 뒷받침한다.
English
Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.