ChatPaper.aiChatPaper

지식-기하학 디커플링: 스트리밍 추천을 위한 갱신 가능한 사전학습 전이

Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation

August 3, 2026
저자: Zixuan Wang, Yuhong Chen, Yuxuan Zhu, Guidong Lei, Zhiluohan Guo, Yu Zhao, Kun Wang, Bangyang Hong, Kangle Wu, Yabo Ni, Anxiang Zeng, Cong Fu, Hui Li
cs.AI

초록

산업용 추천 시스템은 점차 사전학습-후-전이 패러다임을 채택하고 있지만, 행동 분포 변화로 인해 두 가지 질문이 제기된다: 행동 시퀀스에서 무엇을 학습할 것인가, 그리고 사전학습 모델이 지속적으로 갱신되는 상황에서 학습된 지식을 어떻게 전이할 것인가. 이러한 문제를 해결하기 위해 우리는 지식-기하 구조 분리(Knowledge-Geometry Decoupling, KGD)를 제안한다. 무엇을 학습할 것인지에 대해, 기존의 다음 토큰 예측은 인접성을 의존성으로 간주하여 관련 없는 세션 간의 허위 전이를 인코딩할 수 있다. 우리는 협업적 또는 의미적으로 관련된 미래 항목만을 지도 신호로 유지하는 행동 다중 토큰 예측(Behavioral Multi-Token Prediction, BMTP)을 도입하여 더 깨끗하고 전이 가능한 행동 지식을 확보한다. 어떻게 전이할 것인지에 대해, 사전학습된 지식과 작업별 기하 구조는 공유 파라미터에 상충되는 최적화 요구를 부과한다. 이를 처리하기 위해 KGD는 이를 별도의 파라미터 집합에 할당한다: 갱신 가능한 인코더는 행동 지식을 소유하고, 작업 학습자는 읽기 전용 교차 어텐션을 통해 컨텍스트화된 인코더 상태를 읽으며, 사전학습 임베딩에 직교하는 고정 보정 잔차(Anchored Calibration Residual, ACR)를 통해 작업별 기하 구조를 기록한다. 분리된 소유 구조는 작업 그래디언트 간섭이나 다운스트림 적응 무효화 없이 지속적인 지식 갱신을 가능하게 한다. KGD는 여덟 개의 공개 벤치마크에서 강력한 사전학습-전이 기준선 대비 4-12%의 성능 향상을 보이며, 기준선이 이득을 보이지 못하는 90일간의 프로덕션 스트림에서도 그 우위를 유지한다. KGD는 Shopee에 완전히 배포되었다. Shopee 홈페이지 검색에서 진행된 실시간 A/B 테스트에서 KGD는 사용자당 GMV를 1.75%, 광고 수익을 1.53% 증가시켜 높은 실용적 가치를 입증한다. KGD의 핵심 구현은 https://github.com/FuCongResearchSquad/KGD4REC에서 제공한다.
English
Industrial recommenders increasingly adopt the pretrain-then-transfer paradigm, yet behavioral distribution drift raises two questions: what to learn from behavior sequences, and how to transfer the learned knowledge while the pretrained model is continually refreshed. To resolve them, we propose Knowledge-Geometry Decoupling (KGD). For what to learn, conventional next-token prediction treats adjacency as dependency and may encode spurious transitions across unrelated sessions. We introduce Behavioral Multi-Token Prediction (BMTP) to retain only collaboratively or semantically related future items as supervision, yielding cleaner and more transferable behavioral knowledge. For how to transfer, pretrained knowledge and task-specific geometry impose conflicting optimization demands on shared parameters. To handle it, KGD assigns them to separate parameter sets: a refreshable encoder owns behavioral knowledge, while a task learner reads contextualized encoder states through read-only cross-attention and writes task-specific geometry through Anchored Calibration Residual (ACR) orthogonal to the pretrained embedding. The decoupled ownership enables continual knowledge refresh without task-gradient interference or invalidating downstream adaptation. KGD improves over strong pretrain-transfer baselines by 4-12% on eight public benchmarks and sustains its advantage over a 90-day production stream where baselines show no gains. KGD has been fully deployed in Shopee. In a live A/B test on Shopee Homepage Search, it increases GMV per user by 1.75% and advertising revenue by 1.53%, demonstrating its high practical value. We provide the core implementation of KGD at https://github.com/FuCongResearchSquad/KGD4REC.