ChatPaper.aiChatPaper

전환기의 지속적 학습

Continual Learning in Transition

August 6, 2026
저자: Zhiyan Hou, Dan Zhang, Tao Feng, Liyuan Wang, Wei Li, Xiangzhao Hao, Hongyan An, Junfeng Fang, Haokai Ma, Zhaohui Xu, Haiyun Guo, Jinqiao Wang, Tat-Seng Chua
cs.AI

초록

고전적인 지속 학습(CL)은 주로 훈련 전략, 아키텍처 설계, 가중치 적응과 같은 매개변수 중심 메커니즘, 예컨대를 통해 모델이 지식을 갱신하고 유지할 수 있게 하는 데 초점을 맞춰 왔다. 그러나 새로운 패러다임들은 이러한 전통적인 모델 적응 관점을 넘어 CL의 범위를 재편하고 있다. 예를 들어, 온-정책 학습은 갱신 메커니즘의 공간을 확장하고, 테스트 시점 훈련은 CL을 훈련 단계에서 추론 단계로 확장하며, 메모리, 기술 라이브러리, 상호작용 프로토콜과 같은 외부 하네스 구성 요소들은 모델 역량의 진화 경계를 정적 매개변수 공간을 훨씬 넘어 확장한다. 종합하면, 이러한 발전들은 매개변수 중심 학습에서 시스템 수준 적응으로의 전환을 시사한다. 이러한 전환의 특성을 규명하기 위해, 우리는 학습이 언제(When), 어떻게(How), 어디서(Where) 발생하는지라는 세 가지 차원을 통해 지속 학습의 진화를 살펴본다. How 차원은 오프-정책, 온-정책, 그리고 경사 기반을 넘어선 최적화 메커니즘을 포괄한다. When 차원은 사전 훈련, 사후 훈련, 추론 시점 단계에 걸친 진화를 포착한다. Where 차원은 내부 매개변수 내에서 발생하는 갱신과 외부 구조적 제약에서 발생하는 갱신을 구분한다. 이 삼축 프레임워크에 기반하여, 우리는 대표적인 방법들을 체계적으로 조사하고, 지속 학습의 진행 중인 전환을 추적하며, 이러한 패러다임 전환에서 비롯된 핵심 과제, 광범위한 시사점, 그리고 향후 방향을 논의한다.
English
Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g., training strategies, architectural designs, and weight adaptation. However, emerging paradigms are reshaping the scope of CL beyond this traditional model adaptation view. For instance, on-policy learning broadens the space of update mechanisms; test-time training extends CL from the training phase to inference; and external harness components such as memory, skill libraries, and interaction protocols extend the evolutionary boundaries of model capabilities far beyond the static parameter space. Collectively, these developments indicate a transition from parameter-centric learning toward system-level adaptation. To characterize this transition, we examine the evolution of continual learning through three dimensions: When, How, and Where learning occurs. The How dimension encompasses off-policy, on-policy, and beyond-gradient optimization mechanics. The When dimension captures evolution across pre-training, post-training, and inference-time stages. The Where dimension delineates updates occurring within internal parameters versus external structural constraints. Anchored by this tri-axial framework, we systematically survey representative methods, trace the ongoing transition of continual learning, and discuss the key challenges, broader implications, and future directions arising from this paradigm shift.