심층 강화 학습 평가 및 설계 패러다임에 대한 원칙적 분석
Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms
July 8, 2026
저자: Ezgi Korkmaz
cs.AI
초록
가장 도전적인 게임 중 하나를 정복하게 한 심층 신경망을 활용한 상태-행동 가치 함수 근사화에서부터, 당면한 과제의 규칙을 명시적으로 제시하지 않고도 문제를 해결할 수 있게 한 알고리즘적 발전에 이르기까지, 강화 학습 연구는 지난 10년간 놀라운 과학적 진보의 중심에 있어 왔다. 본 논문에서는 이러한 연구 진전의 핵심 요소에 초점을 맞추고, 강화 학습의 정규 설계 및 평가 패러다임을 분석한다. 우리는 강화 학습에서 스케일링 법칙의 이론적 기초를 소개하고, 강화 학습 알고리즘의 점근적 성능이 성능 순위와 데이터 영역 간에 단조로운 관계를 갖지 않음을 보인다. 대규모 실험을 수행한 결과, 정규 설계 및 평가 패러다임 하에서의 일련의 강화 학습 연구가 잘못된 결론을 초래했음을 입증한다. 우리의 분석과 결과는 심층 강화 학습의 스케일링, 용량 및 복잡성에 대한 핵심 분석을 제공한다.
English
Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even explicitly stating the rules of the challenge at hand, reinforcement learning research has been the center of remarkable scientific progress for the past decade. In this paper, we focus on the key ingredients of this research progress and we analyze the canonical evaluation and design paradigms in reinforcement learning. We introduce the theoretical foundations of scaling laws in reinforcement learning and show that the asymptotic performance of reinforcement learning algorithms does not have a monotone relationship between performance rankings and data-regimes. We conduct large-scale experiments and our results demonstrate that a line of reinforcement learning research under the canonical design and evaluation paradigms resulted in incorrect conclusions. Our analysis and results provide a core analysis on scaling, capacity and complexity of deep reinforcement learning.