ChatPaper.aiChatPaper

深度强化学习评估与设计范式的原则性分析

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

July 8, 2026
作者: Ezgi Korkmaz
cs.AI

摘要

从利用深度神经网络逼近状态-动作值函数,并借此攻克最具挑战性的游戏之一,到算法层面的突破使得即使不明确陈述待解决问题的规则也能完成求解,强化学习研究在过去十年间始终是科学进步的核心推动力。本文聚焦于这一研究进展中的关键要素,系统分析了强化学习领域经典的评估与设计范式。我们介绍了强化学习尺度化理论的数学基础,并证明强化学习算法的渐近性能在性能排名与数据规模区间之间并不存在单调关系。通过大规模实验,我们展示出在经典设计与评估范式指导下的一系列强化学习研究得出了错误结论。本文的分析与结果为深度强化学习的尺度化、容量与复杂度研究提供了核心理论依据。
English
Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even explicitly stating the rules of the challenge at hand, reinforcement learning research has been the center of remarkable scientific progress for the past decade. In this paper, we focus on the key ingredients of this research progress and we analyze the canonical evaluation and design paradigms in reinforcement learning. We introduce the theoretical foundations of scaling laws in reinforcement learning and show that the asymptotic performance of reinforcement learning algorithms does not have a monotone relationship between performance rankings and data-regimes. We conduct large-scale experiments and our results demonstrate that a line of reinforcement learning research under the canonical design and evaluation paradigms resulted in incorrect conclusions. Our analysis and results provide a core analysis on scaling, capacity and complexity of deep reinforcement learning.