ChatPaper.aiChatPaper

深度強化學習評估與設計範式的原則性分析

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

July 8, 2026
作者: Ezgi Korkmaz
cs.AI

摘要

從利用深度神經網路逼近狀態-動作價值函數,進而攻克極具挑戰性的遊戲,到演算法上的進展使得無須明確陳述問題規則即可解決問題,強化學習研究在過去十年間一直是科學重大進展的核心。本文聚焦於此研究進展的關鍵要素,並分析強化學習中典型的評估與設計典範。我們引入強化學習中規模法則的理論基礎,並證明在漸近性能上,強化學習演算法的效能排名與數據範疇之間不存在單調關係。透過大規模實驗,我們的結果顯示,在典型設計與評估典範下進行的某些強化學習研究路線,導致了錯誤的結論。我們的分析與結果對深度強化學習的規模化、容量及複雜度提供了核心見解。
English
Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even explicitly stating the rules of the challenge at hand, reinforcement learning research has been the center of remarkable scientific progress for the past decade. In this paper, we focus on the key ingredients of this research progress and we analyze the canonical evaluation and design paradigms in reinforcement learning. We introduce the theoretical foundations of scaling laws in reinforcement learning and show that the asymptotic performance of reinforcement learning algorithms does not have a monotone relationship between performance rankings and data-regimes. We conduct large-scale experiments and our results demonstrate that a line of reinforcement learning research under the canonical design and evaluation paradigms resulted in incorrect conclusions. Our analysis and results provide a core analysis on scaling, capacity and complexity of deep reinforcement learning.