ChatPaper.aiChatPaper

深層強化学習の評価と設計パラダイムの原理的分析

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

July 8, 2026
著者: Ezgi Korkmaz
cs.AI

要旨

深層ニューラルネットワークを用いた状態行動価値関数の近似により、最も困難なゲームの一つでの勝利を達成したことから、問題のルールを明示せずとも課題を解決するアルゴリズムの進歩に至るまで、強化学習研究は過去10年間にわたり顕著な科学的進歩の中心に位置してきた。本稿では、この研究進展の主要な要素に焦点を当て、強化学習における標準的な評価と設計のパラダイムを分析する。強化学習におけるスケーリング則の理論的基礎を導入し、強化学習アルゴリズムの漸近的性能において、性能順位とデータ規模の間に単調な関係が存在しないことを示す。我々は大規模実験を実施し、その結果は、標準的な設計・評価パラダイムの下での一連の強化学習研究が誤った結論を導いたことを実証する。本分析と結果は、深層強化学習のスケーリング、容量、複雑性に関する中核的な分析を提供する。
English
Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even explicitly stating the rules of the challenge at hand, reinforcement learning research has been the center of remarkable scientific progress for the past decade. In this paper, we focus on the key ingredients of this research progress and we analyze the canonical evaluation and design paradigms in reinforcement learning. We introduce the theoretical foundations of scaling laws in reinforcement learning and show that the asymptotic performance of reinforcement learning algorithms does not have a monotone relationship between performance rankings and data-regimes. We conduct large-scale experiments and our results demonstrate that a line of reinforcement learning research under the canonical design and evaluation paradigms resulted in incorrect conclusions. Our analysis and results provide a core analysis on scaling, capacity and complexity of deep reinforcement learning.