ChatPaper.aiChatPaper

R^3-Bench:大規模言語モデルは共有予算下での資源合理的推論を苦手とする

R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

August 17, 2026
著者: Peisong Wang, Zhiwei Ma, Bowen Liu, Feixue Liu, Aochuan Chen, Chenyi Zi, Hongchuan Zeng, Yuhan Li, Jia Li
cs.AI

要旨

認知科学において、資源合理性は、エージェントが限られた計算資源をどのように配分して期待価値を最大化すべきかを問うものである。ほとんどの推論およびエージェントベンチマークは、タスクごとに独立した予算を用いるが、既存の共有予算研究は、同一モデルが単一問題で示した能力に対してスイート性能を較正していない。我々はR^3-Benchを導入する。これは、数学、競技プログラミング、抽象推論にわたる6問題からなるスイートを、共有予算の下で、ツールなし設定とエージェント設定の両方で評価するものである。対応付けられた単一問題の応答曲線は、観測された成功に対するオフラインの経験的オラクルを定義する。6モデルについて主要テーブルの72セル全体で、オラクル平均は全てのセルにおいてコンテスト平均と一致するか上回り、71セルでは厳密に高い。中程度のツールなし圧力の下では、均等配分リプレイも6モデルのうち4モデルでコンテスト性能を上回る。軌跡診断は、限られた戦略更新と圧力依存の失敗パターンを明らかにする。強いエージェント圧力の下での3モデル診断では、少なくとも1つの固定スケジューラが9セルのうち6セルでコンテスト平均を上回るが、どの方針も領域全体を支配しない。これらの結果は、実証された能力と共有予算の実現との間に持続的なギャップが存在することを示している。
English
In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce R^3-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.