ChatPaper.aiChatPaper

R^3-Bench:大型語言模型在共享預算下的資源理性推理面臨挑戰

R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

August 17, 2026
作者: Peisong Wang, Zhiwei Ma, Bowen Liu, Feixue Liu, Aochuan Chen, Chenyi Zi, Hongchuan Zeng, Yuhan Li, Jia Li
cs.AI

摘要

在認知科學中,資源理性探討智能體應如何分配有限的計算資源以最大化期望值。大多數推理與智能體基準採用獨立的每任務預算;現有的共享預算研究並未將套件表現與同一模型在單一問題上展現的能力進行校準。我們引入 R^3-Bench,在無工具與智能體設定中,針對數學、競賽程式設計與抽象推理,評估共享預算下的六問題套件。匹配的單一問題反應曲線根據觀測到的成功定義了離線經驗oracle。在六個模型的72個主表格單元中,oracle平均值在所有單元中達到或超過競賽平均值,其中71個單元嚴格更高。在中等無工具壓力下,等額分配重播在六個模型中有四個也超過競賽表現。軌跡診斷揭示了有限的策略更新與依賴壓力的失敗模式。在強智能體壓力下的三模型診斷中,至少一個固定調度器在九個單元中的六個超過競賽平均值,但沒有任何策略在所有領域中佔優。這些結果揭示了已展現能力與共享預算實現之間的持續差距。
English
In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce R^3-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.