ChatPaper.aiChatPaper

R^3-Bench:大语言模型在共享预算条件下的资源理性推理面临挑战

R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

August 17, 2026
作者: Peisong Wang, Zhiwei Ma, Bowen Liu, Feixue Liu, Aochuan Chen, Chenyi Zi, Hongchuan Zeng, Yuhan Li, Jia Li
cs.AI

摘要

在认知科学中,资源理性研究智能体应如何分配有限计算资源以最大化期望值。大多数推理与智能体基准采用独立的单任务预算;现有的共享预算研究未将套件整体性能与同一模型已展示的单问题能力进行校准。我们提出了R^3-Bench,该基准在数学、竞赛编程和抽象推理领域,以无工具和智能体两种设置评估共享预算下的六问题套件。匹配的单问题响应曲线在观测到的成功结果上定义了离线经验oracle。在六个模型的72个主表单元中,oracle均值在所有单元中均达到或超过竞赛均值,且在71个单元中严格更高。在适度的无工具压力下,均匀分配重放在六个模型中的四个上也超过了竞赛性能。轨迹诊断揭示了有限的策略更新和依赖压力的失败模式。在强智能体压力下的三模型诊断中,九个单元中有六个至少存在一个固定调度器超过竞赛均值,但没有任何策略在所有领域中占据主导。这些结果揭示了已展示能力与共享预算实现之间的持续差距。
English
In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce R^3-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.