R^3-Bench: 대규모 언어 모델은 공유 예산 하에서의 자원 합리적 추론에 어려움을 겪는다
R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets
August 17, 2026
저자: Peisong Wang, Zhiwei Ma, Bowen Liu, Feixue Liu, Aochuan Chen, Chenyi Zi, Hongchuan Zeng, Yuhan Li, Jia Li
cs.AI
초록
인지 과학에서 자원 합리성(resource rationality)은 에이전트가 기대 가치를 최대화하기 위해 제한된 계산을 어떻게 할당해야 하는지를 묻는다. 대부분의 추론 및 에이전트 벤치마크는 작업별 독립 예산을 사용하며, 기존의 공유 예산 연구는 스위트 성능을 동일 모델이 단일 문제에서 입증한 능력과 대비하여 보정하지 않는다. 우리는 수학, 경쟁 프로그래밍, 추상 추론에 걸쳐 도구 미사용 및 에이전트 환경에서 공유 예산 하에 6문제 스위트를 평가하는 R^3-벤치를 도입한다. 정합된 단일 문제 응답 곡선은 관측된 성공들에 대한 오프라인 경험적 오라클을 정의한다. 6개 모델에 대한 72개 주요 표 셀에서 오라클 평균은 모든 셀에서 대회 평균과 같거나 초과하며, 71개 셀에서 엄격히 더 높다. 적당한 도구 미사용 압력 하에서 동등 할당 재실행도 6개 모델 중 4개에서 대회 성능을 초과한다. 궤적 진단은 제한된 전략 갱신과 압력 의존적 실패 패턴을 드러낸다. 강한 에이전트 압력 하의 3모델 진단에서는 9개 셀 중 6개에서 적어도 하나의 고정 스케줄러가 대회 평균을 초과하지만, 어떤 정책도 도메인 전반에서 지배하지 않는다. 이러한 결과는 입증된 능력과 공유 예산 실현 사이의 지속적인 격차를 노출한다.
English
In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce R^3-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.