Papers
arxiv:2608.16033

R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

Published on Aug 17
· Submitted by
PSWang
on Aug 18
Authors:
,
,
,
,
,
,
,

Abstract

R³-Bench reveals that shared computation budgets cause reasoning agents to underperform relative to their single-problem capabilities across math, coding, and abstract reasoning tasks.

In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce R^3-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.

Community

Paper submitter

Can LLMs spend a shared reasoning budget wisely?

  • R³-Bench: A benchmark for resource-rational reasoning under shared computational budgets.
  • Six-problem contests: Math, competitive programming, and abstract reasoning tasks compete for one shared token or action budget.
  • Competence vs. allocation: A same-model empirical oracle compares contest performance with each model's demonstrated single-problem capability.
  • Unrealized headroom in 71/72 cells: The oracle strictly outperforms the model's own contest allocation in 71 main-result cells and ties in one.
  • Tools are not enough: Online strategy updates remain limited; lightweight schedulers help in 6/9 strong-pressure model-domain cases, but no scheduler works universally.

Paper: https://arxiv.org/abs/2608.16033
Code: https://github.com/NineAbyss/R-3-Bench
Dataset: https://huggingface.co/datasets/R-3-Bench/R-3-Bench

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.16033 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.16033 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.