Hugging Face Daily Papers · · 4 min read

R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Can LLMs spend a shared reasoning budget wisely?</p>\n<ul>\n<li><strong>R³-Bench</strong>: A benchmark for resource-rational reasoning under shared computational budgets.</li>\n<li><strong>Six-problem contests</strong>: Math, competitive programming, and abstract reasoning tasks compete for one shared token or action budget.</li>\n<li><strong>Competence vs. allocation</strong>: A same-model empirical oracle compares contest performance with each model's demonstrated single-problem capability.</li>\n<li><strong>Unrealized headroom in 71/72 cells</strong>: The oracle strictly outperforms the model's own contest allocation in 71 main-result cells and ties in one.</li>\n<li><strong>Tools are not enough</strong>: Online strategy updates remain limited; lightweight schedulers help in 6/9 strong-pressure model-domain cases, but no scheduler works universally.</li>\n</ul>\n<p>Paper: <a href=\"https://arxiv.org/abs/2608.16033\" rel=\"nofollow\">https://arxiv.org/abs/2608.16033</a><br>Code: <a href=\"https://github.com/NineAbyss/R-3-Bench\" rel=\"nofollow\">https://github.com/NineAbyss/R-3-Bench</a><br>Dataset: <a href=\"https://huggingface.co/datasets/R-3-Bench/R-3-Bench\">https://huggingface.co/datasets/R-3-Bench/R-3-Bench</a></p>\n","updatedAt":"2026-08-18T02:40:01.262Z","author":{"_id":"64be4408c05a0df0d2b6012e","avatarUrl":"/avatars/09d8427505a418090391dc5a3f8bfef2.svg","fullname":"PSWang","name":"CedarWang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7570841312408447},"editors":["CedarWang"],"editorAvatarUrls":["/avatars/09d8427505a418090391dc5a3f8bfef2.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.16033","authors":[{"_id":"6a83c409675db694db8cd481","name":"Peisong Wang","hidden":false},{"_id":"6a83c409675db694db8cd482","user":{"_id":"68410d555a3b810d2ad6125d","avatarUrl":"/avatars/7aa0cc17a5ceda176e494cc64a35284c.svg","isPro":false,"fullname":"Zhiwei Ma","user":"libliot","type":"user","name":"libliot"},"name":"Zhiwei Ma","status":"claimed_verified","statusLastChangedAt":"2026-08-18T08:45:05.114Z","hidden":false},{"_id":"6a83c409675db694db8cd483","name":"Bowen Liu","hidden":false},{"_id":"6a83c409675db694db8cd484","name":"Feixue Liu","hidden":false},{"_id":"6a83c409675db694db8cd485","name":"Aochuan Chen","hidden":false},{"_id":"6a83c409675db694db8cd486","name":"Chenyi Zi","hidden":false},{"_id":"6a83c409675db694db8cd487","name":"Hongchuan Zeng","hidden":false},{"_id":"6a83c409675db694db8cd488","name":"Yuhan Li","hidden":false},{"_id":"6a83c409675db694db8cd489","name":"Jia Li","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64be4408c05a0df0d2b6012e/Iv5USxvk51osCZEBrrt_B.png","https://cdn-uploads.huggingface.co/production/uploads/64be4408c05a0df0d2b6012e/iNs-kfhmApn_6-CQjSW7U.jpeg"],"publishedAt":"2026-08-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets","submittedOnDailyBy":{"_id":"64be4408c05a0df0d2b6012e","avatarUrl":"/avatars/09d8427505a418090391dc5a3f8bfef2.svg","isPro":false,"fullname":"PSWang","user":"CedarWang","type":"user","name":"CedarWang"},"summary":"In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce R^3-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.","upvotes":11,"discussionId":"6a83c409675db694db8cd48a","projectPage":"https://github.com/NineAbyss/R-3-Bench","githubRepo":"https://github.com/NineAbyss/R-3-Bench","githubRepoAddedBy":"user","ai_summary":"R³-Bench reveals that shared computation budgets cause reasoning agents to underperform relative to their single-problem capabilities across math, coding, and abstract reasoning tasks.","ai_keywords":["resource rationality","R³-Bench","shared budgets","response curves","empirical oracle","strategy updating","fixed scheduler"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"68909d0ec75ddecb2867fcba","name":"HKUST-DSAIL","fullname":"Data Science & Artificial Intelligence Lab at HKUST","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/657dc533aed79486aeffbf55/biw8iB5cUADFtw0_a6LCM.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a5f28c9c3daab50c7ff9f76","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a5f28c9c3daab50c7ff9f76/M-c-NBoMcVsxFJcok0LGF.png","isPro":false,"fullname":"R^3Bench","user":"R-3-Bench","type":"user"},{"_id":"68410d555a3b810d2ad6125d","avatarUrl":"/avatars/7aa0cc17a5ceda176e494cc64a35284c.svg","isPro":false,"fullname":"Zhiwei Ma","user":"libliot","type":"user"},{"_id":"67e8f2845c0900f17bfbb5fd","avatarUrl":"/avatars/9c70427411a5bf43c4a68e11664c3f80.svg","isPro":false,"fullname":"Felicia","user":"ffeeiik","type":"user"},{"_id":"68676ddbb1bd4bd32016d63b","avatarUrl":"/avatars/f9648f8390dc73292ced1a60d5f6d100.svg","isPro":false,"fullname":"RLVER","user":"RLVER","type":"user"},{"_id":"67b80da55e1b74491fc08eb9","avatarUrl":"/avatars/991baa07ce8792b2a04f30e27aabf786.svg","isPro":false,"fullname":"S2R","user":"S2R-data","type":"user"},{"_id":"68d6591428e169473e936a31","avatarUrl":"/avatars/e926e835b6c6f3989c81faa233da55fc.svg","isPro":false,"fullname":"NSPO","user":"ICLR2026NSPO","type":"user"},{"_id":"6a83c7344b4c3155571e91d0","avatarUrl":"/avatars/9a5bca3e984a95e62a13eae3eb36b3a5.svg","isPro":false,"fullname":"ANG","user":"CdricW","type":"user"},{"_id":"64e408524b78ab05968a8493","avatarUrl":"/avatars/6d987ca56eb9c2aa23bf2216f198b991.svg","isPro":false,"fullname":"chenyi","user":"BarristanZi","type":"user"},{"_id":"66bac1248ba60a7a91ba0836","avatarUrl":"/avatars/a92e6fd8b3104842f34f7ceaf5ceaa5f.svg","isPro":false,"fullname":"Wang Yuyao","user":"littlewyy","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"5f32b2367e583543386214d9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1635314457124-5f32b2367e583543386214d9.jpeg","isPro":true,"fullname":"Sergei Averkiev","user":"averoo","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68909d0ec75ddecb2867fcba","name":"HKUST-DSAIL","fullname":"Data Science & Artificial Intelligence Lab at HKUST","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/657dc533aed79486aeffbf55/biw8iB5cUADFtw0_a6LCM.png"},"query":{}}">
Papers
arxiv:2608.16033

R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

Published on Aug 17
· Submitted by
PSWang
on Aug 18
Authors:
,

Abstract

R³-Bench reveals that shared computation budgets cause reasoning agents to underperform relative to their single-problem capabilities across math, coding, and abstract reasoning tasks.

In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce R^3-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.

Community

Paper submitter about 6 hours ago

Can LLMs spend a shared reasoning budget wisely?

  • R³-Bench: A benchmark for resource-rational reasoning under shared computational budgets.
  • Six-problem contests: Math, competitive programming, and abstract reasoning tasks compete for one shared token or action budget.
  • Competence vs. allocation: A same-model empirical oracle compares contest performance with each model's demonstrated single-problem capability.
  • Unrealized headroom in 71/72 cells: The oracle strictly outperforms the model's own contest allocation in 71 main-result cells and ties in one.
  • Tools are not enough: Online strategy updates remain limited; lightweight schedulers help in 6/9 strong-pressure model-domain cases, but no scheduler works universally.

Paper: https://arxiv.org/abs/2608.16033
Code: https://github.com/NineAbyss/R-3-Bench
Dataset: https://huggingface.co/datasets/R-3-Bench/R-3-Bench

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.16033 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.16033 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers