Code: <a href=\"https://github.com/microsoft/thinkingbox\" rel=\"nofollow\">https://github.com/microsoft/thinkingbox</a></p>\n","updatedAt":"2026-08-25T07:06:11.729Z","author":{"_id":"5f1158120c833276f61f1a84","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1608042047613-5f1158120c833276f61f1a84.jpeg","fullname":"Niels Rogge","name":"nielsr","type":"user","isPro":false,"isHf":true,"isHfAdmin":false,"isMod":false,"followerCount":1284,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7292873859405518},"editors":["nielsr"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1608042047613-5f1158120c833276f61f1a84.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.19741","authors":[{"_id":"6a8968463d26296ea30914dc","user":{"_id":"65550ee2f4f9996af318cd36","avatarUrl":"/avatars/1da67f8e1ab5ecde66e99aae00b42a0b.svg","isPro":false,"fullname":"Zhuochun Li","user":"zhuochun","type":"user","name":"zhuochun"},"name":"Zhuochun Li","status":"claimed_verified","statusLastChangedAt":"2026-08-25T00:45:04.208Z","hidden":false},{"_id":"6a8968463d26296ea30914dd","user":{"_id":"6459bf907c80a6105b6e17fc","avatarUrl":"/avatars/821997acf79b7113de576661b1111109.svg","isPro":false,"fullname":"Young Ko","user":"youngko","type":"user","name":"youngko"},"name":"Youngmin Ko","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:14:43.862Z","hidden":false},{"_id":"6a8968463d26296ea30914de","user":{"_id":"6a1e1d1dcf8206a95e45c047","avatarUrl":"/avatars/bee542722d2265f902026c3a0f85096e.svg","isPro":false,"fullname":"Ali Keramati","user":"a-kera","type":"user","name":"a-kera"},"name":"Ali Keramati","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:14:41.822Z","hidden":false},{"_id":"6a8968463d26296ea30914df","name":"Nicola Ferri","hidden":false},{"_id":"6a8968463d26296ea30914e0","name":"Susana Palmaz Lopez Pelaez","hidden":false},{"_id":"6a8968463d26296ea30914e1","user":{"_id":"6a88f9ae330c1bb6659fbd04","avatarUrl":"/avatars/e81a7eb0a351850271d22f183de9458a.svg","isPro":false,"fullname":"Liang-Chun Tsai","user":"ltsai-dev","type":"user","name":"ltsai-dev"},"name":"Liang-Chun Tsai","status":"claimed_verified","statusLastChangedAt":"2026-08-25T00:45:04.200Z","hidden":false},{"_id":"6a8968463d26296ea30914e2","name":"Calvin Wang","hidden":false},{"_id":"6a8968463d26296ea30914e3","name":"Mirco Milletari","hidden":false},{"_id":"6a8968463d26296ea30914e4","user":{"_id":"64b8491203124195cd795cad","avatarUrl":"/avatars/75b0252a17af0de0bd6727f3290577f9.svg","isPro":false,"fullname":"Tuhin Kundu","user":"tuhink","type":"user","name":"tuhink"},"name":"Tuhin Kundu","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:15:06.097Z","hidden":false},{"_id":"6a8968463d26296ea30914e5","name":"Vadim Smolyakov","hidden":false},{"_id":"6a8968463d26296ea30914e6","name":"Kjartan Olafsson","hidden":false},{"_id":"6a8968463d26296ea30914e7","name":"Tommy Guy","hidden":false}],"publishedAt":"2026-08-20T00:00:00.000Z","submittedOnDailyAt":"2026-08-25T00:00:00.000Z","title":"One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows","submittedOnDailyBy":{"_id":"5f1158120c833276f61f1a84","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1608042047613-5f1158120c833276f61f1a84.jpeg","isPro":false,"fullname":"Niels Rogge","user":"nielsr","type":"user","name":"nielsr"},"summary":"Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agents must gather missing information over multiple turns, follow domain policies, coordinate dependent tools, and realize the correct persistent state transition without collateral effects. In this paper, we introduce Thinkingbox, a sandbox for tool-agent-user interaction that provides isolated MCP-compatible tool sessions, complete execution traces, and outcome evaluation over terminal backend state. Built on this sandbox, Thinkingbox-bench contains 507 policy-conditioned workflows across numerous scenarios, including retail, hospitality, auto insurance, neobank internal IT, and consulting IT/HR support. Each attempt is evaluated by task-specific executable checks that accept valid trajectories while rejecting wrong, missing, or extra effects; designated tasks additionally check required properties of the final response. Across proprietary and open-weight models, the strongest achieves 65.36% pass@1, but only 25.25% pass^20. Moreover, many failed trials show clean termination and valid state-changing actions, showing that response or tool-call-level signals are not clear proxies for end-to-end task completion. Thinkingbox-bench reveals a large gap between occasionally finding a successful trajectory and reliably completing stateful business tasks. We release both Thinkingbox and Thinkingbox-Bench: https://github.com/microsoft/thinkingbox","upvotes":6,"discussionId":"6a8968473d26296ea30914e8","githubRepo":"https://github.com/microsoft/thinkingbox","githubRepoAddedBy":"admin","githubStars":19,"organization":{"_id":"5e6485f787403103f9f1055e","name":"microsoft","fullname":"Microsoft","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583646260758-5e64858c87403103f9f1055d.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"65550ee2f4f9996af318cd36","avatarUrl":"/avatars/1da67f8e1ab5ecde66e99aae00b42a0b.svg","isPro":false,"fullname":"Zhuochun Li","user":"zhuochun","type":"user"},{"_id":"6459bf907c80a6105b6e17fc","avatarUrl":"/avatars/821997acf79b7113de576661b1111109.svg","isPro":false,"fullname":"Young Ko","user":"youngko","type":"user"},{"_id":"64b8491203124195cd795cad","avatarUrl":"/avatars/75b0252a17af0de0bd6727f3290577f9.svg","isPro":false,"fullname":"Tuhin Kundu","user":"tuhink","type":"user"},{"_id":"6a88f9ae330c1bb6659fbd04","avatarUrl":"/avatars/e81a7eb0a351850271d22f183de9458a.svg","isPro":false,"fullname":"Liang-Chun Tsai","user":"ltsai-dev","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"5e6485f787403103f9f1055e","name":"microsoft","fullname":"Microsoft","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583646260758-5e64858c87403103f9f1055d.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.19741.md","query":{}}">
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows
Abstract
Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agents must gather missing information over multiple turns, follow domain policies, coordinate dependent tools, and realize the correct persistent state transition without collateral effects. In this paper, we introduce Thinkingbox, a sandbox for tool-agent-user interaction that provides isolated MCP-compatible tool sessions, complete execution traces, and outcome evaluation over terminal backend state. Built on this sandbox, Thinkingbox-bench contains 507 policy-conditioned workflows across numerous scenarios, including retail, hospitality, auto insurance, neobank internal IT, and consulting IT/HR support. Each attempt is evaluated by task-specific executable checks that accept valid trajectories while rejecting wrong, missing, or extra effects; designated tasks additionally check required properties of the final response. Across proprietary and open-weight models, the strongest achieves 65.36% pass@1, but only 25.25% pass^20. Moreover, many failed trials show clean termination and valid state-changing actions, showing that response or tool-call-level signals are not clear proxies for end-to-end task completion. Thinkingbox-bench reveals a large gap between occasionally finding a successful trajectory and reliably completing stateful business tasks. We release both Thinkingbox and Thinkingbox-Bench: https://github.com/microsoft/thinkingbox
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.19741 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.19741 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.19741 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.