Hugging Face Daily Papers · · 5 min read

StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<video src=\"https://cdn-uploads.huggingface.co/production/uploads/655452b8432af1b1116394d1/CUIc4sJ4iR-aKg0FHL3ly.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n\n<p>Harness scaling, a different way to scale agent performance.<br>On Terminal-Bench 2.1, StateM raises GPT-5.5 xhigh to 92.1%, versus 83.1% reference and GPT-5.6 Sol Ultra at 91.9%. The runbook transfers unchanged to GPT-5.6. With GPT-5.6 Sol xhigh, StateM reaches 95.3% raw accuracy across 445 trials and succeeds on all 89 tasks at least once.</p>\n","updatedAt":"2026-08-18T17:15:51.726Z","author":{"_id":"655452b8432af1b1116394d1","avatarUrl":"/avatars/85860fb3c2d09c9c23e7677d7129cca3.svg","fullname":"Kai Wang","name":"VictorKai1996NUS","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8062605857849121},"editors":["VictorKai1996NUS"],"editorAvatarUrls":["/avatars/85860fb3c2d09c9c23e7677d7129cca3.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.15089","authors":[{"_id":"6a848ed0536bdd3bdd48f61f","name":"Ziheng Qin","hidden":false},{"_id":"6a848ed0536bdd3bdd48f620","name":"Yaxin Lu","hidden":false},{"_id":"6a848ed0536bdd3bdd48f621","name":"Zhangyang Atlas Wang","hidden":false},{"_id":"6a848ed0536bdd3bdd48f622","name":"Kai Wang","hidden":false}],"publishedAt":"2026-08-15T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"StateM: Reaching 95.3% Raw Accuracy, or a \\$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling","submittedOnDailyBy":{"_id":"655452b8432af1b1116394d1","avatarUrl":"/avatars/85860fb3c2d09c9c23e7677d7129cca3.svg","isPro":false,"fullname":"Kai Wang","user":"VictorKai1996NUS","type":"user","name":"VictorKai1996NUS"},"summary":"Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track of mutable state, fail to reactivate lessons from earlier executions, skip known procedures, or stop prematurely. We bet on harness scaling to improve the execution system around an agent without changing its model weights. We introduce StateM, an agent-native runtime that organizes execution around durable states, phase-local context, checked transitions, recoverable runbooks, and versioned procedural practices that agents and users can inspect together.\n On Terminal-Bench 2.1, StateM raises GPT-5.5 xhigh to 92.1\\%, versus 83.1\\% reference and GPT-5.6 Sol Ultra at 91.9\\%. The runbook transfers unchanged to GPT-5.6. With GPT-5.6 Sol xhigh, StateM reaches 95.3\\% raw accuracy across 445 trials and succeeds on all 89 tasks at least once. The frozen profile raises GPT-5.6 Luna from 76.7 to 85.4\\%, above the 84.9\\% Sol xhigh reference.\n Using the same runtime, runbook structure, and golden rules, less than \\38 of adaptation raises DeepSeek-V4 Flash from 82.7 to 88.1% under standard timeouts and to 89.1% on an 88-task common core. Extending only the remaining latency-sensitive task matches the reported 88.8% GPT-5.6 Sol max result. Final-score API usage is about 15 versus \\574.68 for the GPT reference; total DeepSeek expenditure is 52.22.\n On BusinessBench, family-specific runbooks built on development sets yield held-out gains of 0.55 macro and 1.34 micro points; two mechanism-matched families improve by 10.04 points. Concrete rules generalize when tasks share execution structure, while the control methodology applies broadly. StateM turns selected postmortem findings into persistent, executable preconditions and practices, making learned controls explicit and enforceable through stateful controls. Code at github.com/henryqin1997/statem.","upvotes":32,"discussionId":"6a848ed1536bdd3bdd48f623","projectPage":"https://henryqin1997.github.io/statem/","githubRepo":"https://github.com/henryqin1997/statem","githubRepoAddedBy":"user","ai_summary":"StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weights.","ai_keywords":["agent-native runtime","durable states","phase-local context","checked transitions","recoverable runbooks","versioned procedural practices","StateM","Terminal-Bench","runbook transfer","frozen profile","DeepSeek-V4 Flash","BusinessBench","mechanism-matched families","stateful controls"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":49},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"674c62293aaa8643615f3f57","avatarUrl":"/avatars/87844cbcc80b5f1bcbfb4bfc88bec4cc.svg","isPro":false,"fullname":"Qin Ziheng","user":"zihengqin","type":"user"},{"_id":"64f5937db6d7050b19c68fec","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64f5937db6d7050b19c68fec/4lWYw6-VxZsbl-1Rd9FSy.jpeg","isPro":false,"fullname":"Xuanlei Zhao","user":"oahzxl","type":"user"},{"_id":"68e5b99fefde515b12716717","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/r4u3k9o_QTKYcza1xFjQY.png","isPro":false,"fullname":"Moyang Li","user":"moyangli","type":"user"},{"_id":"6535759efca2c10e430e2df2","avatarUrl":"/avatars/2f3f33b4827d7caf387caaf9d115f4d5.svg","isPro":false,"fullname":"Yong Liu","user":"sglucas","type":"user"},{"_id":"6440cdd3757aa3c2ad88c024","avatarUrl":"/avatars/7c67e91bab181c0c94b1ab3667a0aa42.svg","isPro":false,"fullname":"Jianing Zhu","user":"Zfancy","type":"user"},{"_id":"655452b8432af1b1116394d1","avatarUrl":"/avatars/85860fb3c2d09c9c23e7677d7129cca3.svg","isPro":false,"fullname":"Kai Wang","user":"VictorKai1996NUS","type":"user"},{"_id":"68bf0b34d99178b62f126a72","avatarUrl":"/avatars/3a0a00cb6a556c1487718166108e17b3.svg","isPro":false,"fullname":"Lucy Mochizuki","user":"lucymochix","type":"user"},{"_id":"66fbb64725d0b1bf22bf0559","avatarUrl":"/avatars/7a0cffc7d15c755ff23a1452f32aa41d.svg","isPro":false,"fullname":"Cora Combe","user":"cocome","type":"user"},{"_id":"65ca358f7cd01b4c7a2ecbb4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ca358f7cd01b4c7a2ecbb4/hve36WXU_MAOd9BhR_amP.jpeg","isPro":false,"fullname":"Liu Jiaxin","user":"liuplus","type":"user"},{"_id":"66fbb3387d19fc194671e8e6","avatarUrl":"/avatars/3bee4ace97643cc0e0843b1693227b53.svg","isPro":false,"fullname":"Angela Fort","user":"angle49","type":"user"},{"_id":"684480b74a480fb947fce601","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/684480b74a480fb947fce601/zm3OrMihRjTDQrj7VDSBn.jpeg","isPro":false,"fullname":"Noel de Boxer","user":"Noliboxer","type":"user"},{"_id":"66b258651ae5b88b8953819a","avatarUrl":"/avatars/b2529f610b7fde8dc215f7c04f817b35.svg","isPro":false,"fullname":"my pull","user":"pullonly","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"query":{}}">
Papers
arxiv:2608.15089

StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

Published on Aug 15
· Submitted by
Kai Wang
on Aug 18
Authors:
,

Abstract

StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weights.

Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track of mutable state, fail to reactivate lessons from earlier executions, skip known procedures, or stop prematurely. We bet on harness scaling to improve the execution system around an agent without changing its model weights. We introduce StateM, an agent-native runtime that organizes execution around durable states, phase-local context, checked transitions, recoverable runbooks, and versioned procedural practices that agents and users can inspect together. On Terminal-Bench 2.1, StateM raises GPT-5.5 xhigh to 92.1\%, versus 83.1\% reference and GPT-5.6 Sol Ultra at 91.9\%. The runbook transfers unchanged to GPT-5.6. With GPT-5.6 Sol xhigh, StateM reaches 95.3\% raw accuracy across 445 trials and succeeds on all 89 tasks at least once. The frozen profile raises GPT-5.6 Luna from 76.7 to 85.4\%, above the 84.9\% Sol xhigh reference. Using the same runtime, runbook structure, and golden rules, less than \38 of adaptation raises DeepSeek-V4 Flash from 82.7 to 88.1% under standard timeouts and to 89.1% on an 88-task common core. Extending only the remaining latency-sensitive task matches the reported 88.8% GPT-5.6 Sol max result. Final-score API usage is about 15 versus \574.68 for the GPT reference; total DeepSeek expenditure is 52.22. On BusinessBench, family-specific runbooks built on development sets yield held-out gains of 0.55 macro and 1.34 micro points; two mechanism-matched families improve by 10.04 points. Concrete rules generalize when tasks share execution structure, while the control methodology applies broadly. StateM turns selected postmortem findings into persistent, executable preconditions and practices, making learned controls explicit and enforceable through stateful controls. Code at github.com/henryqin1997/statem.

Community

Harness scaling, a different way to scale agent performance.
On Terminal-Bench 2.1, StateM raises GPT-5.5 xhigh to 92.1%, versus 83.1% reference and GPT-5.6 Sol Ultra at 91.9%. The runbook transfers unchanged to GPT-5.6. With GPT-5.6 Sol xhigh, StateM reaches 95.3% raw accuracy across 445 trials and succeeds on all 89 tasks at least once.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.15089 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.15089 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.15089 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers