<video src=\"https://cdn-uploads.huggingface.co/production/uploads/655452b8432af1b1116394d1/CUIc4sJ4iR-aKg0FHL3ly.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n\n<p>Harness scaling, a different way to scale agent performance.<br>On Terminal-Bench 2.1, StateM raises GPT-5.5 xhigh to 92.1%, versus 83.1% reference and GPT-5.6 Sol Ultra at 91.9%. The runbook transfers unchanged to GPT-5.6. With GPT-5.6 Sol xhigh, StateM reaches 95.3% raw accuracy across 445 trials and succeeds on all 89 tasks at least once.</p>\n","updatedAt":"2026-08-18T17:15:51.726Z","author":{"_id":"655452b8432af1b1116394d1","avatarUrl":"/avatars/85860fb3c2d09c9c23e7677d7129cca3.svg","fullname":"Kai Wang","name":"VictorKai1996NUS","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8062605857849121},"editors":["VictorKai1996NUS"],"editorAvatarUrls":["/avatars/85860fb3c2d09c9c23e7677d7129cca3.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.15089","authors":[{"_id":"6a848ed0536bdd3bdd48f61f","name":"Ziheng Qin","hidden":false},{"_id":"6a848ed0536bdd3bdd48f620","name":"Yaxin Lu","hidden":false},{"_id":"6a848ed0536bdd3bdd48f621","name":"Zhangyang Atlas Wang","hidden":false},{"_id":"6a848ed0536bdd3bdd48f622","name":"Kai Wang","hidden":false}],"publishedAt":"2026-08-15T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"StateM: Reaching 95.3% Raw Accuracy, or a \\$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling","submittedOnDailyBy":{"_id":"655452b8432af1b1116394d1","avatarUrl":"/avatars/85860fb3c2d09c9c23e7677d7129cca3.svg","isPro":false,"fullname":"Kai Wang","user":"VictorKai1996NUS","type":"user","name":"VictorKai1996NUS"},"summary":"Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track of mutable state, fail to reactivate lessons from earlier executions, skip known procedures, or stop prematurely. We bet on harness scaling to improve the execution system around an agent without changing its model weights. We introduce StateM, an agent-native runtime that organizes execution around durable states, phase-local context, checked transitions, recoverable runbooks, and versioned procedural practices that agents and users can inspect together.\n On Terminal-Bench 2.1, StateM raises GPT-5.5 xhigh to 92.1\\%, versus 83.1\\% reference and GPT-5.6 Sol Ultra at 91.9\\%. The runbook transfers unchanged to GPT-5.6. With GPT-5.6 Sol xhigh, StateM reaches 95.3\\% raw accuracy across 445 trials and succeeds on all 89 tasks at least once. The frozen profile raises GPT-5.6 Luna from 76.7 to 85.4\\%, above the 84.9\\% Sol xhigh reference.\n Using the same runtime, runbook structure, and golden rules, less than \\38 of adaptation raises DeepSeek-V4 Flash from 82.7 to 88.1% under standard timeouts and to 89.1% on an 88-task common core. Extending only the remaining latency-sensitive task matches the reported 88.8% GPT-5.6 Sol max result. Final-score API usage is about 15 versus \\574.68 for the GPT reference; total DeepSeek expenditure is 52.22.\n On BusinessBench, family-specific runbooks built on development sets yield held-out gains of 0.55 macro and 1.34 micro points; two mechanism-matched families improve by 10.04 points. Concrete rules generalize when tasks share execution structure, while the control methodology applies broadly. StateM turns selected postmortem findings into persistent, executable preconditions and practices, making learned controls explicit and enforceable through stateful controls. Code at github.com/henryqin1997/statem.","upvotes":32,"discussionId":"6a848ed1536bdd3bdd48f623","projectPage":"https://henryqin1997.github.io/statem/","githubRepo":"https://github.com/henryqin1997/statem","githubRepoAddedBy":"user","ai_summary":"StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weights.","ai_keywords":["agent-native runtime","durable states","phase-local context","checked transitions","recoverable runbooks","versioned procedural practices","StateM","Terminal-Bench","runbook transfer","frozen profile","DeepSeek-V4 Flash","BusinessBench","mechanism-matched families","stateful controls"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":49},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"674c62293aaa8643615f3f57","avatarUrl":"/avatars/87844cbcc80b5f1bcbfb4bfc88bec4cc.svg","isPro":false,"fullname":"Qin Ziheng","user":"zihengqin","type":"user"},{"_id":"64f5937db6d7050b19c68fec","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64f5937db6d7050b19c68fec/4lWYw6-VxZsbl-1Rd9FSy.jpeg","isPro":false,"fullname":"Xuanlei Zhao","user":"oahzxl","type":"user"},{"_id":"68e5b99fefde515b12716717","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/r4u3k9o_QTKYcza1xFjQY.png","isPro":false,"fullname":"Moyang Li","user":"moyangli","type":"user"},{"_id":"6535759efca2c10e430e2df2","avatarUrl":"/avatars/2f3f33b4827d7caf387caaf9d115f4d5.svg","isPro":false,"fullname":"Yong Liu","user":"sglucas","type":"user"},{"_id":"6440cdd3757aa3c2ad88c024","avatarUrl":"/avatars/7c67e91bab181c0c94b1ab3667a0aa42.svg","isPro":false,"fullname":"Jianing Zhu","user":"Zfancy","type":"user"},{"_id":"655452b8432af1b1116394d1","avatarUrl":"/avatars/85860fb3c2d09c9c23e7677d7129cca3.svg","isPro":false,"fullname":"Kai Wang","user":"VictorKai1996NUS","type":"user"},{"_id":"68bf0b34d99178b62f126a72","avatarUrl":"/avatars/3a0a00cb6a556c1487718166108e17b3.svg","isPro":false,"fullname":"Lucy Mochizuki","user":"lucymochix","type":"user"},{"_id":"66fbb64725d0b1bf22bf0559","avatarUrl":"/avatars/7a0cffc7d15c755ff23a1452f32aa41d.svg","isPro":false,"fullname":"Cora Combe","user":"cocome","type":"user"},{"_id":"65ca358f7cd01b4c7a2ecbb4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ca358f7cd01b4c7a2ecbb4/hve36WXU_MAOd9BhR_amP.jpeg","isPro":false,"fullname":"Liu Jiaxin","user":"liuplus","type":"user"},{"_id":"66fbb3387d19fc194671e8e6","avatarUrl":"/avatars/3bee4ace97643cc0e0843b1693227b53.svg","isPro":false,"fullname":"Angela Fort","user":"angle49","type":"user"},{"_id":"684480b74a480fb947fce601","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/684480b74a480fb947fce601/zm3OrMihRjTDQrj7VDSBn.jpeg","isPro":false,"fullname":"Noel de Boxer","user":"Noliboxer","type":"user"},{"_id":"66b258651ae5b88b8953819a","avatarUrl":"/avatars/b2529f610b7fde8dc215f7c04f817b35.svg","isPro":false,"fullname":"my pull","user":"pullonly","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"query":{}}">
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
Abstract
StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weights.
Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track of mutable state, fail to reactivate lessons from earlier executions, skip known procedures, or stop prematurely. We bet on harness scaling to improve the execution system around an agent without changing its model weights. We introduce StateM, an agent-native runtime that organizes execution around durable states, phase-local context, checked transitions, recoverable runbooks, and versioned procedural practices that agents and users can inspect together.
On Terminal-Bench 2.1, StateM raises GPT-5.5 xhigh to 92.1\%, versus 83.1\% reference and GPT-5.6 Sol Ultra at 91.9\%. The runbook transfers unchanged to GPT-5.6. With GPT-5.6 Sol xhigh, StateM reaches 95.3\% raw accuracy across 445 trials and succeeds on all 89 tasks at least once. The frozen profile raises GPT-5.6 Luna from 76.7 to 85.4\%, above the 84.9\% Sol xhigh reference.
Using the same runtime, runbook structure, and golden rules, less than \38 of adaptation raises DeepSeek-V4 Flash from 82.7 to 88.1% under standard timeouts and to 89.1% on an 88-task common core. Extending only the remaining latency-sensitive task matches the reported 88.8% GPT-5.6 Sol max result. Final-score API usage is about 15 versus \574.68 for the GPT reference; total DeepSeek expenditure is 52.22.
On BusinessBench, family-specific runbooks built on development sets yield held-out gains of 0.55 macro and 1.34 micro points; two mechanism-matched families improve by 10.04 points. Concrete rules generalize when tasks share execution structure, while the control methodology applies broadly. StateM turns selected postmortem findings into persistent, executable preconditions and practices, making learned controls explicit and enforceable through stateful controls. Code at github.com/henryqin1997/statem.
Community
Harness scaling, a different way to scale agent performance.
On Terminal-Bench 2.1, StateM raises GPT-5.5 xhigh to 92.1%, versus 83.1% reference and GPT-5.6 Sol Ultra at 91.9%. The runbook transfers unchanged to GPT-5.6. With GPT-5.6 Sol xhigh, StateM reaches 95.3% raw accuracy across 445 trials and succeeds on all 89 tasks at least once.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.15089 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.15089 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.15089 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.