We freeze the base model and evolve only the harness — prompts, tools, skills, control flow — by label-free natural selection: variants are scored on measured fitness (avg@k, no gold solutions), survivors are kept, and complementary ones are merged.</p>\n<p>Matched-model, harness alone:</p>\n<ul>\n<li><strong>Terminal-Bench 2.1</strong> 75.5 → <strong>83.2</strong> avg@5 on GPT-5.5, past Codex at 83.1 (84.7 on GPT-5.6 Sol)</li>\n<li><strong>TerminalWorld</strong> 61.0 → <strong>68.3</strong> pass@1 on Opus 4.8, 41 held-out tasks, past Claude Code at 65.9</li>\n<li><strong>WebArena-Infinity</strong> 43.5 → <strong>93.0</strong> audit-clean pass@1 on 1,260 real tasks</li>\n<li><strong>SWE-bench Verified</strong> 80.8 → <strong>84.2</strong> by zero-shot transfer of the terminal harness</li>\n</ul>\n<p>Two results we did not expect. The in-loop proxy saturates (0.505 → 1.000) while held-out pass@1 is 68.3 — a 31.7-point gap — and the variant that best fits the proxy is <em>not</em> the best generalizer, so keeping a population rather than following the single best lineage is what converts an overfit proxy into held-out gain. And we audited every WebArena-Infinity trajectory for validity: the base agent's raw 53.0 falls to 43.5 while ours goes 94.4 → 93.0, so the gap <em>widens</em> under scrutiny instead of narrowing.</p>\n<p>Project page with interactive figures, every number regenerated from the run artifacts: <a href=\"https://huggingface.co/spaces/CoderDoge/darwinx\">https://huggingface.co/spaces/CoderDoge/darwinx</a></p>\n","updatedAt":"2026-08-14T02:53:29.038Z","author":{"_id":"638a765ef560ea995b580cc4","avatarUrl":"/avatars/16dd0363d6c99e35f8a6522a29f4c9b2.svg","fullname":"Yifan Zhang","name":"CoderDoge","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.822934627532959},"editors":["CoderDoge"],"editorAvatarUrls":["/avatars/16dd0363d6c99e35f8a6522a29f4c9b2.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.07545","authors":[{"_id":"6a7da6fd42823931a1f173df","name":"Yifan Zhang","hidden":false},{"_id":"6a7da6fd42823931a1f173e0","name":"Yutong Dai","hidden":false},{"_id":"6a7da6fd42823931a1f173e1","name":"Juntao Tan","hidden":false},{"_id":"6a7da6fd42823931a1f173e2","name":"Luyu Yang","hidden":false},{"_id":"6a7da6fd42823931a1f173e3","name":"Rishi Mullur","hidden":false},{"_id":"6a7da6fd42823931a1f173e4","name":"Thai Hoang","hidden":false},{"_id":"6a7da6fd42823931a1f173e5","name":"Zhiyuan Hu","hidden":false},{"_id":"6a7da6fd42823931a1f173e6","name":"James Zhu","hidden":false},{"_id":"6a7da6fd42823931a1f173e7","name":"Phil Mui","hidden":false},{"_id":"6a7da6fd42823931a1f173e8","name":"Silvio Savarese","hidden":false},{"_id":"6a7da6fd42823931a1f173e9","name":"Ran Xu","hidden":false},{"_id":"6a7da6fd42823931a1f173ea","name":"Zeyuan Chen","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/638a765ef560ea995b580cc4/MUjDLUz6Uyes2wkpKGuz9.png"],"publishedAt":"2026-07-31T00:00:00.000Z","submittedOnDailyAt":"2026-08-14T00:00:00.000Z","title":"DarwinX: Evolving Agent Harnesses Through Natural Selection","submittedOnDailyBy":{"_id":"638a765ef560ea995b580cc4","avatarUrl":"/avatars/16dd0363d6c99e35f8a6522a29f4c9b2.svg","isPro":false,"fullname":"Yifan Zhang","user":"CoderDoge","type":"user","name":"CoderDoge"},"summary":"An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. We introduce DarwinX, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence share one edit interface. Fitness comes from each benchmark's own verifier: no gold solutions, no hand-picked winners. Across four benchmarks that progressively separate the evolution signal from the test, one loop adds about 17 points on average: Terminal-Bench 2.1 rises +7.7 to 83.2% on a matched base and to the verified frontier at 84.7% on a stronger one; TerminalWorld's held-out split reaches 68.3%, ahead of every off-the-shelf agent; WebArena-Infinity real-task pass@1 rises from 43.5% to 93.0% audit-clean; and a Terminal-Bench 2.1 harness transfers unchanged to SWE-bench Verified. What evolves is general agent competence, not benchmark-specific patches, so it survives changes of task, verifier, and base model. A frozen model need not be a fixed agent: harness selection turns evaluation compute into durable capability.","upvotes":39,"discussionId":"6a7da6fe42823931a1f173eb","projectPage":"https://huggingface.co/spaces/CoderDoge/darwinx","ai_summary":"DarwinX evolves agent harnesses via population selection with frozen models, improving verified performance across benchmarks without benchmark-specific patches.","ai_keywords":["LLM agent","harness","self-improvement","population selection","preserve-and-extend contract","archive","recombination","verifier","DarwinX"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"5f6d64475e78cc6b0ed31e4c","name":"Salesforce","fullname":"Salesforce AI Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1602756670970-noauth.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"638a765ef560ea995b580cc4","avatarUrl":"/avatars/16dd0363d6c99e35f8a6522a29f4c9b2.svg","isPro":false,"fullname":"Yifan Zhang","user":"CoderDoge","type":"user"},{"_id":"632405844624707127217677","avatarUrl":"/avatars/af5deba9ec51814bc5ccfd04f68d76d3.svg","isPro":false,"fullname":"Zeyuan Chen","user":"bradypus","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"64c49396bf1954890192bfbf","avatarUrl":"/avatars/b2dbd1d601911440788efcea2a2b77d3.svg","isPro":false,"fullname":"Yutong Dai","user":"UncleFish","type":"user"},{"_id":"65cb46f919683f9817dd6909","avatarUrl":"/avatars/610e7de911497a3b43add837aaedf6f7.svg","isPro":false,"fullname":"rui zhang","user":"frankzhangrui","type":"user"},{"_id":"6465c4c863e7e09dd02e3e1b","avatarUrl":"/avatars/200b029184d2616f98296a2c212f0785.svg","isPro":false,"fullname":"Ran Xu","user":"xurantju","type":"user"},{"_id":"651b35c254134dcaad03758d","avatarUrl":"/avatars/0230e088c0c73184bdcae47240be86c5.svg","isPro":false,"fullname":"Thai Hoang","user":"qthai912uw","type":"user"},{"_id":"6a0763957577a33e8c39da90","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/bRyJvUqc4-ZpKtA_A_bIa.png","isPro":false,"fullname":"Frank Wang","user":"frankwang611","type":"user"},{"_id":"64b682e99a338c25127af067","avatarUrl":"/avatars/2c72eff485c2f265839eb6108e79ef0a.svg","isPro":false,"fullname":"Chen Huang","user":"ChenZiChun","type":"user"},{"_id":"627a124ffe55fa0f8ce0eaf7","avatarUrl":"/avatars/41e0dc029faed6dc45d620c5fe2652a5.svg","isPro":false,"fullname":"Serendipity","user":"Yuhan","type":"user"},{"_id":"65f8b846d761741567ecfe92","avatarUrl":"/avatars/eb09761dfe7ccc94ca09f94cc4e6e6b8.svg","isPro":false,"fullname":"Juntao Tan","user":"chrisjtan23","type":"user"},{"_id":"6a7e971e560102d6b0f425a6","avatarUrl":"/avatars/c61240d7f58683312b3cb46cae8eba3c.svg","isPro":false,"fullname":"Alyssa","user":"alyssa1728593","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"5f6d64475e78cc6b0ed31e4c","name":"Salesforce","fullname":"Salesforce AI Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1602756670970-noauth.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.07545.md","query":{}}">
DarwinX: Evolving Agent Harnesses Through Natural Selection
Abstract
DarwinX evolves agent harnesses via population selection with frozen models, improving verified performance across benchmarks without benchmark-specific patches.
An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. We introduce DarwinX, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence share one edit interface. Fitness comes from each benchmark's own verifier: no gold solutions, no hand-picked winners. Across four benchmarks that progressively separate the evolution signal from the test, one loop adds about 17 points on average: Terminal-Bench 2.1 rises +7.7 to 83.2% on a matched base and to the verified frontier at 84.7% on a stronger one; TerminalWorld's held-out split reaches 68.3%, ahead of every off-the-shelf agent; WebArena-Infinity real-task pass@1 rises from 43.5% to 93.0% audit-clean; and a Terminal-Bench 2.1 harness transfers unchanged to SWE-bench Verified. What evolves is general agent competence, not benchmark-specific patches, so it survives changes of task, verifier, and base model. A frozen model need not be a fixed agent: harness selection turns evaluation compute into durable capability.
Community
We freeze the base model and evolve only the harness — prompts, tools, skills, control flow — by label-free natural selection: variants are scored on measured fitness (avg@k, no gold solutions), survivors are kept, and complementary ones are merged.
Matched-model, harness alone:
- Terminal-Bench 2.1 75.5 → 83.2 avg@5 on GPT-5.5, past Codex at 83.1 (84.7 on GPT-5.6 Sol)
- TerminalWorld 61.0 → 68.3 pass@1 on Opus 4.8, 41 held-out tasks, past Claude Code at 65.9
- WebArena-Infinity 43.5 → 93.0 audit-clean pass@1 on 1,260 real tasks
- SWE-bench Verified 80.8 → 84.2 by zero-shot transfer of the terminal harness
Two results we did not expect. The in-loop proxy saturates (0.505 → 1.000) while held-out pass@1 is 68.3 — a 31.7-point gap — and the variant that best fits the proxy is not the best generalizer, so keeping a population rather than following the single best lineage is what converts an overfit proxy into held-out gain. And we audited every WebArena-Infinity trajectory for validity: the base agent's raw 53.0 falls to 43.5 while ours goes 94.4 → 93.0, so the gap widens under scrutiny instead of narrowing.
Project page with interactive figures, every number regenerated from the run artifacts: https://huggingface.co/spaces/CoderDoge/darwinx
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.07545 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.07545 in a dataset README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.