Hugging Face Daily Papers · · 5 min read

DarwinX: Evolving Agent Harnesses Through Natural Selection

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We freeze the base model and evolve only the harness — prompts, tools, skills, control flow — by label-free natural selection: variants are scored on measured fitness (avg@k, no gold solutions), survivors are kept, and complementary ones are merged.</p>\n<p>Matched-model, harness alone:</p>\n<ul>\n<li><strong>Terminal-Bench 2.1</strong> 75.5 → <strong>83.2</strong> avg@5 on GPT-5.5, past Codex at 83.1 (84.7 on GPT-5.6 Sol)</li>\n<li><strong>TerminalWorld</strong> 61.0 → <strong>68.3</strong> pass@1 on Opus 4.8, 41 held-out tasks, past Claude Code at 65.9</li>\n<li><strong>WebArena-Infinity</strong> 43.5 → <strong>93.0</strong> audit-clean pass@1 on 1,260 real tasks</li>\n<li><strong>SWE-bench Verified</strong> 80.8 → <strong>84.2</strong> by zero-shot transfer of the terminal harness</li>\n</ul>\n<p>Two results we did not expect. The in-loop proxy saturates (0.505 → 1.000) while held-out pass@1 is 68.3 — a 31.7-point gap — and the variant that best fits the proxy is <em>not</em> the best generalizer, so keeping a population rather than following the single best lineage is what converts an overfit proxy into held-out gain. And we audited every WebArena-Infinity trajectory for validity: the base agent's raw 53.0 falls to 43.5 while ours goes 94.4 → 93.0, so the gap <em>widens</em> under scrutiny instead of narrowing.</p>\n<p>Project page with interactive figures, every number regenerated from the run artifacts: <a href=\"https://huggingface.co/spaces/CoderDoge/darwinx\">https://huggingface.co/spaces/CoderDoge/darwinx</a></p>\n","updatedAt":"2026-08-14T02:53:29.038Z","author":{"_id":"638a765ef560ea995b580cc4","avatarUrl":"/avatars/16dd0363d6c99e35f8a6522a29f4c9b2.svg","fullname":"Yifan Zhang","name":"CoderDoge","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.822934627532959},"editors":["CoderDoge"],"editorAvatarUrls":["/avatars/16dd0363d6c99e35f8a6522a29f4c9b2.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.07545","authors":[{"_id":"6a7da6fd42823931a1f173df","name":"Yifan Zhang","hidden":false},{"_id":"6a7da6fd42823931a1f173e0","name":"Yutong Dai","hidden":false},{"_id":"6a7da6fd42823931a1f173e1","name":"Juntao Tan","hidden":false},{"_id":"6a7da6fd42823931a1f173e2","name":"Luyu Yang","hidden":false},{"_id":"6a7da6fd42823931a1f173e3","name":"Rishi Mullur","hidden":false},{"_id":"6a7da6fd42823931a1f173e4","name":"Thai Hoang","hidden":false},{"_id":"6a7da6fd42823931a1f173e5","name":"Zhiyuan Hu","hidden":false},{"_id":"6a7da6fd42823931a1f173e6","name":"James Zhu","hidden":false},{"_id":"6a7da6fd42823931a1f173e7","name":"Phil Mui","hidden":false},{"_id":"6a7da6fd42823931a1f173e8","name":"Silvio Savarese","hidden":false},{"_id":"6a7da6fd42823931a1f173e9","name":"Ran Xu","hidden":false},{"_id":"6a7da6fd42823931a1f173ea","name":"Zeyuan Chen","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/638a765ef560ea995b580cc4/MUjDLUz6Uyes2wkpKGuz9.png"],"publishedAt":"2026-07-31T00:00:00.000Z","submittedOnDailyAt":"2026-08-14T00:00:00.000Z","title":"DarwinX: Evolving Agent Harnesses Through Natural Selection","submittedOnDailyBy":{"_id":"638a765ef560ea995b580cc4","avatarUrl":"/avatars/16dd0363d6c99e35f8a6522a29f4c9b2.svg","isPro":false,"fullname":"Yifan Zhang","user":"CoderDoge","type":"user","name":"CoderDoge"},"summary":"An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. We introduce DarwinX, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence share one edit interface. Fitness comes from each benchmark's own verifier: no gold solutions, no hand-picked winners. Across four benchmarks that progressively separate the evolution signal from the test, one loop adds about 17 points on average: Terminal-Bench 2.1 rises +7.7 to 83.2% on a matched base and to the verified frontier at 84.7% on a stronger one; TerminalWorld's held-out split reaches 68.3%, ahead of every off-the-shelf agent; WebArena-Infinity real-task pass@1 rises from 43.5% to 93.0% audit-clean; and a Terminal-Bench 2.1 harness transfers unchanged to SWE-bench Verified. What evolves is general agent competence, not benchmark-specific patches, so it survives changes of task, verifier, and base model. A frozen model need not be a fixed agent: harness selection turns evaluation compute into durable capability.","upvotes":39,"discussionId":"6a7da6fe42823931a1f173eb","projectPage":"https://huggingface.co/spaces/CoderDoge/darwinx","ai_summary":"DarwinX evolves agent harnesses via population selection with frozen models, improving verified performance across benchmarks without benchmark-specific patches.","ai_keywords":["LLM agent","harness","self-improvement","population selection","preserve-and-extend contract","archive","recombination","verifier","DarwinX"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"5f6d64475e78cc6b0ed31e4c","name":"Salesforce","fullname":"Salesforce AI Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1602756670970-noauth.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"638a765ef560ea995b580cc4","avatarUrl":"/avatars/16dd0363d6c99e35f8a6522a29f4c9b2.svg","isPro":false,"fullname":"Yifan Zhang","user":"CoderDoge","type":"user"},{"_id":"632405844624707127217677","avatarUrl":"/avatars/af5deba9ec51814bc5ccfd04f68d76d3.svg","isPro":false,"fullname":"Zeyuan Chen","user":"bradypus","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"64c49396bf1954890192bfbf","avatarUrl":"/avatars/b2dbd1d601911440788efcea2a2b77d3.svg","isPro":false,"fullname":"Yutong Dai","user":"UncleFish","type":"user"},{"_id":"65cb46f919683f9817dd6909","avatarUrl":"/avatars/610e7de911497a3b43add837aaedf6f7.svg","isPro":false,"fullname":"rui zhang","user":"frankzhangrui","type":"user"},{"_id":"6465c4c863e7e09dd02e3e1b","avatarUrl":"/avatars/200b029184d2616f98296a2c212f0785.svg","isPro":false,"fullname":"Ran Xu","user":"xurantju","type":"user"},{"_id":"651b35c254134dcaad03758d","avatarUrl":"/avatars/0230e088c0c73184bdcae47240be86c5.svg","isPro":false,"fullname":"Thai Hoang","user":"qthai912uw","type":"user"},{"_id":"6a0763957577a33e8c39da90","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/bRyJvUqc4-ZpKtA_A_bIa.png","isPro":false,"fullname":"Frank Wang","user":"frankwang611","type":"user"},{"_id":"64b682e99a338c25127af067","avatarUrl":"/avatars/2c72eff485c2f265839eb6108e79ef0a.svg","isPro":false,"fullname":"Chen Huang","user":"ChenZiChun","type":"user"},{"_id":"627a124ffe55fa0f8ce0eaf7","avatarUrl":"/avatars/41e0dc029faed6dc45d620c5fe2652a5.svg","isPro":false,"fullname":"Serendipity","user":"Yuhan","type":"user"},{"_id":"65f8b846d761741567ecfe92","avatarUrl":"/avatars/eb09761dfe7ccc94ca09f94cc4e6e6b8.svg","isPro":false,"fullname":"Juntao Tan","user":"chrisjtan23","type":"user"},{"_id":"6a7e971e560102d6b0f425a6","avatarUrl":"/avatars/c61240d7f58683312b3cb46cae8eba3c.svg","isPro":false,"fullname":"Alyssa","user":"alyssa1728593","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"5f6d64475e78cc6b0ed31e4c","name":"Salesforce","fullname":"Salesforce AI Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1602756670970-noauth.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.07545.md","query":{}}">
Papers
arxiv:2608.07545

DarwinX: Evolving Agent Harnesses Through Natural Selection

Published on Jul 31
· Submitted by
Yifan Zhang
on Aug 14
#2 Paper of the day
Authors:
,

Abstract

DarwinX evolves agent harnesses via population selection with frozen models, improving verified performance across benchmarks without benchmark-specific patches.

An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. We introduce DarwinX, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence share one edit interface. Fitness comes from each benchmark's own verifier: no gold solutions, no hand-picked winners. Across four benchmarks that progressively separate the evolution signal from the test, one loop adds about 17 points on average: Terminal-Bench 2.1 rises +7.7 to 83.2% on a matched base and to the verified frontier at 84.7% on a stronger one; TerminalWorld's held-out split reaches 68.3%, ahead of every off-the-shelf agent; WebArena-Infinity real-task pass@1 rises from 43.5% to 93.0% audit-clean; and a Terminal-Bench 2.1 harness transfers unchanged to SWE-bench Verified. What evolves is general agent competence, not benchmark-specific patches, so it survives changes of task, verifier, and base model. A frozen model need not be a fixed agent: harness selection turns evaluation compute into durable capability.

Community

Paper submitter about 9 hours ago

We freeze the base model and evolve only the harness — prompts, tools, skills, control flow — by label-free natural selection: variants are scored on measured fitness (avg@k, no gold solutions), survivors are kept, and complementary ones are merged.

Matched-model, harness alone:

  • Terminal-Bench 2.1 75.5 → 83.2 avg@5 on GPT-5.5, past Codex at 83.1 (84.7 on GPT-5.6 Sol)
  • TerminalWorld 61.0 → 68.3 pass@1 on Opus 4.8, 41 held-out tasks, past Claude Code at 65.9
  • WebArena-Infinity 43.5 → 93.0 audit-clean pass@1 on 1,260 real tasks
  • SWE-bench Verified 80.8 → 84.2 by zero-shot transfer of the terminal harness

Two results we did not expect. The in-loop proxy saturates (0.505 → 1.000) while held-out pass@1 is 68.3 — a 31.7-point gap — and the variant that best fits the proxy is not the best generalizer, so keeping a population rather than following the single best lineage is what converts an overfit proxy into held-out gain. And we audited every WebArena-Infinity trajectory for validity: the base agent's raw 53.0 falls to 43.5 while ours goes 94.4 → 93.0, so the gap widens under scrutiny instead of narrowing.

Project page with interactive figures, every number regenerated from the run artifacts: https://huggingface.co/spaces/CoderDoge/darwinx

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.07545
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.07545 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.07545 in a dataset README.md to link it from this page.

Spaces citing this paper

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers