webpage:<a href=\"https://envharness.com/\" rel=\"nofollow\">https://envharness.com/</a><br>code are all released.</p>\n","updatedAt":"2026-08-21T02:30:55.075Z","author":{"_id":"62ea79dd01ed9b0e8f61ccd3","avatarUrl":"/avatars/70af83e0e267be39fcd5f23b85e2dafa.svg","fullname":"Chengsong Huang","name":"ChengsongHuang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":15,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.98542320728302},"editors":["ChengsongHuang"],"editorAvatarUrls":["/avatars/70af83e0e267be39fcd5f23b85e2dafa.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.19880","authors":[{"_id":"6a87afc589e517cbfd75dbcf","name":"Chengsong Huang","hidden":false},{"_id":"6a87afc589e517cbfd75dbd0","name":"Zifeng Wang","hidden":false},{"_id":"6a87afc589e517cbfd75dbd1","name":"Rujun Han","hidden":false},{"_id":"6a87afc589e517cbfd75dbd2","name":"Jun Yan","hidden":false},{"_id":"6a87afc589e517cbfd75dbd3","name":"Yanfei Chen","hidden":false},{"_id":"6a87afc589e517cbfd75dbd4","name":"Zoey CuiZhu","hidden":false},{"_id":"6a87afc589e517cbfd75dbd5","name":"Ke Jiang","hidden":false},{"_id":"6a87afc589e517cbfd75dbd6","name":"Peng Xia","hidden":false},{"_id":"6a87afc589e517cbfd75dbd7","name":"Han Yu","hidden":false},{"_id":"6a87afc589e517cbfd75dbd8","name":"Yufan Zhuang","hidden":false},{"_id":"6a87afc589e517cbfd75dbd9","name":"Yifei Ming","hidden":false},{"_id":"6a87afc589e517cbfd75dbda","name":"Jiaqi Pan","hidden":false},{"_id":"6a87afc589e517cbfd75dbdb","name":"Bhavana Dalvi Mishra","hidden":false},{"_id":"6a87afc589e517cbfd75dbdc","name":"Jiaxin Huang","hidden":false},{"_id":"6a87afc589e517cbfd75dbdd","name":"Burak Gokturk","hidden":false},{"_id":"6a87afc589e517cbfd75dbde","name":"Tomas Pfister","hidden":false},{"_id":"6a87afc589e517cbfd75dbdf","name":"Chen-Yu Lee","hidden":false}],"publishedAt":"2026-08-20T00:00:00.000Z","submittedOnDailyAt":"2026-08-21T00:00:00.000Z","title":"EnvHarness: Awakening Static Worlds for Agent Learning","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and quickly left behind as it improves. While recent environment generation methods attempt to address this, they require domain-specific pipelines, rely on expensive or unreliable verifiers, and still produce static environments. To alleviate the engineering burden of rebuilding environments from scratch, we propose Environment Harness (EnvHarness), a programmable layer of plug-in components that wraps a static environment to reshape its behavior without modifying the underlying logic. Operating through standard interfaces, EnvHarness applies across diverse domains while ensuring every reshaped environment retains its original verifier. To automate this process, we introduce EnvRigger, which treats the target policy as a black box, observing its execution trajectories to synthesize EnvHarness components targeting diagnosed flaws, and validating them via fresh rollouts. Across five benchmarks in four domains, EnvHarness outperforms both original environments and domain-specific environment generation pipelines, achieving up to a 9.0-point improvement on held-out instances with 9.8% fewer execution steps. Furthermore, EnvHarness provides a superior optimization signal for reinforcement learning, enabling continuous, targeted co-evolution of the policy and its environment.","upvotes":158,"discussionId":"6a87afc589e517cbfd75dbe0","projectPage":"https://envharness.com/","githubRepo":"https://github.com/google-research/envharness","githubRepoAddedBy":"user","ai_summary":"EnvHarness and EnvRigger dynamically reshape static environments via programmable plugins to target agent weaknesses and improve reinforcement learning co-evolution.","ai_keywords":["LLM agents","Environment Harness","EnvHarness","EnvRigger","black-box policy","execution trajectories","reinforcement learning","co-evolution"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":8,"organization":{"_id":"5e6aca39878b8b2bf9806447","name":"google","fullname":"Google","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/WtA3YYitedOr9n02eHfJe.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"62ea79dd01ed9b0e8f61ccd3","avatarUrl":"/avatars/70af83e0e267be39fcd5f23b85e2dafa.svg","isPro":false,"fullname":"Chengsong Huang","user":"ChengsongHuang","type":"user"},{"_id":"64fc20d899123d7698a30e61","avatarUrl":"/avatars/9231982cf70a0689f50accedf1004702.svg","isPro":false,"fullname":"Jinyuan Li","user":"jinyuan222","type":"user"},{"_id":"6452faa03f80ad88c77c0efc","avatarUrl":"/avatars/2ce498d6a88f643dd91b6d56e14cb66e.svg","isPro":false,"fullname":"YUYI YANG","user":"yyuyi","type":"user"},{"_id":"68db3115f071f8164ecf3db5","avatarUrl":"/avatars/6e1de35de929bcea0e10441b7bd27375.svg","isPro":false,"fullname":"88170edison","user":"8817edison","type":"user"},{"_id":"6852f69c7b834665dc38c837","avatarUrl":"/avatars/bf9c5fc72300756e88319ec45c75f337.svg","isPro":true,"fullname":"Jianliang He","user":"JLiangHe","type":"user"},{"_id":"6850ab3fb73a72cdc6159f6f","avatarUrl":"/avatars/74c58336655cdfe7d3fb9854d426240d.svg","isPro":false,"fullname":"Haolin Liu","user":"lhl616","type":"user"},{"_id":"662b6a8f0b7f23f3c000559e","avatarUrl":"/avatars/0a5b4e09ac9a8e40342131319ff32b29.svg","isPro":false,"fullname":"Zisu Huang","user":"zisuh","type":"user"},{"_id":"69213a2d2e9ab0fed46648ae","avatarUrl":"/avatars/1815ef10cebb95b1f9cfe9f79e6a7434.svg","isPro":false,"fullname":"Peng Xia","user":"xpxpxpxp","type":"user"},{"_id":"6526307af06ac0cf9a922e86","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/nchCipX-XWw2cnzYsU_Cv.jpeg","isPro":false,"fullname":"Zhepei Wei","user":"weizhepei","type":"user"},{"_id":"650b077bac21d7087db628db","avatarUrl":"/avatars/0dbd611e8f5b3f01c09c5199ba88ca5a.svg","isPro":false,"fullname":"Mingtian Tan","user":"Tomo1916","type":"user"},{"_id":"64cb2d2bf9c57cdcb2616dae","avatarUrl":"/avatars/81c60d5b682d444e4988df400ba7f8ca.svg","isPro":false,"fullname":"Ying Xu","user":"FionaXu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"5e6aca39878b8b2bf9806447","name":"google","fullname":"Google","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/WtA3YYitedOr9n02eHfJe.png"},"query":{}}">
EnvHarness: Awakening Static Worlds for Agent Learning
Abstract
EnvHarness and EnvRigger dynamically reshape static environments via programmable plugins to target agent weaknesses and improve reinforcement learning co-evolution.
LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and quickly left behind as it improves. While recent environment generation methods attempt to address this, they require domain-specific pipelines, rely on expensive or unreliable verifiers, and still produce static environments. To alleviate the engineering burden of rebuilding environments from scratch, we propose Environment Harness (EnvHarness), a programmable layer of plug-in components that wraps a static environment to reshape its behavior without modifying the underlying logic. Operating through standard interfaces, EnvHarness applies across diverse domains while ensuring every reshaped environment retains its original verifier. To automate this process, we introduce EnvRigger, which treats the target policy as a black box, observing its execution trajectories to synthesize EnvHarness components targeting diagnosed flaws, and validating them via fresh rollouts. Across five benchmarks in four domains, EnvHarness outperforms both original environments and domain-specific environment generation pipelines, achieving up to a 9.0-point improvement on held-out instances with 9.8% fewer execution steps. Furthermore, EnvHarness provides a superior optimization signal for reinforcement learning, enabling continuous, targeted co-evolution of the policy and its environment.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.19880 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.19880 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.19880 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.