Will Agentic ESOpt Be the Future of Long-Horizon Agentic Fine-Tuning?</p>\n<p>Excited to share our new work on Arxiv.<br>Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Memory Requirements</p>\n<p>As LLM agents become longer-horizon and more branching, conventional Agentic RL faces three growing bottlenecks: high training/memory cost and increasingly challenging long-horizon credit assignment.</p>\n<p>Our Agentic ESOpt approaches agent fine-tuning from the parameter space instead, targeting three advantages:</p>\n<p>⚡ Model Scalability — only minimal inference-level GPU memory requirement<br>🔧 Flexibility — prompt/skill–parameter co-evolution<br>⏳ Long-Horizon Scalability — trajectory-level parameter attribution without horizon-wise decomposition</p>\n","updatedAt":"2026-08-19T01:59:00.197Z","author":{"_id":"67a1d21e33e92b4a1183f3bb","avatarUrl":"/avatars/43f9dd3fcb7d58ddc69562fd1fc12957.svg","fullname":"Zhi Zheng","name":"zz1358m","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6855172514915466},"editors":["zz1358m"],"editorAvatarUrls":["/avatars/43f9dd3fcb7d58ddc69562fd1fc12957.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.17310","authors":[{"_id":"6a850de3536bdd3bdd48f717","name":"Zhi Zheng","hidden":false},{"_id":"6a850de3536bdd3bdd48f718","name":"Rongsheng Chen","hidden":false},{"_id":"6a850de3536bdd3bdd48f719","name":"Yunpeng Ba","hidden":false},{"_id":"6a850de3536bdd3bdd48f71a","name":"Zhenkun Wang","hidden":false},{"_id":"6a850de3536bdd3bdd48f71b","name":"Yee Whye Teh","hidden":false},{"_id":"6a850de3536bdd3bdd48f71c","name":"Wee Sun Lee","hidden":false}],"publishedAt":"2026-08-18T00:00:00.000Z","submittedOnDailyAt":"2026-08-19T00:00:00.000Z","title":"Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements","submittedOnDailyBy":{"_id":"67a1d21e33e92b4a1183f3bb","avatarUrl":"/avatars/43f9dd3fcb7d58ddc69562fd1fc12957.svg","isPro":false,"fullname":"Zhi Zheng","user":"zz1358m","type":"user","name":"zz1358m"},"summary":"Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack makes it impractical to fine-tune larger LLMs, and longer-horizon trajectories make credit assignment in RL substantially harder. This paper argues that evolution strategies (ES) can be a better choice for fine-tuning long-horizon LLM agents. Compared with agentic RL, ES offers three key advantages: 1) Model Scalability: ES enables full-parameter optimization with only minimal, inference-level GPU memory, making it possible to fine-tune large LLMs. 2) Flexibility: its lightweight, black-box feedback interface makes ES fine-tuning easy to compose with prompt-space evolution (e.g., skill optimization & test-time compute); and 3) Long-Horizon Scalability: ES performs trajectory-level parameter attribution without decomposing rewards across horizons, yielding better scalability than Agentic RL as the horizon length grows. Based on this insight, we propose Agentic ESOpt, a full-parameter agentic fine-tuning framework tailored to flexible parameter--context co-evolution. At each step, Agentic ESOpt samples perturbations around the current LLM parameters, evaluates the resulting agents with rewards, and applies an online reward-weighted update. To improve the exploration--adaptation trade-off, Agentic ESOpt further introduces a cosine decay schedule of the perturbation scale σ. On WebArena-Lite, full-parameter optimization of Qwen-3.5-27B improves the No Skill baseline by 6.69%. In test-time automatic heuristic design, Agentic ESOpt performs online prompt--parameter co-evolution, improving its matched baseline in 28 of 36 settings.","upvotes":28,"discussionId":"6a850de3536bdd3bdd48f71d","projectPage":"https://zz1358m.github.io/Project-Agentic-ESOpt","githubRepo":"https://github.com/zz1358m/Agentic-ESOpt","githubRepoAddedBy":"user","ai_summary":"Agentic ESOpt uses evolution strategies for scalable full-parameter fine-tuning of long-horizon LLM agents via trajectory-level reward-weighted updates and parameter-context co-evolution.","ai_keywords":["evolution strategies","reinforcement learning","long-horizon agentic reasoning","credit assignment","full-parameter optimization","Agentic ESOpt","parameter-context co-evolution","reward-weighted update","cosine decay schedule","WebArena-Lite"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":3,"organization":{"_id":"6508ab2b349930913196378b","name":"NationalUniversityofSingapore","fullname":"National University of Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/630ca0817dacb93b33506ce7/ZYUmpSMsa5Whihw3me2Bw.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67a1d21e33e92b4a1183f3bb","avatarUrl":"/avatars/43f9dd3fcb7d58ddc69562fd1fc12957.svg","isPro":false,"fullname":"Zhi Zheng","user":"zz1358m","type":"user"},{"_id":"6a1481b614524dbef227e9cf","avatarUrl":"/avatars/6950a2495e270f72af553803c7abbdb8.svg","isPro":false,"fullname":"Mario Ba","user":"marioba7","type":"user"},{"_id":"6a1ef6772a65642f97ab44e1","avatarUrl":"/avatars/d17285f1efdb5e5973dd6f6bec34c1dd.svg","isPro":false,"fullname":"Jiaqing Li","user":"ljq34952","type":"user"},{"_id":"67b2d9cb48d466697ba54563","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/N7zTMQvNNOJL4TMNmRNxt.png","isPro":false,"fullname":"Guyu","user":"kuangrepi","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"675d25fe61ac4f52fa074211","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/q1Q0IMYUtmAumUP7oSDWq.png","isPro":false,"fullname":"Shuyuan","user":"ShuyuanNan","type":"user"},{"_id":"67cd4815af5349a6171f072f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/BEo-6dNNMbFwhKStCjkAz.png","isPro":false,"fullname":"Li Jiaqing","user":"Mambaout0824","type":"user"},{"_id":"63885f1d0bebb233d8ad6e5b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1669881620925-noauth.jpeg","isPro":false,"fullname":"Penghui Qi","user":"QPHutu","type":"user"},{"_id":"677dd7e23b7a7dcf9579324f","avatarUrl":"/avatars/306c1796d072abc5f467dc58489eae5b.svg","isPro":false,"fullname":"Pengfei Yang","user":"Yang2005","type":"user"},{"_id":"66ec09285d5bc73473746fb6","avatarUrl":"/avatars/8363f086734f05161a1fbab702a4f284.svg","isPro":false,"fullname":"Liu Yifan","user":"Eclipse76","type":"user"},{"_id":"6926965987352e7a6d4bef8a","avatarUrl":"/avatars/7e5ef18d2f02e5d99c04d346d9ad35d1.svg","isPro":false,"fullname":"Zhiyu Hou","user":"KevinHuge","type":"user"},{"_id":"6850d7dc2a1b88bdf7e76227","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ugGrRqt3QxdQhVc-oy76d.png","isPro":false,"fullname":"Wang Zimo","user":"ChevalierDeSangreal","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"6508ab2b349930913196378b","name":"NationalUniversityofSingapore","fullname":"National University of Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/630ca0817dacb93b33506ce7/ZYUmpSMsa5Whihw3me2Bw.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.17310.md","query":{}}">
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
Abstract
Agentic ESOpt uses evolution strategies for scalable full-parameter fine-tuning of long-horizon LLM agents via trajectory-level reward-weighted updates and parameter-context co-evolution.
Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack makes it impractical to fine-tune larger LLMs, and longer-horizon trajectories make credit assignment in RL substantially harder. This paper argues that evolution strategies (ES) can be a better choice for fine-tuning long-horizon LLM agents. Compared with agentic RL, ES offers three key advantages: 1) Model Scalability: ES enables full-parameter optimization with only minimal, inference-level GPU memory, making it possible to fine-tune large LLMs. 2) Flexibility: its lightweight, black-box feedback interface makes ES fine-tuning easy to compose with prompt-space evolution (e.g., skill optimization & test-time compute); and 3) Long-Horizon Scalability: ES performs trajectory-level parameter attribution without decomposing rewards across horizons, yielding better scalability than Agentic RL as the horizon length grows. Based on this insight, we propose Agentic ESOpt, a full-parameter agentic fine-tuning framework tailored to flexible parameter--context co-evolution. At each step, Agentic ESOpt samples perturbations around the current LLM parameters, evaluates the resulting agents with rewards, and applies an online reward-weighted update. To improve the exploration--adaptation trade-off, Agentic ESOpt further introduces a cosine decay schedule of the perturbation scale σ. On WebArena-Lite, full-parameter optimization of Qwen-3.5-27B improves the No Skill baseline by 6.69%. In test-time automatic heuristic design, Agentic ESOpt performs online prompt--parameter co-evolution, improving its matched baseline in 28 of 36 settings.
Community
Will Agentic ESOpt Be the Future of Long-Horizon Agentic Fine-Tuning?
Excited to share our new work on Arxiv.
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Memory Requirements
As LLM agents become longer-horizon and more branching, conventional Agentic RL faces three growing bottlenecks: high training/memory cost and increasingly challenging long-horizon credit assignment.
Our Agentic ESOpt approaches agent fine-tuning from the parameter space instead, targeting three advantages:
⚡ Model Scalability — only minimal inference-level GPU memory requirement
🔧 Flexibility — prompt/skill–parameter co-evolution
⏳ Long-Horizon Scalability — trajectory-level parameter attribution without horizon-wise decomposition
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.17310 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.17310 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.17310 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.