Hugging Face Daily Papers · · 5 min read

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Will Agentic ESOpt Be the Future of Long-Horizon Agentic Fine-Tuning?</p>\n<p>Excited to share our new work on Arxiv.<br>Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Memory Requirements</p>\n<p>As LLM agents become longer-horizon and more branching, conventional Agentic RL faces three growing bottlenecks: high training/memory cost and increasingly challenging long-horizon credit assignment.</p>\n<p>Our Agentic ESOpt approaches agent fine-tuning from the parameter space instead, targeting three advantages:</p>\n<p>⚡ Model Scalability — only minimal inference-level GPU memory requirement<br>🔧 Flexibility — prompt/skill–parameter co-evolution<br>⏳ Long-Horizon Scalability — trajectory-level parameter attribution without horizon-wise decomposition</p>\n","updatedAt":"2026-08-19T01:59:00.197Z","author":{"_id":"67a1d21e33e92b4a1183f3bb","avatarUrl":"/avatars/43f9dd3fcb7d58ddc69562fd1fc12957.svg","fullname":"Zhi Zheng","name":"zz1358m","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6855172514915466},"editors":["zz1358m"],"editorAvatarUrls":["/avatars/43f9dd3fcb7d58ddc69562fd1fc12957.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.17310","authors":[{"_id":"6a850de3536bdd3bdd48f717","name":"Zhi Zheng","hidden":false},{"_id":"6a850de3536bdd3bdd48f718","name":"Rongsheng Chen","hidden":false},{"_id":"6a850de3536bdd3bdd48f719","name":"Yunpeng Ba","hidden":false},{"_id":"6a850de3536bdd3bdd48f71a","name":"Zhenkun Wang","hidden":false},{"_id":"6a850de3536bdd3bdd48f71b","name":"Yee Whye Teh","hidden":false},{"_id":"6a850de3536bdd3bdd48f71c","name":"Wee Sun Lee","hidden":false}],"publishedAt":"2026-08-18T00:00:00.000Z","submittedOnDailyAt":"2026-08-19T00:00:00.000Z","title":"Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements","submittedOnDailyBy":{"_id":"67a1d21e33e92b4a1183f3bb","avatarUrl":"/avatars/43f9dd3fcb7d58ddc69562fd1fc12957.svg","isPro":false,"fullname":"Zhi Zheng","user":"zz1358m","type":"user","name":"zz1358m"},"summary":"Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack makes it impractical to fine-tune larger LLMs, and longer-horizon trajectories make credit assignment in RL substantially harder. This paper argues that evolution strategies (ES) can be a better choice for fine-tuning long-horizon LLM agents. Compared with agentic RL, ES offers three key advantages: 1) Model Scalability: ES enables full-parameter optimization with only minimal, inference-level GPU memory, making it possible to fine-tune large LLMs. 2) Flexibility: its lightweight, black-box feedback interface makes ES fine-tuning easy to compose with prompt-space evolution (e.g., skill optimization & test-time compute); and 3) Long-Horizon Scalability: ES performs trajectory-level parameter attribution without decomposing rewards across horizons, yielding better scalability than Agentic RL as the horizon length grows. Based on this insight, we propose Agentic ESOpt, a full-parameter agentic fine-tuning framework tailored to flexible parameter--context co-evolution. At each step, Agentic ESOpt samples perturbations around the current LLM parameters, evaluates the resulting agents with rewards, and applies an online reward-weighted update. To improve the exploration--adaptation trade-off, Agentic ESOpt further introduces a cosine decay schedule of the perturbation scale σ. On WebArena-Lite, full-parameter optimization of Qwen-3.5-27B improves the No Skill baseline by 6.69%. In test-time automatic heuristic design, Agentic ESOpt performs online prompt--parameter co-evolution, improving its matched baseline in 28 of 36 settings.","upvotes":28,"discussionId":"6a850de3536bdd3bdd48f71d","projectPage":"https://zz1358m.github.io/Project-Agentic-ESOpt","githubRepo":"https://github.com/zz1358m/Agentic-ESOpt","githubRepoAddedBy":"user","ai_summary":"Agentic ESOpt uses evolution strategies for scalable full-parameter fine-tuning of long-horizon LLM agents via trajectory-level reward-weighted updates and parameter-context co-evolution.","ai_keywords":["evolution strategies","reinforcement learning","long-horizon agentic reasoning","credit assignment","full-parameter optimization","Agentic ESOpt","parameter-context co-evolution","reward-weighted update","cosine decay schedule","WebArena-Lite"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":3,"organization":{"_id":"6508ab2b349930913196378b","name":"NationalUniversityofSingapore","fullname":"National University of Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/630ca0817dacb93b33506ce7/ZYUmpSMsa5Whihw3me2Bw.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67a1d21e33e92b4a1183f3bb","avatarUrl":"/avatars/43f9dd3fcb7d58ddc69562fd1fc12957.svg","isPro":false,"fullname":"Zhi Zheng","user":"zz1358m","type":"user"},{"_id":"6a1481b614524dbef227e9cf","avatarUrl":"/avatars/6950a2495e270f72af553803c7abbdb8.svg","isPro":false,"fullname":"Mario Ba","user":"marioba7","type":"user"},{"_id":"6a1ef6772a65642f97ab44e1","avatarUrl":"/avatars/d17285f1efdb5e5973dd6f6bec34c1dd.svg","isPro":false,"fullname":"Jiaqing Li","user":"ljq34952","type":"user"},{"_id":"67b2d9cb48d466697ba54563","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/N7zTMQvNNOJL4TMNmRNxt.png","isPro":false,"fullname":"Guyu","user":"kuangrepi","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"675d25fe61ac4f52fa074211","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/q1Q0IMYUtmAumUP7oSDWq.png","isPro":false,"fullname":"Shuyuan","user":"ShuyuanNan","type":"user"},{"_id":"67cd4815af5349a6171f072f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/BEo-6dNNMbFwhKStCjkAz.png","isPro":false,"fullname":"Li Jiaqing","user":"Mambaout0824","type":"user"},{"_id":"63885f1d0bebb233d8ad6e5b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1669881620925-noauth.jpeg","isPro":false,"fullname":"Penghui Qi","user":"QPHutu","type":"user"},{"_id":"677dd7e23b7a7dcf9579324f","avatarUrl":"/avatars/306c1796d072abc5f467dc58489eae5b.svg","isPro":false,"fullname":"Pengfei Yang","user":"Yang2005","type":"user"},{"_id":"66ec09285d5bc73473746fb6","avatarUrl":"/avatars/8363f086734f05161a1fbab702a4f284.svg","isPro":false,"fullname":"Liu Yifan","user":"Eclipse76","type":"user"},{"_id":"6926965987352e7a6d4bef8a","avatarUrl":"/avatars/7e5ef18d2f02e5d99c04d346d9ad35d1.svg","isPro":false,"fullname":"Zhiyu Hou","user":"KevinHuge","type":"user"},{"_id":"6850d7dc2a1b88bdf7e76227","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ugGrRqt3QxdQhVc-oy76d.png","isPro":false,"fullname":"Wang Zimo","user":"ChevalierDeSangreal","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"6508ab2b349930913196378b","name":"NationalUniversityofSingapore","fullname":"National University of Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/630ca0817dacb93b33506ce7/ZYUmpSMsa5Whihw3me2Bw.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.17310.md","query":{}}">
Papers
arxiv:2608.17310

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Published on Aug 18
· Submitted by
Zhi Zheng
on Aug 19
#3 Paper of the day
Authors:
,

Abstract

Agentic ESOpt uses evolution strategies for scalable full-parameter fine-tuning of long-horizon LLM agents via trajectory-level reward-weighted updates and parameter-context co-evolution.

Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack makes it impractical to fine-tune larger LLMs, and longer-horizon trajectories make credit assignment in RL substantially harder. This paper argues that evolution strategies (ES) can be a better choice for fine-tuning long-horizon LLM agents. Compared with agentic RL, ES offers three key advantages: 1) Model Scalability: ES enables full-parameter optimization with only minimal, inference-level GPU memory, making it possible to fine-tune large LLMs. 2) Flexibility: its lightweight, black-box feedback interface makes ES fine-tuning easy to compose with prompt-space evolution (e.g., skill optimization & test-time compute); and 3) Long-Horizon Scalability: ES performs trajectory-level parameter attribution without decomposing rewards across horizons, yielding better scalability than Agentic RL as the horizon length grows. Based on this insight, we propose Agentic ESOpt, a full-parameter agentic fine-tuning framework tailored to flexible parameter--context co-evolution. At each step, Agentic ESOpt samples perturbations around the current LLM parameters, evaluates the resulting agents with rewards, and applies an online reward-weighted update. To improve the exploration--adaptation trade-off, Agentic ESOpt further introduces a cosine decay schedule of the perturbation scale σ. On WebArena-Lite, full-parameter optimization of Qwen-3.5-27B improves the No Skill baseline by 6.69%. In test-time automatic heuristic design, Agentic ESOpt performs online prompt--parameter co-evolution, improving its matched baseline in 28 of 36 settings.

Community

Paper submitter about 6 hours ago

Will Agentic ESOpt Be the Future of Long-Horizon Agentic Fine-Tuning?

Excited to share our new work on Arxiv.
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Memory Requirements

As LLM agents become longer-horizon and more branching, conventional Agentic RL faces three growing bottlenecks: high training/memory cost and increasingly challenging long-horizon credit assignment.

Our Agentic ESOpt approaches agent fine-tuning from the parameter space instead, targeting three advantages:

⚡ Model Scalability — only minimal inference-level GPU memory requirement
🔧 Flexibility — prompt/skill–parameter co-evolution
⏳ Long-Horizon Scalability — trajectory-level parameter attribution without horizon-wise decomposition

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.17310
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.17310 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.17310 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.17310 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers