Hugging Face Daily Papers · · 4 min read

Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We propose Agent-G^2, a Gaussian guidance framework that draws the depth per task from a Gaussian whose center and spread are estimated online from rollouts already collected for policy optimization, requiring no probe rollouts or learned depth predictor.</p>\n","updatedAt":"2026-08-27T02:18:31.741Z","author":{"_id":"676127cf11b19ea602bb202a","avatarUrl":"/avatars/dfd802a24bd63e509728159ebb1769f6.svg","fullname":"Zhengxi Lu","name":"LZXzju","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":13,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.830083429813385},"editors":["LZXzju"],"editorAvatarUrls":["/avatars/dfd802a24bd63e509728159ebb1769f6.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.23318","authors":[{"_id":"6a8d0bb25add2537c32e9706","user":{"_id":"677d3523a6918748bf81d0e9","avatarUrl":"/avatars/c9961ac54d089efb36db20c421c2bea2.svg","isPro":false,"fullname":"wangzixuan","user":"wangzx1210","type":"user","name":"wangzx1210"},"name":"Zixuan Wang","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:13:50.781Z","hidden":false},{"_id":"6a8d0bb25add2537c32e9707","name":"Yanrui Miao","hidden":false},{"_id":"6a8d0bb25add2537c32e9708","user":{"_id":"676127cf11b19ea602bb202a","avatarUrl":"/avatars/dfd802a24bd63e509728159ebb1769f6.svg","isPro":false,"fullname":"Zhengxi Lu","user":"LZXzju","type":"user","name":"LZXzju"},"name":"Zhengxi Lu","status":"claimed_verified","statusLastChangedAt":"2026-08-27T08:45:04.860Z","hidden":false},{"_id":"6a8d0bb25add2537c32e9709","name":"Teng Pan","hidden":false},{"_id":"6a8d0bb25add2537c32e970a","name":"Yiwen Qiu","hidden":false},{"_id":"6a8d0bb25add2537c32e970b","name":"Hongxing Li","hidden":false},{"_id":"6a8d0bb25add2537c32e970c","name":"Peng Qiu","hidden":false},{"_id":"6a8d0bb25add2537c32e970d","name":"Ruiqing Zhang","hidden":false},{"_id":"6a8d0bb25add2537c32e970e","name":"Yongliang Shen","hidden":false}],"publishedAt":"2026-08-24T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning","submittedOnDailyBy":{"_id":"676127cf11b19ea602bb202a","avatarUrl":"/avatars/dfd802a24bd63e509728159ebb1769f6.svg","isPro":false,"fullname":"Zhengxi Lu","user":"LZXzju","type":"user","name":"LZXzju"},"summary":"Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory before each rollout, letting the policy explore from a state closer to success. Its effectiveness hinges on the guidance depth: how much of the trajectory to keep. Existing methods treat this depth as a deterministic scalar. Scheduled approaches share one value across samples and ignore per-task heterogeneity; per-sample probing estimates it separately at the cost of extra rollouts. We find that useful guidance occupies a band of depths whose informativeness profile is approximately Gaussian around the band center, rather than concentrating at a single optimal point. We propose Agent-G^2, a Gaussian guidance framework that draws the depth per task from a Gaussian whose center and spread are estimated online from rollouts already collected for policy optimization, requiring no probe rollouts or learned depth predictor. The center combines a global baseline with per-cluster difficulty, and the spread tracks within-cluster variance. We evaluate Agent-G^2 on ALFWorld and WebShop on Qwen2.5-1.5B / 7B-Instruct. Agent-G^2 outperforms the strongest hint-based, hint-free, and Aux-RL baselines on ALFWorld by 2.3 / 3.9 / 7.4 points at under one-third the rollout cost of per-sample probing.","upvotes":17,"discussionId":"6a8d0bb25add2537c32e970f","projectPage":"https://zju-real.github.io/Agent-G2/","githubRepo":"https://github.com/ZJU-REAL/Agent-G2","githubRepoAddedBy":"user","ai_summary":"Agent-G² models hint depth as a Gaussian distribution estimated online from existing rollouts, improving reinforcement learning on long-horizon tasks without extra probing.","ai_keywords":["hint-based reinforcement learning","reward sparsity","expert trajectory","guidance depth","Gaussian guidance","Agent-G²","ALFWorld","WebShop"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":18,"organization":{"_id":"6345aadf5efccdc07f1365a5","name":"ZhejiangUniversity","fullname":"Zhejiang University","avatar":"https://www.gravatar.com/avatar/d1d414628877bec2958f95ad283c15e7?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"677d3523a6918748bf81d0e9","avatarUrl":"/avatars/c9961ac54d089efb36db20c421c2bea2.svg","isPro":false,"fullname":"wangzixuan","user":"wangzx1210","type":"user"},{"_id":"66bc0ebd9c13cd4047a9bee0","avatarUrl":"/avatars/3bcf7d614ba6d6e975734eecc94fe037.svg","isPro":false,"fullname":"luka","user":"077lukamagic","type":"user"},{"_id":"698d83748068f1c868b8c0aa","avatarUrl":"/avatars/fb861fd898613f570478885512d046c1.svg","isPro":false,"fullname":"Yue","user":"WwwweweYue","type":"user"},{"_id":"6a0fe30ffb597becf7f6ea4f","avatarUrl":"/avatars/90c6a41e88f556a022566391ee1593fa.svg","isPro":false,"fullname":"Yue","user":"zhengwuyue","type":"user"},{"_id":"6a8f87a3056f4cd846f25ab2","avatarUrl":"/avatars/0afa5d4b4e23ac3e43cd7228aa17f07d.svg","isPro":false,"fullname":"Jinzhi Zhao","user":"Bonitoz","type":"user"},{"_id":"682b396401efbfd69d69a4ba","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/8VxCrRbuyuqmRVtXl386n.png","isPro":false,"fullname":"Jiang","user":"RyanJiang0817","type":"user"},{"_id":"67543820c3af453d7b3e1d5e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67543820c3af453d7b3e1d5e/RbAZ9AQlxpy5E5Is-QN8b.jpeg","isPro":false,"fullname":"Dingming Li","user":"lidingm","type":"user"},{"_id":"676127cf11b19ea602bb202a","avatarUrl":"/avatars/dfd802a24bd63e509728159ebb1769f6.svg","isPro":false,"fullname":"Zhengxi Lu","user":"LZXzju","type":"user"},{"_id":"5e1058e9fcf41d740b69966d","avatarUrl":"/avatars/ce74839ba871f2b54313a670a233ba82.svg","isPro":false,"fullname":"Yongliang Shen","user":"tricktreat","type":"user"},{"_id":"67b970e414b1af2ce915c906","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/2l7CwmVEs8NVZWcYgHk36.png","isPro":false,"fullname":"yiwen qiu","user":"qywMichelle","type":"user"},{"_id":"682c0fcebbe0c6fa323f531b","avatarUrl":"/avatars/953e6f5ce0361c2f9693bc0ca82787b7.svg","isPro":false,"fullname":"yy","user":"yuy07","type":"user"},{"_id":"68c7a335a7a25ec4cddf62be","avatarUrl":"/avatars/6cda5245a8a046aa455f9552888ed56d.svg","isPro":false,"fullname":"Kaiwen Zhang","user":"KeepKevin","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6345aadf5efccdc07f1365a5","name":"ZhejiangUniversity","fullname":"Zhejiang University","avatar":"https://www.gravatar.com/avatar/d1d414628877bec2958f95ad283c15e7?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.23318.md","query":{}}">
Papers
arxiv:2608.23318

Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning

Published on Aug 24
· Submitted by
Zhengxi Lu
on Aug 27
Authors:

Abstract

Agent-G² models hint depth as a Gaussian distribution estimated online from existing rollouts, improving reinforcement learning on long-horizon tasks without extra probing.

Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory before each rollout, letting the policy explore from a state closer to success. Its effectiveness hinges on the guidance depth: how much of the trajectory to keep. Existing methods treat this depth as a deterministic scalar. Scheduled approaches share one value across samples and ignore per-task heterogeneity; per-sample probing estimates it separately at the cost of extra rollouts. We find that useful guidance occupies a band of depths whose informativeness profile is approximately Gaussian around the band center, rather than concentrating at a single optimal point. We propose Agent-G^2, a Gaussian guidance framework that draws the depth per task from a Gaussian whose center and spread are estimated online from rollouts already collected for policy optimization, requiring no probe rollouts or learned depth predictor. The center combines a global baseline with per-cluster difficulty, and the spread tracks within-cluster variance. We evaluate Agent-G^2 on ALFWorld and WebShop on Qwen2.5-1.5B / 7B-Instruct. Agent-G^2 outperforms the strongest hint-based, hint-free, and Aux-RL baselines on ALFWorld by 2.3 / 3.9 / 7.4 points at under one-third the rollout cost of per-sample probing.

Community

Paper author Paper submitter about 7 hours ago

We propose Agent-G^2, a Gaussian guidance framework that draws the depth per task from a Gaussian whose center and spread are estimated online from rollouts already collected for policy optimization, requiring no probe rollouts or learned depth predictor.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.23318
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.23318 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers