Hugging Face Daily Papers · · 4 min read

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Up~</p>\n","updatedAt":"2026-08-14T05:01:36.806Z","author":{"_id":"6a798ec103457d254b610960","avatarUrl":"/avatars/dce4680e6f4068243f1f6ed4b94df494.svg","fullname":"yan","name":"Nikoyan","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"ja","probability":0.9383794069290161},"editors":["Nikoyan"],"editorAvatarUrls":["/avatars/dce4680e6f4068243f1f6ed4b94df494.svg"],"reactions":[],"isReport":false}},{"id":"6a87cc2a36e378ccd79f4bbc","author":{"_id":"68fc3ddcdc9e5cbf49cbc716","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68fc3ddcdc9e5cbf49cbc716/gzvksq-XgWnekB6Xl25pw.jpeg","fullname":"EasonYe","name":"EasonUwU","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-08-21T03:55:22.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Nice paper~","html":"<p>Nice paper~</p>\n","updatedAt":"2026-08-21T03:55:22.635Z","author":{"_id":"68fc3ddcdc9e5cbf49cbc716","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68fc3ddcdc9e5cbf49cbc716/gzvksq-XgWnekB6Xl25pw.jpeg","fullname":"EasonYe","name":"EasonUwU","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"ru","probability":0.23798467218875885},"editors":["EasonUwU"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/68fc3ddcdc9e5cbf49cbc716/gzvksq-XgWnekB6Xl25pw.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.13120","authors":[{"_id":"6a7e812a42823931a1f1765c","name":"Qianxi Yan","hidden":false},{"_id":"6a7e812a42823931a1f1765d","name":"Chunrong Chen","hidden":false},{"_id":"6a7e812a42823931a1f1765e","name":"Jiuzhou Zhao","hidden":false},{"_id":"6a7e812a42823931a1f1765f","name":"Min Zhang","hidden":false},{"_id":"6a7e812a42823931a1f17660","name":"Yongzhou Xu","hidden":false},{"_id":"6a7e812a42823931a1f17661","name":"Xiaochuan Xu","hidden":false}],"publishedAt":"2026-08-13T00:00:00.000Z","submittedOnDailyAt":"2026-08-21T00:00:00.000Z","title":"SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback","submittedOnDailyBy":{"_id":"68fc3ddcdc9e5cbf49cbc716","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68fc3ddcdc9e5cbf49cbc716/gzvksq-XgWnekB6Xl25pw.jpeg","isPro":false,"fullname":"EasonYe","user":"EasonUwU","type":"user","name":"EasonUwU"},"summary":"Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only across multiple turns remain invisible, and evolution stalls. Governance in these systems is likewise driven by an end-to-end verification score, a scalar gate that can reject a degraded candidate but can neither localize nor repair its structural cause. We argue that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients. We introduce SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction. The first component recasts multi-turn user simulation from an evaluation endpoint into a feedback generator: follow-up questions expose defects layer by layer, so that every round of revision both consumes feedback and produces new feedback. The second replaces the passive rejection of a scalar gate with an independent governance layer that actively repairs factual degradation and structural bloat, preventing the gradient from drifting as degradation accumulates. Across six categories of cloud services, 9 production Skills, and 98 skill-reference files, SkillEvo surpasses self-reflection-based evolution by 23.0 points and single- turn-QA-driven evolution by 15.4 points.","upvotes":6,"discussionId":"6a7e812a42823931a1f17662","ai_summary":"SkillEvo improves agent skills through multi-turn feedback and active governance to sustain evolution gradients.","ai_keywords":["multi-turn user simulation","feedback generator","governance layer","evolution gradients","self-reflection","single-turn QA"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"61bac2af530e5c78d7b99667","name":"zju","fullname":"Zhejiang University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5e1058e9fcf41d740b69966d/7G1xjlxwCdMEmKcxNR0n5.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"651c0477e315e8a224ad8824","avatarUrl":"/avatars/276239281d452e8005972d02a3bcf9e2.svg","isPro":false,"fullname":"charent","user":"charent","type":"user"},{"_id":"6a798ec103457d254b610960","avatarUrl":"/avatars/dce4680e6f4068243f1f6ed4b94df494.svg","isPro":false,"fullname":"yan","user":"Nikoyan","type":"user"},{"_id":"68fc3ddcdc9e5cbf49cbc716","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68fc3ddcdc9e5cbf49cbc716/gzvksq-XgWnekB6Xl25pw.jpeg","isPro":false,"fullname":"EasonYe","user":"EasonUwU","type":"user"},{"_id":"66d8512c54209e9101811e8e","avatarUrl":"/avatars/62dfd8e6261108f2508efe678d5a2a57.svg","isPro":false,"fullname":"M Saad Salman","user":"MSS444","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"62f6065049c083e86649c9db","avatarUrl":"/avatars/8c329eab4e2216c2c51565b48376cb03.svg","isPro":false,"fullname":"joskazhao","user":"joska","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61bac2af530e5c78d7b99667","name":"zju","fullname":"Zhejiang University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5e1058e9fcf41d740b69966d/7G1xjlxwCdMEmKcxNR0n5.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.13120.md","query":{}}">
Papers
arxiv:2608.13120

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

Published on Aug 13
· Submitted by
EasonYe
on Aug 21
Authors:
,

Abstract

SkillEvo improves agent skills through multi-turn feedback and active governance to sustain evolution gradients.

Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only across multiple turns remain invisible, and evolution stalls. Governance in these systems is likewise driven by an end-to-end verification score, a scalar gate that can reject a degraded candidate but can neither localize nor repair its structural cause. We argue that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients. We introduce SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction. The first component recasts multi-turn user simulation from an evaluation endpoint into a feedback generator: follow-up questions expose defects layer by layer, so that every round of revision both consumes feedback and produces new feedback. The second replaces the passive rejection of a scalar gate with an independent governance layer that actively repairs factual degradation and structural bloat, preventing the gradient from drifting as degradation accumulates. Across six categories of cloud services, 9 production Skills, and 98 skill-reference files, SkillEvo surpasses self-reflection-based evolution by 23.0 points and single- turn-QA-driven evolution by 15.4 points.

Community

Paper submitter about 4 hours ago

Nice paper~

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.13120
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.13120 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.13120 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.13120 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers