Hugging Face Daily Papers · · 4 min read

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We propose PILOT, a supervisor–worker harness that makes agent self-improvement live rather than post-hoc. Two coupled mechanisms: a separate supervisor redirects or aborts the active worker mid-run (live steering), while runtime-discovered procedures and failure modes are distilled into reusable skills and memory (live self-evolution). With frozen GLM-5.1 and Kimi-K2.6 backbones, PILOT ranks first in 5 of 6 configurations across three benchmarks. In the self-improvement setting it gains +14.6 / +12.4 points while cutting mean output tokens by 42.9% / 47.4%, and successful evaluations per million tokens rise by 110.3% / 134.0%.</p>\n<p>Code will be released soon on GitHub.</p>\n","updatedAt":"2026-08-28T05:26:05.892Z","author":{"_id":"6002c316698168af3bb9f4a6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6002c316698168af3bb9f4a6/M2J2QFCRc5RhYRoz7OuZN.png","fullname":"yangxiao","name":"YangXiao-nlp","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":7,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8757508993148804},"editors":["YangXiao-nlp"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6002c316698168af3bb9f4a6/M2J2QFCRc5RhYRoz7OuZN.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.26530","authors":[{"_id":"6a90f454a64059bab69c3609","user":{"_id":"6002c316698168af3bb9f4a6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6002c316698168af3bb9f4a6/M2J2QFCRc5RhYRoz7OuZN.png","isPro":false,"fullname":"yangxiao","user":"YangXiao-nlp","type":"user","name":"YangXiao-nlp"},"name":"Yang Xiao","status":"claimed_verified","statusLastChangedAt":"2026-08-28T08:45:04.622Z","hidden":false},{"_id":"6a90f454a64059bab69c360a","name":"Yusong Sun","hidden":false},{"_id":"6a90f454a64059bab69c360b","name":"Haoyi Wu","hidden":false},{"_id":"6a90f454a64059bab69c360c","name":"Wenyang Hui","hidden":false},{"_id":"6a90f454a64059bab69c360d","name":"Wen Da","hidden":false},{"_id":"6a90f454a64059bab69c360e","name":"Zhaokai Luo","hidden":false},{"_id":"6a90f454a64059bab69c360f","name":"Mu Chuan","hidden":false},{"_id":"6a90f454a64059bab69c3610","name":"Yao Hu","hidden":false},{"_id":"6a90f454a64059bab69c3611","name":"Wenjie Li","hidden":false},{"_id":"6a90f454a64059bab69c3612","name":"Chengyue Jiang","hidden":false}],"publishedAt":"2026-08-27T00:00:00.000Z","submittedOnDailyAt":"2026-08-28T00:00:00.000Z","title":"PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents","submittedOnDailyBy":{"_id":"6002c316698168af3bb9f4a6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6002c316698168af3bb9f4a6/M2J2QFCRc5RhYRoz7OuZN.png","isPro":false,"fullname":"yangxiao","user":"YangXiao-nlp","type":"user","name":"YangXiao-nlp"},"summary":"Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvement methods process this experience only after execution ends, so they cannot redirect the active run or immediately apply and validate lessons learned from it. We argue that self-improvement should instead be live, using emerging experience both to redirect the active run and to update the persistent harness. Existing agent architectures do not fully support this goal. Single-agent self-correction combines task execution and trajectory assessment within one context, while subagent delegation separates execution but typically cannot redirect an active subagent. We present PILOT, a supervisor-worker harness for live self-improvement through two coupled mechanisms: (1) live steering lets a separate supervisor redirect or abort the active worker during execution; and (2) live self-evolution distils procedures and failure modes revealed during execution into reusable skills and memory. Across two frozen backbones and three benchmarks, PILOT ranks first in five of six configurations. On Terminal-Bench 2.0, PILOT outperforms counterpart harnesses by up to 9.8 percentage points. In the self-improvement setting, PILOT gains 14.6 points with GLM-5.1 and 12.4 points with Kimi-K2.6. Mean output tokens fall by 42.9% and 47.4%, while successful evaluations per million output tokens rise by 110.3% and 134.0%, respectively.","upvotes":21,"discussionId":"6a90f454a64059bab69c3613","ai_summary":"PILOT enables live self-improvement by allowing a supervisor to steer active workers and distilling execution experience into reusable skills, improving accuracy and efficiency.","ai_keywords":["live steering","supervisor-worker harness","live self-evolution","reusable skills","memory","self-correction","subagent delegation"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"646ecc368d316fde87b3b6e3","name":"PolyUHK","fullname":"The Hong Kong Polytechnic University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/646ecbc0cbb7bb996513e298/Akb4zKqIP9kb9PQoUPUmj.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6002c316698168af3bb9f4a6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6002c316698168af3bb9f4a6/M2J2QFCRc5RhYRoz7OuZN.png","isPro":false,"fullname":"yangxiao","user":"YangXiao-nlp","type":"user"},{"_id":"69537e9c313d96feb976d0d1","avatarUrl":"/avatars/80a8295a67d4d7dfec80ab8359bd5612.svg","isPro":false,"fullname":"Yusong Sun","user":"YusongSun","type":"user"},{"_id":"64fd753e5ca946a010d61b54","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/HbbwJZqxghoR-57v3Bq8l.jpeg","isPro":false,"fullname":"Hui Zhang","user":"Julie11","type":"user"},{"_id":"6596534fd61d4cbc1bf08683","avatarUrl":"/avatars/ab37f077bcebccb52a1eb2b0b7567552.svg","isPro":false,"fullname":"Jiashuo WANG","user":"Jessie09","type":"user"},{"_id":"64e6c617ecce34cb442cb208","avatarUrl":"/avatars/ebc61bf6a043314cb2089b1efd5e6a18.svg","isPro":false,"fullname":"JieSun(SII)","user":"Sunshine279","type":"user"},{"_id":"6a911cdb8f844ed4db749a6a","avatarUrl":"/avatars/877f0a6b62535b8cb15f0af9a2c776ab.svg","isPro":false,"fullname":"Jeff","user":"Jeff0628","type":"user"},{"_id":"671b8777ac4168f79848b282","avatarUrl":"/avatars/b4b8062a8fe890fb6bac8917630bfb5a.svg","isPro":false,"fullname":"Zechen Sun","user":"Mintszc","type":"user"},{"_id":"6a911df0dd062b16f11e400e","avatarUrl":"/avatars/da2dc0be8747cc4f9f69b0e42330f995.svg","isPro":false,"fullname":"huiwy","user":"huiwy1","type":"user"},{"_id":"6469f91531fbfc5df81b9154","avatarUrl":"/avatars/20658b7341f181dbe702b4665d08780d.svg","isPro":false,"fullname":"whywhy","user":"whynlp","type":"user"},{"_id":"6a911ff3c72880ee4efb22e1","avatarUrl":"/avatars/7405273092bd756a0f8584c363414382.svg","isPro":false,"fullname":"Reiko","user":"Reiko777","type":"user"},{"_id":"69a6944d0a20e2f2f984aaee","avatarUrl":"/avatars/1c60bf82930125e2cb7afe0b94ca4288.svg","isPro":false,"fullname":"yuaofan","user":"AofaYu71","type":"user"},{"_id":"630341a4ef6f432d84e97db9","avatarUrl":"/avatars/dc9910b5b9053123bad5ff848f677e12.svg","isPro":false,"fullname":"tomsawyer","user":"tomhu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"646ecc368d316fde87b3b6e3","name":"PolyUHK","fullname":"The Hong Kong Polytechnic University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/646ecbc0cbb7bb996513e298/Akb4zKqIP9kb9PQoUPUmj.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.26530.md","query":{}}">
Papers
arxiv:2608.26530

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents

Published on Aug 27
· Submitted by
yangxiao
on Aug 28
Authors:

Abstract

PILOT enables live self-improvement by allowing a supervisor to steer active workers and distilling execution experience into reusable skills, improving accuracy and efficiency.

Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvement methods process this experience only after execution ends, so they cannot redirect the active run or immediately apply and validate lessons learned from it. We argue that self-improvement should instead be live, using emerging experience both to redirect the active run and to update the persistent harness. Existing agent architectures do not fully support this goal. Single-agent self-correction combines task execution and trajectory assessment within one context, while subagent delegation separates execution but typically cannot redirect an active subagent. We present PILOT, a supervisor-worker harness for live self-improvement through two coupled mechanisms: (1) live steering lets a separate supervisor redirect or abort the active worker during execution; and (2) live self-evolution distils procedures and failure modes revealed during execution into reusable skills and memory. Across two frozen backbones and three benchmarks, PILOT ranks first in five of six configurations. On Terminal-Bench 2.0, PILOT outperforms counterpart harnesses by up to 9.8 percentage points. In the self-improvement setting, PILOT gains 14.6 points with GLM-5.1 and 12.4 points with Kimi-K2.6. Mean output tokens fall by 42.9% and 47.4%, while successful evaluations per million output tokens rise by 110.3% and 134.0%, respectively.

Community

Paper author Paper submitter about 5 hours ago

We propose PILOT, a supervisor–worker harness that makes agent self-improvement live rather than post-hoc. Two coupled mechanisms: a separate supervisor redirects or aborts the active worker mid-run (live steering), while runtime-discovered procedures and failure modes are distilled into reusable skills and memory (live self-evolution). With frozen GLM-5.1 and Kimi-K2.6 backbones, PILOT ranks first in 5 of 6 configurations across three benchmarks. In the self-improvement setting it gains +14.6 / +12.4 points while cutting mean output tokens by 42.9% / 47.4%, and successful evaluations per million tokens rise by 110.3% / 134.0%.

Code will be released soon on GitHub.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.26530
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.26530 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.26530 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.26530 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers