Hugging Face Daily Papers · · 8 min read

Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<strong>AFTER</strong> is a benchmark for studying procedural memory in LLM agents: 382 realistic tasks across 6 professional roles and 22 procedural skills, mixing single- and multi-skill workflows over difficulty tiers. Its controlled splits separate how much a skill helps in its original context from how well it transfers to held-out tasks, other professional roles, and different model backbones, giving four evaluation regimes: local improvement, cross-task, cross-role, and cross-model transfer. This lets procedural-memory and skill-evolution methods be measured by whether learned skills actually generalize, treating skills as evolving artifacts learned from experience rather than static, hand-written prompts. Check full description <a href=\"https://huggingface.co/datasets/DavydenkoGr/AFTER\">here</a>.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/661a2d0af17e918674639599/90B2TuF_tZYakxohS-fQY.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/661a2d0af17e918674639599/90B2TuF_tZYakxohS-fQY.png\" alt=\"bench_overview\"></a></p>\n","updatedAt":"2026-07-01T10:44:38.574Z","author":{"_id":"661a2d0af17e918674639599","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661a2d0af17e918674639599/prBuSpEUbhxjsmAm9Fql-.jpeg","fullname":"Davydenko Grigorii","name":"DavydenkoGr","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8986873030662537},"editors":["DavydenkoGr"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/661a2d0af17e918674639599/prBuSpEUbhxjsmAm9Fql-.jpeg"],"reactions":[],"isReport":false}},{"id":"6a450157625eb0ea40de9412","author":{"_id":"658412f93a84a40185adaf37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg","fullname":"Aamer Mihaysi","name":"O96a","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-07-01T12:00:23.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"The focus on procedural memory in LLM agents is a critical pivot from simple RAG to actual operational competence. Most 'agentic' frameworks today just wrap a prompt around a tool, but the real bottleneck is the lack of a durable, reusable skill layer that doesn't require a full context window refill every time. The AFTER benchmark's approach to testing cross-role transfer is exactly the kind of engineering rigor we need to move past 'demo-ware' and toward deployable enterprise systems. If we can't quantify how a skill transfers from one professional role to another, we're just guessing at the reliability of the system. This is a solid step toward treating agent skills as first-class software artifacts rather than stochastic accidents.","html":"<p>The focus on procedural memory in LLM agents is a critical pivot from simple RAG to actual operational competence. Most 'agentic' frameworks today just wrap a prompt around a tool, but the real bottleneck is the lack of a durable, reusable skill layer that doesn't require a full context window refill every time. The AFTER benchmark's approach to testing cross-role transfer is exactly the kind of engineering rigor we need to move past 'demo-ware' and toward deployable enterprise systems. If we can't quantify how a skill transfers from one professional role to another, we're just guessing at the reliability of the system. This is a solid step toward treating agent skills as first-class software artifacts rather than stochastic accidents.</p>\n","updatedAt":"2026-07-01T12:00:23.243Z","author":{"_id":"658412f93a84a40185adaf37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg","fullname":"Aamer Mihaysi","name":"O96a","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9240942597389221},"editors":["O96a"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg"],"reactions":[],"isReport":false}},{"id":"6a45c2f4308a40a29ad3f034","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":372,"isUserFollowing":false},"createdAt":"2026-07-02T01:46:28.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision](https://huggingface.co/papers/2606.01139) (2026)\n* [SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents](https://huggingface.co/papers/2605.18693) (2026)\n* [SEAGym: An Evaluation Environment for Self-Evolving LLM Agents](https://huggingface.co/papers/2606.17546) (2026)\n* [Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows](https://huggingface.co/papers/2605.27922) (2026)\n* [SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories](https://huggingface.co/papers/2606.01311) (2026)\n* [Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents](https://huggingface.co/papers/2605.30621) (2026)\n* [SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills](https://huggingface.co/papers/2605.24117) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2606.01139\">SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2605.18693\">SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.17546\">SEAGym: An Evaluation Environment for Self-Evolving LLM Agents</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2605.27922\">Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.01311\">SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2605.30621\">Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2605.24117\">SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-07-02T01:46:28.384Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":372,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7396476864814758},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2606.23127","authors":[{"_id":"6a438139763f63ca3757eb84","name":"Julia Belikova","hidden":false},{"_id":"6a438139763f63ca3757eb85","name":"Rauf Parchiev","hidden":false},{"_id":"6a438139763f63ca3757eb86","name":"Evgeny Egorov","hidden":false},{"_id":"6a438139763f63ca3757eb87","user":{"_id":"661a2d0af17e918674639599","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661a2d0af17e918674639599/prBuSpEUbhxjsmAm9Fql-.jpeg","isPro":false,"fullname":"Davydenko Grigorii","user":"DavydenkoGr","type":"user","name":"DavydenkoGr"},"name":"Grigorii Davydenko","status":"claimed_verified","statusLastChangedAt":"2026-07-01T08:45:55.637Z","hidden":false},{"_id":"6a438139763f63ca3757eb88","name":"Gleb Gusev","hidden":false},{"_id":"6a438139763f63ca3757eb89","name":"Andrey Savchenko","hidden":false},{"_id":"6a438139763f63ca3757eb8a","user":{"_id":"683c210c9c6d4e639fa827a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/6-uRkimtBedgU4eh85e9k.png","isPro":false,"fullname":"Maksim Makarenko","user":"makamoa","type":"user","name":"makamoa"},"name":"Maksim Makarenko","status":"claimed_verified","statusLastChangedAt":"2026-07-01T08:45:53.819Z","hidden":false}],"publishedAt":"2026-06-22T00:00:00.000Z","submittedOnDailyAt":"2026-07-01T00:00:00.000Z","title":"Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation","submittedOnDailyBy":{"_id":"661a2d0af17e918674639599","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661a2d0af17e918674639599/prBuSpEUbhxjsmAm9Fql-.jpeg","isPro":false,"fullname":"Davydenko Grigorii","user":"DavydenkoGr","type":"user","name":"DavydenkoGr"},"summary":"Procedural memory is increasingly used to improve LLM agents on recurring workplace tasks, yet its ability to produce reusable skills remains poorly understood. We introduce AFTER, a benchmark of 382 realistic enterprise tasks spanning six professional roles and 22 procedural skills, designed to evaluate how skills transfer across tasks, roles, and model backbones. The benchmark includes controlled evaluation settings for local improvement, cross-task transfer, cross-role transfer, and cross-model generalization. Experiments show that procedural memory delivers consistent gains in industrial workflows: a single refinement round improves aggregate performance by 3.7-6.7 points, while skills evolved from diverse multi-model execution traces achieve 73.1% cross-model test accuracy, outperforming all single-model trace sources. We further find that some skills generalize broadly across tasks and models, whereas others become specialized to role-specific workflows and lose effectiveness under transfer. These results provide practical guidance for building, evaluating, and deploying procedural memory systems in production agent platforms.","upvotes":17,"discussionId":"6a438139763f63ca3757eb8b","githubRepo":"https://github.com/DavydenkoGr/AFTER","githubRepoAddedBy":"user","ai_summary":"Procedural memory enhances LLM agents on workplace tasks through skill transfer across roles and models, with varying generalization capabilities affecting deployment strategies.","ai_keywords":["procedural memory","LLM agents","enterprise tasks","procedural skills","cross-task transfer","cross-role transfer","cross-model generalization","skill evolution","multi-model execution traces","aggregate performance"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":2},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"683c210c9c6d4e639fa827a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/6-uRkimtBedgU4eh85e9k.png","isPro":false,"fullname":"Maksim Makarenko","user":"makamoa","type":"user"},{"_id":"661a2d0af17e918674639599","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661a2d0af17e918674639599/prBuSpEUbhxjsmAm9Fql-.jpeg","isPro":false,"fullname":"Davydenko Grigorii","user":"DavydenkoGr","type":"user"},{"_id":"6270324ebecab9e2dcf245de","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6270324ebecab9e2dcf245de/cMbtWSasyNlYc9hvsEEzt.jpeg","isPro":false,"fullname":"Kye Gomez","user":"kye","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"67c7ee99567ddbad6f86be3d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/nvTEyXOl6_Z89taJBD_ol.png","isPro":false,"fullname":"Ivan","user":"ivio05","type":"user"},{"_id":"6904703c9c3d523381fe6220","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6904703c9c3d523381fe6220/tEiI8NYsSz5li4v2G31-r.png","isPro":false,"fullname":"Nail","user":"gafurovnail","type":"user"},{"_id":"65cdde3046a98b4598e69d00","avatarUrl":"/avatars/9f61e0c9e5754737d3cc71c8437eacf9.svg","isPro":false,"fullname":"Ivan Poddiakov","user":"ivanpodd","type":"user"},{"_id":"6a197310633544c42baac073","avatarUrl":"/avatars/ac044a09c33fe67718e342f7547e32f0.svg","isPro":false,"fullname":"Artem Sakhno","user":"warofgam","type":"user"},{"_id":"683485048ec8d9888bd1cd6e","avatarUrl":"/avatars/ef7e48e605632fc6bfed06545f1f7286.svg","isPro":false,"fullname":"Milhail Filimonov","user":"advanced12iq","type":"user"},{"_id":"654a49d76167ff03f7ff1638","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/0q1PM80eI2F8aYFcfA2RT.jpeg","isPro":false,"fullname":"Ruslan","user":"Karifannaa","type":"user"},{"_id":"643fd8f427dc46cca58c661e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/klNlGqneH-2Dzvm1fqJ57.jpeg","isPro":false,"fullname":"Parchiev Rauf","user":"parchiev","type":"user"},{"_id":"61dedb1b2066746d68b63adb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/61dedb1b2066746d68b63adb/PLBHQxvbcay3qDJjY7HM3.jpeg","isPro":false,"fullname":"Ivan Sviridov","user":"univanxx","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2606/2606.23127.md","query":{}}">
Papers
arxiv:2606.23127

Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation

Published on Jun 22
· Submitted by
Davydenko Grigorii
on Jul 1
Authors:
,
,
,
,
,

Abstract

Procedural memory enhances LLM agents on workplace tasks through skill transfer across roles and models, with varying generalization capabilities affecting deployment strategies.

Procedural memory is increasingly used to improve LLM agents on recurring workplace tasks, yet its ability to produce reusable skills remains poorly understood. We introduce AFTER, a benchmark of 382 realistic enterprise tasks spanning six professional roles and 22 procedural skills, designed to evaluate how skills transfer across tasks, roles, and model backbones. The benchmark includes controlled evaluation settings for local improvement, cross-task transfer, cross-role transfer, and cross-model generalization. Experiments show that procedural memory delivers consistent gains in industrial workflows: a single refinement round improves aggregate performance by 3.7-6.7 points, while skills evolved from diverse multi-model execution traces achieve 73.1% cross-model test accuracy, outperforming all single-model trace sources. We further find that some skills generalize broadly across tasks and models, whereas others become specialized to role-specific workflows and lose effectiveness under transfer. These results provide practical guidance for building, evaluating, and deploying procedural memory systems in production agent platforms.

Community

Paper author Paper submitter about 15 hours ago

AFTER is a benchmark for studying procedural memory in LLM agents: 382 realistic tasks across 6 professional roles and 22 procedural skills, mixing single- and multi-skill workflows over difficulty tiers. Its controlled splits separate how much a skill helps in its original context from how well it transfers to held-out tasks, other professional roles, and different model backbones, giving four evaluation regimes: local improvement, cross-task, cross-role, and cross-model transfer. This lets procedural-memory and skill-evolution methods be measured by whether learned skills actually generalize, treating skills as evolving artifacts learned from experience rather than static, hand-written prompts. Check full description here.

bench_overview

The focus on procedural memory in LLM agents is a critical pivot from simple RAG to actual operational competence. Most 'agentic' frameworks today just wrap a prompt around a tool, but the real bottleneck is the lack of a durable, reusable skill layer that doesn't require a full context window refill every time. The AFTER benchmark's approach to testing cross-role transfer is exactly the kind of engineering rigor we need to move past 'demo-ware' and toward deployable enterprise systems. If we can't quantify how a skill transfers from one professional role to another, we're just guessing at the reliability of the system. This is a solid step toward treating agent skills as first-class software artifacts rather than stochastic accidents.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.23127
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2606.23127 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.23127 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers