<a href=\"https://github.com/Gen-Verse/Recuris\" rel=\"nofollow\">https://github.com/Gen-Verse/Recuris</a></p>\n","updatedAt":"2026-08-26T02:16:56.078Z","author":{"_id":"64fde4e252e82dd432b74ce9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64fde4e252e82dd432b74ce9/-CQZbBP7FsPPyawYrsi4z.jpeg","fullname":"Ling Yang","name":"Lingaaaaaaa","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":16,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.764903724193573},"editors":["Lingaaaaaaa"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64fde4e252e82dd432b74ce9/-CQZbBP7FsPPyawYrsi4z.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.24876","authors":[{"_id":"6a8e4c657bc881afa25f3091","name":"Zhaochen Yu","hidden":false},{"_id":"6a8e4c657bc881afa25f3092","name":"Yingcheng Wu","hidden":false},{"_id":"6a8e4c657bc881afa25f3093","name":"Zhenfei Yin","hidden":false},{"_id":"6a8e4c657bc881afa25f3094","name":"Kaiyuan Chen","hidden":false},{"_id":"6a8e4c657bc881afa25f3095","name":"Zhe Zhao","hidden":false},{"_id":"6a8e4c657bc881afa25f3096","name":"Mengdi Wang","hidden":false},{"_id":"6a8e4c657bc881afa25f3097","name":"Shuicheng Yan","hidden":false},{"_id":"6a8e4c657bc881afa25f3098","name":"Ling Yang","hidden":false}],"publishedAt":"2026-08-25T00:00:00.000Z","submittedOnDailyAt":"2026-08-26T00:00:00.000Z","title":"Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses","submittedOnDailyBy":{"_id":"64fde4e252e82dd432b74ce9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64fde4e252e82dd432b74ce9/-CQZbBP7FsPPyawYrsi4z.jpeg","isPro":false,"fullname":"Ling Yang","user":"Lingaaaaaaa","type":"user","name":"Lingaaaaaaa"},"summary":"Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation. We introduce Recuris, a recursive Experiential-Working Memory architecture for long-horizon agent harnesses, in which Working Memory tracks task progress and guides skill selection from Experiential Memory, grounding skill use in current needs rather than the full history. This coupling also turns execution into structured evidence that localizes failures to specific memory components. Across tasks, a fixed Meta-Agent turns that evidence into localized, validation-gated updates to Skill Memory that reshape execution and yield new evidence, forming a bounded recursive memory-evolution loop. Across four long-horizon benchmarks and ten models, Recuris improves task success in 35 of the 37 completed model-benchmark pairs, carrying frontier models to SOTA-level task success: on tau-bench it adds +17.8 points to GPT-5.6 Sol and +15.6 to Claude Opus 5, taking Opus 5 to 87.9%, and +16.6/+13.5 points on Qwen3.6-27B/35B on SkillFlow. The advantage widens as the interaction horizon grows, to +32.2 points on the longest tasks, and common long-horizon failures fall by up to 80%. These results position recursively evolving memory as a scalable foundation for RSI, enabling agents to continuously transform accumulated experience into increasingly effective long-horizon behavior. Code: https://github.com/Gen-Verse/Recuris","upvotes":15,"discussionId":"6a8e4c657bc881afa25f3099","projectPage":"https://github.com/Gen-Verse/Recuris","githubRepo":"https://github.com/Gen-Verse/Recuris","githubRepoAddedBy":"user","ai_summary":"Recuris introduces a recursive memory architecture that tracks progress and guides skill selection to improve long-horizon agent success through localized, validation-gated updates.","ai_keywords":["recursive self-improvement","Experiential-Working Memory","Working Memory","Experiential Memory","Skill Memory","Meta-Agent","validation-gated updates","long-horizon agent harnesses"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":4,"organization":{"_id":"64374111a701a7e744c02b0e","name":"princetonu","fullname":"Princeton University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/b3xXusq8Zz3ej8Z6fRTSZ.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64fde4e252e82dd432b74ce9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64fde4e252e82dd432b74ce9/-CQZbBP7FsPPyawYrsi4z.jpeg","isPro":false,"fullname":"Ling Yang","user":"Lingaaaaaaa","type":"user"},{"_id":"6662a23c2f86097c6d828b96","avatarUrl":"/avatars/2aa31ab30874257529861f2e4024acc2.svg","isPro":false,"fullname":"liu","user":"miao6","type":"user"},{"_id":"69817bd8a819c22bf570fae4","avatarUrl":"/avatars/00642747bfb2bcf443165ae7a4175c0c.svg","isPro":false,"fullname":"Yang","user":"TonyYang1","type":"user"},{"_id":"6981942081c01373225279f5","avatarUrl":"/avatars/b84b75986a3889ccb241b8a25cb09df8.svg","isPro":false,"fullname":"ma","user":"sherryma23","type":"user"},{"_id":"6662a59cf8d1fcc749cbc5de","avatarUrl":"/avatars/0e965b6b996c154b8d39106c0cc5178d.svg","isPro":false,"fullname":"liu","user":"miao99","type":"user"},{"_id":"6981928e9dcf301b73fe1fcb","avatarUrl":"/avatars/42d73f5498c65a4ebe67c168abafa218.svg","isPro":false,"fullname":"yang","user":"martinyang1","type":"user"},{"_id":"641b26ed1911d3be6743e8d0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/641b26ed1911d3be6743e8d0/oybjQjcEgiMBC-qGKCQVR.png","isPro":false,"fullname":"余昭辰","user":"BitStarWalkin","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"698196075cc02c68e5555ea5","avatarUrl":"/avatars/b36383b891f37275386b18f19d0a6ef3.svg","isPro":false,"fullname":"liu","user":"2linda","type":"user"},{"_id":"6a8e4d98a6cfbc25447acd4b","avatarUrl":"/avatars/c6098454bad143112ef467791e0c6795.svg","isPro":false,"fullname":"zcy","user":"trafalgaryu","type":"user"},{"_id":"6a8e4ddef8b010c865604e6f","avatarUrl":"/avatars/e65b79944fa205555d8ceaeadcc6d24b.svg","isPro":false,"fullname":"sada","user":"fakersda","type":"user"},{"_id":"6a8e4e3d86120332b78d4306","avatarUrl":"/avatars/cb89da704273577a0e62852a8d1ea193.svg","isPro":false,"fullname":"Johnny","user":"Johnny872","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"64374111a701a7e744c02b0e","name":"princetonu","fullname":"Princeton University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/b3xXusq8Zz3ej8Z6fRTSZ.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.24876.md","query":{}}">
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
Abstract
Recuris introduces a recursive memory architecture that tracks progress and guides skill selection to improve long-horizon agent success through localized, validation-gated updates.
Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation. We introduce Recuris, a recursive Experiential-Working Memory architecture for long-horizon agent harnesses, in which Working Memory tracks task progress and guides skill selection from Experiential Memory, grounding skill use in current needs rather than the full history. This coupling also turns execution into structured evidence that localizes failures to specific memory components. Across tasks, a fixed Meta-Agent turns that evidence into localized, validation-gated updates to Skill Memory that reshape execution and yield new evidence, forming a bounded recursive memory-evolution loop. Across four long-horizon benchmarks and ten models, Recuris improves task success in 35 of the 37 completed model-benchmark pairs, carrying frontier models to SOTA-level task success: on tau-bench it adds +17.8 points to GPT-5.6 Sol and +15.6 to Claude Opus 5, taking Opus 5 to 87.9%, and +16.6/+13.5 points on Qwen3.6-27B/35B on SkillFlow. The advantage widens as the interaction horizon grows, to +32.2 points on the longest tasks, and common long-horizon failures fall by up to 80%. These results position recursively evolving memory as a scalable foundation for RSI, enabling agents to continuously transform accumulated experience into increasingly effective long-horizon behavior. Code: https://github.com/Gen-Verse/Recuris
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.24876 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.24876 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.24876 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.