Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing editors inject such a passage yet fail to use it: the edited model can recall the passage, but can neither answer atomic questions about its facts nor compose them into multi-hop reasoning. We attribute this missing property, which we term composability, to editors' passive reliance on the fixed passage as the sole learning source. In response, we cast editing as a proactive self-distillation from a privileged in-context state of the same model, which requires no external supervision. We further reveal that due to the novelty of the injected knowledge, the pre-edited model's own rollouts rarely cover it, which limits the effectiveness of pure on-policy distillation. To close this gap, we propose HPSE, which builds a hybrid rollout that steps in to place missing facts onto the student's own trajectory precisely where its coverage fails, while staying on-policy elsewhere. We theoretically analyze HPSE's advantage over pure on-policy distillation, and empirically establish its plug-and-play improvements across four LLM backbones and two KE editors under various scenarios.</p>\n","updatedAt":"2026-08-14T15:01:18.116Z","author":{"_id":"64bf811d76a6e2efcceabc00","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64bf811d76a6e2efcceabc00/0p3zSIVqzoME25Zmfh7SD.png","fullname":"Tianci Liu","name":"lliutianc","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9275947213172913},"editors":["lliutianc"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64bf811d76a6e2efcceabc00/0p3zSIVqzoME25Zmfh7SD.png"],"reactions":[],"isReport":false}},{"id":"6a7fc1f472f87f0631aab30f","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-15T01:33:40.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [KARLA: Knowledge-base Augmented Retrieval for Language Models](https://huggingface.co/papers/2606.26807) (2026)\n* [PRISM Edit: One Vector for All Temporal Answers](https://huggingface.co/papers/2607.11327) (2026)\n* [On-Policy Self-Distillation without Any Supervision](https://huggingface.co/papers/2608.06296) (2026)\n* [Regime-Aware Peer Specialization for Robust RAG under Heterogeneous Knowledge Conflicts](https://huggingface.co/papers/2606.30518) (2026)\n* [Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing](https://huggingface.co/papers/2607.20433) (2026)\n* [LeAct: Learning to Reason from Expert Actions](https://huggingface.co/papers/2607.21856) (2026)\n* [Training-Free Token-Level Steering for LLM Personalized Co-Writing](https://huggingface.co/papers/2608.06069) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2606.26807\">KARLA: Knowledge-base Augmented Retrieval for Language Models</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.11327\">PRISM Edit: One Vector for All Temporal Answers</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.06296\">On-Policy Self-Distillation without Any Supervision</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.30518\">Regime-Aware Peer Specialization for Robust RAG under Heterogeneous Knowledge Conflicts</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.20433\">Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.21856\">LeAct: Learning to Reason from Expert Actions</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.06069\">Training-Free Token-Level Steering for LLM Personalized Co-Writing</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-15T01:33:40.619Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7320732474327087},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.11660","authors":[{"_id":"6a7f2db5f747ea94019af4e8","user":{"_id":"64bf811d76a6e2efcceabc00","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64bf811d76a6e2efcceabc00/0p3zSIVqzoME25Zmfh7SD.png","isPro":false,"fullname":"Tianci Liu","user":"lliutianc","type":"user","name":"lliutianc"},"name":"Tianci Liu","status":"claimed_verified","statusLastChangedAt":"2026-08-15T08:45:04.160Z","hidden":false},{"_id":"6a7f2db5f747ea94019af4e9","name":"Zihan Dong","hidden":false},{"_id":"6a7f2db5f747ea94019af4ea","name":"Tianchun Li","hidden":false},{"_id":"6a7f2db5f747ea94019af4eb","name":"Yi-Chung Chen","hidden":false},{"_id":"6a7f2db5f747ea94019af4ec","name":"Qiming Cao","hidden":false},{"_id":"6a7f2db5f747ea94019af4ed","name":"Xingchen Wang","hidden":false},{"_id":"6a7f2db5f747ea94019af4ee","name":"Shiyang Wang","hidden":false},{"_id":"6a7f2db5f747ea94019af4ef","name":"Zichen Miao","hidden":false},{"_id":"6a7f2db5f747ea94019af4f0","name":"Linjun Zhang","hidden":false},{"_id":"6a7f2db5f747ea94019af4f1","name":"Haoyu Wang","hidden":false},{"_id":"6a7f2db5f747ea94019af4f2","name":"Jing Gao","hidden":false}],"publishedAt":"2026-08-12T00:00:00.000Z","submittedOnDailyAt":"2026-08-14T00:00:00.000Z","title":"Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing","submittedOnDailyBy":{"_id":"64bf811d76a6e2efcceabc00","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64bf811d76a6e2efcceabc00/0p3zSIVqzoME25Zmfh7SD.png","isPro":false,"fullname":"Tianci Liu","user":"lliutianc","type":"user","name":"lliutianc"},"summary":"Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing editors inject such a passage yet fail to use it: the edited model can recall the passage, but can neither answer atomic questions about its facts nor compose them into multi-hop reasoning. We attribute this missing property, which we term composability, to editors' passive reliance on the fixed passage as the sole learning source. In response, we cast editing as a proactive self-distillation from a privileged in-context state of the same model, which requires no external supervision. We further reveal that due to the novelty of the injected knowledge, the pre-edited model's own rollouts rarely cover it, which limits the effectiveness of pure on-policy distillation. To close this gap, we propose HPSE, which builds a hybrid rollout that steps in to place missing facts onto the student's own trajectory precisely where its coverage fails, while staying on-policy elsewhere. We theoretically analyze HPSE's advantage over pure on-policy distillation, and empirically establish its plug-and-play improvements across four LLM backbones and two KE editors under various scenarios.","upvotes":4,"discussionId":"6a7f2db5f747ea94019af4f3","ai_summary":"HPSE improves unstructured knowledge editing by distilling from hybrid rollouts that insert missing facts into the model's reasoning paths, enabling composable multi-hop reasoning.","ai_keywords":["large language models","knowledge editing","unstructured knowledge editing","composability","self-distillation","in-context learning","hybrid rollout","on-policy distillation","multi-hop reasoning"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64bf811d76a6e2efcceabc00","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64bf811d76a6e2efcceabc00/0p3zSIVqzoME25Zmfh7SD.png","isPro":false,"fullname":"Tianci Liu","user":"lliutianc","type":"user"},{"_id":"661ab1f1fa3b144a381fa454","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661ab1f1fa3b144a381fa454/IlpZBb9NCjo7ntFwMIH53.png","isPro":false,"fullname":"Urro","user":"urroxyz","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"631e14ac473a6825f285e89d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/631e14ac473a6825f285e89d/K-6QnoeGLg8XFvbTMMdqA.jpeg","isPro":false,"fullname":"Yury Panikov","user":"panikov","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.11660.md","query":{}}">
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing
Abstract
HPSE improves unstructured knowledge editing by distilling from hybrid rollouts that insert missing facts into the model's reasoning paths, enabling composable multi-hop reasoning.
Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing editors inject such a passage yet fail to use it: the edited model can recall the passage, but can neither answer atomic questions about its facts nor compose them into multi-hop reasoning. We attribute this missing property, which we term composability, to editors' passive reliance on the fixed passage as the sole learning source. In response, we cast editing as a proactive self-distillation from a privileged in-context state of the same model, which requires no external supervision. We further reveal that due to the novelty of the injected knowledge, the pre-edited model's own rollouts rarely cover it, which limits the effectiveness of pure on-policy distillation. To close this gap, we propose HPSE, which builds a hybrid rollout that steps in to place missing facts onto the student's own trajectory precisely where its coverage fails, while staying on-policy elsewhere. We theoretically analyze HPSE's advantage over pure on-policy distillation, and empirically establish its plug-and-play improvements across four LLM backbones and two KE editors under various scenarios.
Community
Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing editors inject such a passage yet fail to use it: the edited model can recall the passage, but can neither answer atomic questions about its facts nor compose them into multi-hop reasoning. We attribute this missing property, which we term composability, to editors' passive reliance on the fixed passage as the sole learning source. In response, we cast editing as a proactive self-distillation from a privileged in-context state of the same model, which requires no external supervision. We further reveal that due to the novelty of the injected knowledge, the pre-edited model's own rollouts rarely cover it, which limits the effectiveness of pure on-policy distillation. To close this gap, we propose HPSE, which builds a hybrid rollout that steps in to place missing facts onto the student's own trajectory precisely where its coverage fails, while staying on-policy elsewhere. We theoretically analyze HPSE's advantage over pure on-policy distillation, and empirically establish its plug-and-play improvements across four LLM backbones and two KE editors under various scenarios.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.11660 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.11660 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.11660 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.