Hugging Face Daily Papers · · 8 min read

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing editors inject such a passage yet fail to use it: the edited model can recall the passage, but can neither answer atomic questions about its facts nor compose them into multi-hop reasoning. We attribute this missing property, which we term composability, to editors' passive reliance on the fixed passage as the sole learning source. In response, we cast editing as a proactive self-distillation from a privileged in-context state of the same model, which requires no external supervision. We further reveal that due to the novelty of the injected knowledge, the pre-edited model's own rollouts rarely cover it, which limits the effectiveness of pure on-policy distillation. To close this gap, we propose HPSE, which builds a hybrid rollout that steps in to place missing facts onto the student's own trajectory precisely where its coverage fails, while staying on-policy elsewhere. We theoretically analyze HPSE's advantage over pure on-policy distillation, and empirically establish its plug-and-play improvements across four LLM backbones and two KE editors under various scenarios.</p>\n","updatedAt":"2026-08-14T15:01:18.116Z","author":{"_id":"64bf811d76a6e2efcceabc00","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64bf811d76a6e2efcceabc00/0p3zSIVqzoME25Zmfh7SD.png","fullname":"Tianci Liu","name":"lliutianc","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9275947213172913},"editors":["lliutianc"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64bf811d76a6e2efcceabc00/0p3zSIVqzoME25Zmfh7SD.png"],"reactions":[],"isReport":false}},{"id":"6a7fc1f472f87f0631aab30f","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-15T01:33:40.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [KARLA: Knowledge-base Augmented Retrieval for Language Models](https://huggingface.co/papers/2606.26807) (2026)\n* [PRISM Edit: One Vector for All Temporal Answers](https://huggingface.co/papers/2607.11327) (2026)\n* [On-Policy Self-Distillation without Any Supervision](https://huggingface.co/papers/2608.06296) (2026)\n* [Regime-Aware Peer Specialization for Robust RAG under Heterogeneous Knowledge Conflicts](https://huggingface.co/papers/2606.30518) (2026)\n* [Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing](https://huggingface.co/papers/2607.20433) (2026)\n* [LeAct: Learning to Reason from Expert Actions](https://huggingface.co/papers/2607.21856) (2026)\n* [Training-Free Token-Level Steering for LLM Personalized Co-Writing](https://huggingface.co/papers/2608.06069) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2606.26807\">KARLA: Knowledge-base Augmented Retrieval for Language Models</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.11327\">PRISM Edit: One Vector for All Temporal Answers</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.06296\">On-Policy Self-Distillation without Any Supervision</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.30518\">Regime-Aware Peer Specialization for Robust RAG under Heterogeneous Knowledge Conflicts</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.20433\">Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.21856\">LeAct: Learning to Reason from Expert Actions</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.06069\">Training-Free Token-Level Steering for LLM Personalized Co-Writing</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-15T01:33:40.619Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7320732474327087},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.11660","authors":[{"_id":"6a7f2db5f747ea94019af4e8","user":{"_id":"64bf811d76a6e2efcceabc00","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64bf811d76a6e2efcceabc00/0p3zSIVqzoME25Zmfh7SD.png","isPro":false,"fullname":"Tianci Liu","user":"lliutianc","type":"user","name":"lliutianc"},"name":"Tianci Liu","status":"claimed_verified","statusLastChangedAt":"2026-08-15T08:45:04.160Z","hidden":false},{"_id":"6a7f2db5f747ea94019af4e9","name":"Zihan Dong","hidden":false},{"_id":"6a7f2db5f747ea94019af4ea","name":"Tianchun Li","hidden":false},{"_id":"6a7f2db5f747ea94019af4eb","name":"Yi-Chung Chen","hidden":false},{"_id":"6a7f2db5f747ea94019af4ec","name":"Qiming Cao","hidden":false},{"_id":"6a7f2db5f747ea94019af4ed","name":"Xingchen Wang","hidden":false},{"_id":"6a7f2db5f747ea94019af4ee","name":"Shiyang Wang","hidden":false},{"_id":"6a7f2db5f747ea94019af4ef","name":"Zichen Miao","hidden":false},{"_id":"6a7f2db5f747ea94019af4f0","name":"Linjun Zhang","hidden":false},{"_id":"6a7f2db5f747ea94019af4f1","name":"Haoyu Wang","hidden":false},{"_id":"6a7f2db5f747ea94019af4f2","name":"Jing Gao","hidden":false}],"publishedAt":"2026-08-12T00:00:00.000Z","submittedOnDailyAt":"2026-08-14T00:00:00.000Z","title":"Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing","submittedOnDailyBy":{"_id":"64bf811d76a6e2efcceabc00","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64bf811d76a6e2efcceabc00/0p3zSIVqzoME25Zmfh7SD.png","isPro":false,"fullname":"Tianci Liu","user":"lliutianc","type":"user","name":"lliutianc"},"summary":"Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing editors inject such a passage yet fail to use it: the edited model can recall the passage, but can neither answer atomic questions about its facts nor compose them into multi-hop reasoning. We attribute this missing property, which we term composability, to editors' passive reliance on the fixed passage as the sole learning source. In response, we cast editing as a proactive self-distillation from a privileged in-context state of the same model, which requires no external supervision. We further reveal that due to the novelty of the injected knowledge, the pre-edited model's own rollouts rarely cover it, which limits the effectiveness of pure on-policy distillation. To close this gap, we propose HPSE, which builds a hybrid rollout that steps in to place missing facts onto the student's own trajectory precisely where its coverage fails, while staying on-policy elsewhere. We theoretically analyze HPSE's advantage over pure on-policy distillation, and empirically establish its plug-and-play improvements across four LLM backbones and two KE editors under various scenarios.","upvotes":4,"discussionId":"6a7f2db5f747ea94019af4f3","ai_summary":"HPSE improves unstructured knowledge editing by distilling from hybrid rollouts that insert missing facts into the model's reasoning paths, enabling composable multi-hop reasoning.","ai_keywords":["large language models","knowledge editing","unstructured knowledge editing","composability","self-distillation","in-context learning","hybrid rollout","on-policy distillation","multi-hop reasoning"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64bf811d76a6e2efcceabc00","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64bf811d76a6e2efcceabc00/0p3zSIVqzoME25Zmfh7SD.png","isPro":false,"fullname":"Tianci Liu","user":"lliutianc","type":"user"},{"_id":"661ab1f1fa3b144a381fa454","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661ab1f1fa3b144a381fa454/IlpZBb9NCjo7ntFwMIH53.png","isPro":false,"fullname":"Urro","user":"urroxyz","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"631e14ac473a6825f285e89d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/631e14ac473a6825f285e89d/K-6QnoeGLg8XFvbTMMdqA.jpeg","isPro":false,"fullname":"Yury Panikov","user":"panikov","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.11660.md","query":{}}">
Papers
arxiv:2608.11660

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

Published on Aug 12
· Submitted by
Tianci Liu
on Aug 14
Authors:

Abstract

HPSE improves unstructured knowledge editing by distilling from hybrid rollouts that insert missing facts into the model's reasoning paths, enabling composable multi-hop reasoning.

Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing editors inject such a passage yet fail to use it: the edited model can recall the passage, but can neither answer atomic questions about its facts nor compose them into multi-hop reasoning. We attribute this missing property, which we term composability, to editors' passive reliance on the fixed passage as the sole learning source. In response, we cast editing as a proactive self-distillation from a privileged in-context state of the same model, which requires no external supervision. We further reveal that due to the novelty of the injected knowledge, the pre-edited model's own rollouts rarely cover it, which limits the effectiveness of pure on-policy distillation. To close this gap, we propose HPSE, which builds a hybrid rollout that steps in to place missing facts onto the student's own trajectory precisely where its coverage fails, while staying on-policy elsewhere. We theoretically analyze HPSE's advantage over pure on-policy distillation, and empirically establish its plug-and-play improvements across four LLM backbones and two KE editors under various scenarios.

Community

Paper author Paper submitter 1 day ago

Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing editors inject such a passage yet fail to use it: the edited model can recall the passage, but can neither answer atomic questions about its facts nor compose them into multi-hop reasoning. We attribute this missing property, which we term composability, to editors' passive reliance on the fixed passage as the sole learning source. In response, we cast editing as a proactive self-distillation from a privileged in-context state of the same model, which requires no external supervision. We further reveal that due to the novelty of the injected knowledge, the pre-edited model's own rollouts rarely cover it, which limits the effectiveness of pure on-policy distillation. To close this gap, we propose HPSE, which builds a hybrid rollout that steps in to place missing facts onto the student's own trajectory precisely where its coverage fails, while staying on-policy elsewhere. We theoretically analyze HPSE's advantage over pure on-policy distillation, and empirically establish its plug-and-play improvements across four LLM backbones and two KE editors under various scenarios.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.11660
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.11660 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.11660 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.11660 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers