Hugging Face Daily Papers · · 6 min read

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

.\" To prevent arbitrary manipulation of agents, defenders try to train secure LLMs, which, however, still suffer from near 100% attack success rates (ASRs) against adaptive prompt injections. We note that this is because existing defensive finetuning recipes rely on sequence-level feedback signals (in DPO or GRPO). Treating an entire output equally prevents the model from learning precisely which output tokens are insecure. In this paper, we propose Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning. The LLM receives an injected sample and produces a rollout, whose tokens are scored by the initialization model given the corresponding clean input. With more fine-grained training signals, our defended Qwen3.6-27B achieves a 9.0% ASR against the SoTA PISmith adaptive prompt injections, compared to 94.0% for the prior SoTA, Meta-SecAlign. The obtained security generalizes to domains completely unseen in training: in agentic tool calling, SecOPD achieves a 4.7% ASR compared to 5.5% for Meta-SecAlign. Code and the model are available at https://github.com/pppyb/SecOPD and pybbb/Qwen3.6-27B-SecOPD.","html":"<p>Prompt injection is listed as the #1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inject a prompt into the data, saying, \"Ignore all prior instructions and perform &lt;an attacker’s task&gt;.\" To prevent arbitrary manipulation of agents, defenders try to train secure LLMs, which, however, still suffer from near 100% attack success rates (ASRs) against adaptive prompt injections. We note that this is because existing defensive finetuning recipes rely on sequence-level feedback signals (in DPO or GRPO). Treating an entire output equally prevents the model from learning precisely which output tokens are insecure. In this paper, we propose Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning. The LLM receives an injected sample and produces a rollout, whose tokens are scored by the initialization model given the corresponding clean input. With more fine-grained training signals, our defended Qwen3.6-27B achieves a 9.0% ASR against the SoTA PISmith adaptive prompt injections, compared to 94.0% for the prior SoTA, Meta-SecAlign. The obtained security generalizes to domains completely unseen in training: in agentic tool calling, SecOPD achieves a 4.7% ASR compared to 5.5% for Meta-SecAlign. Code and the model are available at <a href=\"https://github.com/pppyb/SecOPD\" rel=\"nofollow\">https://github.com/pppyb/SecOPD</a> and pybbb/Qwen3.6-27B-SecOPD.</p>\n","updatedAt":"2026-08-26T15:52:15.663Z","author":{"_id":"660f670d0ba017647fa60508","avatarUrl":"/avatars/085303c35170b50241d61501aed5f206.svg","fullname":"Yibo Peng","name":"pybbb","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9007771611213684},"editors":["pybbb"],"editorAvatarUrls":["/avatars/085303c35170b50241d61501aed5f206.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.21500","authors":[{"_id":"6a8cfc608dd056518b7f544a","user":{"_id":"660f670d0ba017647fa60508","avatarUrl":"/avatars/085303c35170b50241d61501aed5f206.svg","isPro":false,"fullname":"Yibo Peng","user":"pybbb","type":"user","name":"pybbb"},"name":"Yibo Peng","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:14:23.645Z","hidden":false},{"_id":"6a8cfc608dd056518b7f544b","name":"Long Lian","hidden":false},{"_id":"6a8cfc608dd056518b7f544c","name":"David Wagner","hidden":false},{"_id":"6a8cfc608dd056518b7f544d","name":"Sizhe Chen","hidden":false}],"publishedAt":"2026-08-21T00:00:00.000Z","submittedOnDailyAt":"2026-08-26T00:00:00.000Z","title":"SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation","submittedOnDailyBy":{"_id":"660f670d0ba017647fa60508","avatarUrl":"/avatars/085303c35170b50241d61501aed5f206.svg","isPro":false,"fullname":"Yibo Peng","user":"pybbb","type":"user","name":"pybbb"},"summary":"Prompt injection is listed as the \\#1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inject a prompt into the data, saying, \"Ignore all prior instructions and perform <an attacker's task>.\" To prevent arbitrary manipulation of agents, defenders try to train secure LLMs, which, however, still suffer from near 100% attack success rates (ASRs) against adaptive prompt injections. We note that this is because existing defensive finetuning recipes rely on sequence-level feedback signals (in DPO or GRPO). Treating an entire output equally prevents the model from learning precisely which output tokens are insecure. In this paper, we propose Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning. The LLM receives an injected sample and produces a rollout, whose tokens are scored by the initialization model given the corresponding clean input. With more fine-grained training signals, our defended Qwen3.6-27B achieves a 9.0% ASR against the SoTA PISmith adaptive prompt injections, compared to 94.0% for the prior SoTA, Meta-SecAlign. The obtained security generalizes to domains completely unseen in training: in agentic tool calling, SecOPD achieves a 4.7% ASR compared to 5.5% for Meta-SecAlign. Code and the model are available at https://github.com/pppyb/SecOPD and https://huggingface.co/pybbb/Qwen3.6-27B-SecOPD.","upvotes":36,"discussionId":"6a8cfc608dd056518b7f544e","projectPage":"https://pppyb.github.io/SecOPD/","githubRepo":"https://github.com/pppyb/SecOPD","githubRepoAddedBy":"user","ai_summary":"SecOPD improves defense against adaptive prompt injection by using token-level feedback during fine-tuning, sharply reducing attack success rates on language models.","ai_keywords":["prompt injection","secure LLMs","defensive fine-tuning","DPO","GRPO","token-level feedback","Secure On-Policy Distillation","SecOPD","adaptive prompt injections","ASR"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":3},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"660f670d0ba017647fa60508","avatarUrl":"/avatars/085303c35170b50241d61501aed5f206.svg","isPro":false,"fullname":"Yibo Peng","user":"pybbb","type":"user"},{"_id":"65e66eb107e593c50aa36708","avatarUrl":"/avatars/762a5145bd65aae0549520bee714b3a1.svg","isPro":false,"fullname":"Sam Huang","user":"SammySpace","type":"user"},{"_id":"63797c273f575acc2f6893c0","avatarUrl":"/avatars/32d7a6a8881c8c4d80a097b732ed24b6.svg","isPro":false,"fullname":"Long(Tony) Lian","user":"longlian","type":"user"},{"_id":"6841e21c1e7aca4238fc6895","avatarUrl":"/avatars/728460dac92c8e96eee2f15764c72b24.svg","isPro":false,"fullname":"Xiwen Min","user":"alexis-mmm","type":"user"},{"_id":"646f59bc753be77a8e94bd62","avatarUrl":"/avatars/b34fd36e01d431a69de62c497f989079.svg","isPro":false,"fullname":"Xiyuxing Zhang","user":"zx-explorer","type":"user"},{"_id":"64affbb951e993d33be0ddd1","avatarUrl":"/avatars/36eb12f5c74cfe69d16a954f96819308.svg","isPro":false,"fullname":"Haizhong","user":"haizhongzheng","type":"user"},{"_id":"650c4efa5085c0ce1ffdf9b4","avatarUrl":"/avatars/148c9f643cd23a59ca10a449c6fa6a75.svg","isPro":true,"fullname":"Yang Zhou","user":"YangZhoumill","type":"user"},{"_id":"65401959cdc151e46270f6fb","avatarUrl":"/avatars/b7e71a9c68cdabd432a7ae9e8e640c4f.svg","isPro":false,"fullname":"Shu Liu","user":"lynnliu030","type":"user"},{"_id":"69ca36e9908caa7b8531b1ab","avatarUrl":"/avatars/667ffed98f74706d3178674eca94362d.svg","isPro":false,"fullname":"XunZou","user":"XunZOU","type":"user"},{"_id":"6a8d0d8f148ec81cdcf45091","avatarUrl":"/avatars/ed665272aa28d14b91c7b129f4a13140.svg","isPro":false,"fullname":"Kaixin Ji","user":"kkkkk2017","type":"user"},{"_id":"67e25d84995c9444d01e59da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/IoydlrrKXpWKyt2ytPMT9.png","isPro":false,"fullname":"Anonymous00159","user":"Anonymous00159","type":"user"},{"_id":"67f37798d16d55e4bb77f8b4","avatarUrl":"/avatars/865043bd56a183ce9fe634ce872b9bbc.svg","isPro":false,"fullname":"Dono","user":"AnonDono","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.21500.md","query":{}}">
Papers
arxiv:2608.21500

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

Published on Aug 21
· Submitted by
Yibo Peng
on Aug 26
Authors:

Abstract

SecOPD improves defense against adaptive prompt injection by using token-level feedback during fine-tuning, sharply reducing attack success rates on language models.

Prompt injection is listed as the \#1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inject a prompt into the data, saying, "Ignore all prior instructions and perform <an attacker's task>." To prevent arbitrary manipulation of agents, defenders try to train secure LLMs, which, however, still suffer from near 100% attack success rates (ASRs) against adaptive prompt injections. We note that this is because existing defensive finetuning recipes rely on sequence-level feedback signals (in DPO or GRPO). Treating an entire output equally prevents the model from learning precisely which output tokens are insecure. In this paper, we propose Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning. The LLM receives an injected sample and produces a rollout, whose tokens are scored by the initialization model given the corresponding clean input. With more fine-grained training signals, our defended Qwen3.6-27B achieves a 9.0% ASR against the SoTA PISmith adaptive prompt injections, compared to 94.0% for the prior SoTA, Meta-SecAlign. The obtained security generalizes to domains completely unseen in training: in agentic tool calling, SecOPD achieves a 4.7% ASR compared to 5.5% for Meta-SecAlign. Code and the model are available at https://github.com/pppyb/SecOPD and https://huggingface.co/pybbb/Qwen3.6-27B-SecOPD.

Community

Paper author Paper submitter about 10 hours ago

Prompt injection is listed as the #1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inject a prompt into the data, saying, "Ignore all prior instructions and perform <an attacker’s task>." To prevent arbitrary manipulation of agents, defenders try to train secure LLMs, which, however, still suffer from near 100% attack success rates (ASRs) against adaptive prompt injections. We note that this is because existing defensive finetuning recipes rely on sequence-level feedback signals (in DPO or GRPO). Treating an entire output equally prevents the model from learning precisely which output tokens are insecure. In this paper, we propose Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning. The LLM receives an injected sample and produces a rollout, whose tokens are scored by the initialization model given the corresponding clean input. With more fine-grained training signals, our defended Qwen3.6-27B achieves a 9.0% ASR against the SoTA PISmith adaptive prompt injections, compared to 94.0% for the prior SoTA, Meta-SecAlign. The obtained security generalizes to domains completely unseen in training: in agentic tool calling, SecOPD achieves a 4.7% ASR compared to 5.5% for Meta-SecAlign. Code and the model are available at https://github.com/pppyb/SecOPD and pybbb/Qwen3.6-27B-SecOPD.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.21500
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.21500 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.21500 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.21500 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers