.\" To prevent arbitrary manipulation of agents, defenders try to train secure LLMs, which, however, still suffer from near 100% attack success rates (ASRs) against adaptive prompt injections. We note that this is because existing defensive finetuning recipes rely on sequence-level feedback signals (in DPO or GRPO). Treating an entire output equally prevents the model from learning precisely which output tokens are insecure. In this paper, we propose Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning. The LLM receives an injected sample and produces a rollout, whose tokens are scored by the initialization model given the corresponding clean input. With more fine-grained training signals, our defended Qwen3.6-27B achieves a 9.0% ASR against the SoTA PISmith adaptive prompt injections, compared to 94.0% for the prior SoTA, Meta-SecAlign. The obtained security generalizes to domains completely unseen in training: in agentic tool calling, SecOPD achieves a 4.7% ASR compared to 5.5% for Meta-SecAlign. Code and the model are available at https://github.com/pppyb/SecOPD and pybbb/Qwen3.6-27B-SecOPD.","html":"<p>Prompt injection is listed as the #1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inject a prompt into the data, saying, \"Ignore all prior instructions and perform <an attacker’s task>.\" To prevent arbitrary manipulation of agents, defenders try to train secure LLMs, which, however, still suffer from near 100% attack success rates (ASRs) against adaptive prompt injections. We note that this is because existing defensive finetuning recipes rely on sequence-level feedback signals (in DPO or GRPO). Treating an entire output equally prevents the model from learning precisely which output tokens are insecure. In this paper, we propose Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning. The LLM receives an injected sample and produces a rollout, whose tokens are scored by the initialization model given the corresponding clean input. With more fine-grained training signals, our defended Qwen3.6-27B achieves a 9.0% ASR against the SoTA PISmith adaptive prompt injections, compared to 94.0% for the prior SoTA, Meta-SecAlign. The obtained security generalizes to domains completely unseen in training: in agentic tool calling, SecOPD achieves a 4.7% ASR compared to 5.5% for Meta-SecAlign. Code and the model are available at <a href=\"https://github.com/pppyb/SecOPD\" rel=\"nofollow\">https://github.com/pppyb/SecOPD</a> and pybbb/Qwen3.6-27B-SecOPD.</p>\n","updatedAt":"2026-08-26T15:52:15.663Z","author":{"_id":"660f670d0ba017647fa60508","avatarUrl":"/avatars/085303c35170b50241d61501aed5f206.svg","fullname":"Yibo Peng","name":"pybbb","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9007771611213684},"editors":["pybbb"],"editorAvatarUrls":["/avatars/085303c35170b50241d61501aed5f206.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.21500","authors":[{"_id":"6a8cfc608dd056518b7f544a","user":{"_id":"660f670d0ba017647fa60508","avatarUrl":"/avatars/085303c35170b50241d61501aed5f206.svg","isPro":false,"fullname":"Yibo Peng","user":"pybbb","type":"user","name":"pybbb"},"name":"Yibo Peng","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:14:23.645Z","hidden":false},{"_id":"6a8cfc608dd056518b7f544b","name":"Long Lian","hidden":false},{"_id":"6a8cfc608dd056518b7f544c","name":"David Wagner","hidden":false},{"_id":"6a8cfc608dd056518b7f544d","name":"Sizhe Chen","hidden":false}],"publishedAt":"2026-08-21T00:00:00.000Z","submittedOnDailyAt":"2026-08-26T00:00:00.000Z","title":"SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation","submittedOnDailyBy":{"_id":"660f670d0ba017647fa60508","avatarUrl":"/avatars/085303c35170b50241d61501aed5f206.svg","isPro":false,"fullname":"Yibo Peng","user":"pybbb","type":"user","name":"pybbb"},"summary":"Prompt injection is listed as the \\#1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inject a prompt into the data, saying, \"Ignore all prior instructions and perform <an attacker's task>.\" To prevent arbitrary manipulation of agents, defenders try to train secure LLMs, which, however, still suffer from near 100% attack success rates (ASRs) against adaptive prompt injections. We note that this is because existing defensive finetuning recipes rely on sequence-level feedback signals (in DPO or GRPO). Treating an entire output equally prevents the model from learning precisely which output tokens are insecure. In this paper, we propose Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning. The LLM receives an injected sample and produces a rollout, whose tokens are scored by the initialization model given the corresponding clean input. With more fine-grained training signals, our defended Qwen3.6-27B achieves a 9.0% ASR against the SoTA PISmith adaptive prompt injections, compared to 94.0% for the prior SoTA, Meta-SecAlign. The obtained security generalizes to domains completely unseen in training: in agentic tool calling, SecOPD achieves a 4.7% ASR compared to 5.5% for Meta-SecAlign. Code and the model are available at https://github.com/pppyb/SecOPD and https://huggingface.co/pybbb/Qwen3.6-27B-SecOPD.","upvotes":36,"discussionId":"6a8cfc608dd056518b7f544e","projectPage":"https://pppyb.github.io/SecOPD/","githubRepo":"https://github.com/pppyb/SecOPD","githubRepoAddedBy":"user","ai_summary":"SecOPD improves defense against adaptive prompt injection by using token-level feedback during fine-tuning, sharply reducing attack success rates on language models.","ai_keywords":["prompt injection","secure LLMs","defensive fine-tuning","DPO","GRPO","token-level feedback","Secure On-Policy Distillation","SecOPD","adaptive prompt injections","ASR"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":3},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"660f670d0ba017647fa60508","avatarUrl":"/avatars/085303c35170b50241d61501aed5f206.svg","isPro":false,"fullname":"Yibo Peng","user":"pybbb","type":"user"},{"_id":"65e66eb107e593c50aa36708","avatarUrl":"/avatars/762a5145bd65aae0549520bee714b3a1.svg","isPro":false,"fullname":"Sam Huang","user":"SammySpace","type":"user"},{"_id":"63797c273f575acc2f6893c0","avatarUrl":"/avatars/32d7a6a8881c8c4d80a097b732ed24b6.svg","isPro":false,"fullname":"Long(Tony) Lian","user":"longlian","type":"user"},{"_id":"6841e21c1e7aca4238fc6895","avatarUrl":"/avatars/728460dac92c8e96eee2f15764c72b24.svg","isPro":false,"fullname":"Xiwen Min","user":"alexis-mmm","type":"user"},{"_id":"646f59bc753be77a8e94bd62","avatarUrl":"/avatars/b34fd36e01d431a69de62c497f989079.svg","isPro":false,"fullname":"Xiyuxing Zhang","user":"zx-explorer","type":"user"},{"_id":"64affbb951e993d33be0ddd1","avatarUrl":"/avatars/36eb12f5c74cfe69d16a954f96819308.svg","isPro":false,"fullname":"Haizhong","user":"haizhongzheng","type":"user"},{"_id":"650c4efa5085c0ce1ffdf9b4","avatarUrl":"/avatars/148c9f643cd23a59ca10a449c6fa6a75.svg","isPro":true,"fullname":"Yang Zhou","user":"YangZhoumill","type":"user"},{"_id":"65401959cdc151e46270f6fb","avatarUrl":"/avatars/b7e71a9c68cdabd432a7ae9e8e640c4f.svg","isPro":false,"fullname":"Shu Liu","user":"lynnliu030","type":"user"},{"_id":"69ca36e9908caa7b8531b1ab","avatarUrl":"/avatars/667ffed98f74706d3178674eca94362d.svg","isPro":false,"fullname":"XunZou","user":"XunZOU","type":"user"},{"_id":"6a8d0d8f148ec81cdcf45091","avatarUrl":"/avatars/ed665272aa28d14b91c7b129f4a13140.svg","isPro":false,"fullname":"Kaixin Ji","user":"kkkkk2017","type":"user"},{"_id":"67e25d84995c9444d01e59da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/IoydlrrKXpWKyt2ytPMT9.png","isPro":false,"fullname":"Anonymous00159","user":"Anonymous00159","type":"user"},{"_id":"67f37798d16d55e4bb77f8b4","avatarUrl":"/avatars/865043bd56a183ce9fe634ce872b9bbc.svg","isPro":false,"fullname":"Dono","user":"AnonDono","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.21500.md","query":{}}">
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
Abstract
SecOPD improves defense against adaptive prompt injection by using token-level feedback during fine-tuning, sharply reducing attack success rates on language models.
Prompt injection is listed as the \#1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inject a prompt into the data, saying, "Ignore all prior instructions and perform <an attacker's task>." To prevent arbitrary manipulation of agents, defenders try to train secure LLMs, which, however, still suffer from near 100% attack success rates (ASRs) against adaptive prompt injections. We note that this is because existing defensive finetuning recipes rely on sequence-level feedback signals (in DPO or GRPO). Treating an entire output equally prevents the model from learning precisely which output tokens are insecure. In this paper, we propose Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning. The LLM receives an injected sample and produces a rollout, whose tokens are scored by the initialization model given the corresponding clean input. With more fine-grained training signals, our defended Qwen3.6-27B achieves a 9.0% ASR against the SoTA PISmith adaptive prompt injections, compared to 94.0% for the prior SoTA, Meta-SecAlign. The obtained security generalizes to domains completely unseen in training: in agentic tool calling, SecOPD achieves a 4.7% ASR compared to 5.5% for Meta-SecAlign. Code and the model are available at https://github.com/pppyb/SecOPD and https://huggingface.co/pybbb/Qwen3.6-27B-SecOPD.
Community
Prompt injection is listed as the #1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inject a prompt into the data, saying, "Ignore all prior instructions and perform <an attacker’s task>." To prevent arbitrary manipulation of agents, defenders try to train secure LLMs, which, however, still suffer from near 100% attack success rates (ASRs) against adaptive prompt injections. We note that this is because existing defensive finetuning recipes rely on sequence-level feedback signals (in DPO or GRPO). Treating an entire output equally prevents the model from learning precisely which output tokens are insecure. In this paper, we propose Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning. The LLM receives an injected sample and produces a rollout, whose tokens are scored by the initialization model given the corresponding clean input. With more fine-grained training signals, our defended Qwen3.6-27B achieves a 9.0% ASR against the SoTA PISmith adaptive prompt injections, compared to 94.0% for the prior SoTA, Meta-SecAlign. The obtained security generalizes to domains completely unseen in training: in agentic tool calling, SecOPD achieves a 4.7% ASR compared to 5.5% for Meta-SecAlign. Code and the model are available at https://github.com/pppyb/SecOPD and pybbb/Qwen3.6-27B-SecOPD.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.21500 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.21500 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.21500 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.