Hugging Face Daily Papers · · 5 min read

Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We assess indirect prompt injection in DeepSeek Harness (DSH), using AI-Infra-Guard (A.I.G) to construct tests, deliver controlled taint, execute DSH, collect traces, and judge outcomes. The study covers 14,560 controlled executions over 16 indirect-content channels, text and file carrier modes, 35 payload objectives, one unmodified baseline, and 12 attack methods. The experiment preserves DSH's agent loop, tool registry, model adapter, and session-event path; source tools and sensitive sinks are local fixtures, so attempted actions are recorded without external side effects. We evaluate each trace with a deterministic rule-based judge, \\JudgeR{} (RuleJudge), and a semantic LLM-based judge, \\JudgeL{} (LLMJudge). The strongest observed attack success rates are 17.0% under \\JudgeL{} for fake-completion attack in text mode, 25.5% under \\JudgeR{} for hidden Unicode in file mode, and 16.0% under \\JudgeR{} for the skills channel in file mode. \\JudgeL{} also assigns partial compliance more often than \\JudgeR{} (7.3% versus 2.0%). We relate these results to DSH's treatment of tool results, additional contexts, and tool-call policy hooks, then identify controls that should sit between untrusted content and sensitive actions. Our code is available at this https URL .</p>\n","updatedAt":"2026-08-19T02:37:16.358Z","author":{"_id":"650fc649dc509ae7d7d58215","avatarUrl":"/avatars/7bb8d9048ff604cf7475e52c33ec780e.svg","fullname":"Zonghao Ying","name":"Zonghao2025","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8815545439720154},"editors":["Zonghao2025"],"editorAvatarUrls":["/avatars/7bb8d9048ff604cf7475e52c33ec780e.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.16393","authors":[{"_id":"6a8516b0536bdd3bdd48f7de","name":"Zonghao Ying","hidden":false},{"_id":"6a8516b0536bdd3bdd48f7df","name":"Xiangfan Wu","hidden":false},{"_id":"6a8516b0536bdd3bdd48f7e0","name":"Huiyu Wu","hidden":false},{"_id":"6a8516b0536bdd3bdd48f7e1","name":"Xing Zheng","hidden":false},{"_id":"6a8516b0536bdd3bdd48f7e2","name":"Huangsheng Cheng","hidden":false},{"_id":"6a8516b0536bdd3bdd48f7e3","name":"Xiaorong Shi","hidden":false},{"_id":"6a8516b0536bdd3bdd48f7e4","name":"Jing Guo","hidden":false}],"publishedAt":"2026-08-18T00:00:00.000Z","submittedOnDailyAt":"2026-08-19T00:00:00.000Z","title":"Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection","submittedOnDailyBy":{"_id":"650fc649dc509ae7d7d58215","avatarUrl":"/avatars/7bb8d9048ff604cf7475e52c33ec780e.svg","isPro":false,"fullname":"Zonghao Ying","user":"Zonghao2025","type":"user","name":"Zonghao2025"},"summary":"We assess indirect prompt injection in DeepSeek Harness (DSH), using AI-Infra-Guard (A.I.G) to construct tests, deliver controlled taint, execute DSH, collect traces, and judge outcomes. The study covers 14,560 controlled executions over 16 indirect-content channels, text and file carrier modes, 35 payload objectives, one unmodified baseline, and 12 attack methods. The experiment preserves DSH's agent loop, tool registry, model adapter, and session-event path; source tools and sensitive sinks are local fixtures, so attempted actions are recorded without external side effects. We evaluate each trace with a deterministic rule-based judge, (RuleJudge), and a semantic LLM-based judge, (LLMJudge). The strongest observed attack success rates are 17.0% under for fake-completion attack in text mode, 25.5% under for hidden Unicode in file mode, and 16.0% under for the skills channel in file mode. also assigns partial compliance more often than (7.3% versus 2.0%). We relate these results to DSH's treatment of tool results, additional contexts, and tool-call policy hooks, then identify controls that should sit between untrusted content and sensitive actions. Our code is available at https://github.com/Tencent/AI-Infra-Guard/tree/main/Research/deepseek-harness-security-assessment .","upvotes":3,"discussionId":"6a8516b0536bdd3bdd48f7e5","ai_summary":"Researchers evaluate indirect prompt injection risks in DeepSeek Harness using controlled taint and dual judges, finding notable success rates across text and file channels and recommending controls between untrusted content and sensitive actions.","ai_keywords":["indirect prompt injection","DeepSeek Harness","AI-Infra-Guard","agent loop","tool registry","model adapter","session-event path","rule-based judge","LLM-based judge","fake-completion attack","hidden Unicode","skills channel","tool-call policy hooks"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"650fc649dc509ae7d7d58215","avatarUrl":"/avatars/7bb8d9048ff604cf7475e52c33ec780e.svg","isPro":false,"fullname":"Zonghao Ying","user":"Zonghao2025","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"66d8512c54209e9101811e8e","avatarUrl":"/avatars/62dfd8e6261108f2508efe678d5a2a57.svg","isPro":false,"fullname":"M Saad Salman","user":"MSS444","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.16393.md","query":{}}">
Papers
arxiv:2608.16393

Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection

Published on Aug 18
· Submitted by
Zonghao Ying
on Aug 19
Authors:
,

Abstract

Researchers evaluate indirect prompt injection risks in DeepSeek Harness using controlled taint and dual judges, finding notable success rates across text and file channels and recommending controls between untrusted content and sensitive actions.

We assess indirect prompt injection in DeepSeek Harness (DSH), using AI-Infra-Guard (A.I.G) to construct tests, deliver controlled taint, execute DSH, collect traces, and judge outcomes. The study covers 14,560 controlled executions over 16 indirect-content channels, text and file carrier modes, 35 payload objectives, one unmodified baseline, and 12 attack methods. The experiment preserves DSH's agent loop, tool registry, model adapter, and session-event path; source tools and sensitive sinks are local fixtures, so attempted actions are recorded without external side effects. We evaluate each trace with a deterministic rule-based judge, (RuleJudge), and a semantic LLM-based judge, (LLMJudge). The strongest observed attack success rates are 17.0% under for fake-completion attack in text mode, 25.5% under for hidden Unicode in file mode, and 16.0% under for the skills channel in file mode. also assigns partial compliance more often than (7.3% versus 2.0%). We relate these results to DSH's treatment of tool results, additional contexts, and tool-call policy hooks, then identify controls that should sit between untrusted content and sensitive actions. Our code is available at https://github.com/Tencent/AI-Infra-Guard/tree/main/Research/deepseek-harness-security-assessment .

Community

Paper submitter about 5 hours ago

We assess indirect prompt injection in DeepSeek Harness (DSH), using AI-Infra-Guard (A.I.G) to construct tests, deliver controlled taint, execute DSH, collect traces, and judge outcomes. The study covers 14,560 controlled executions over 16 indirect-content channels, text and file carrier modes, 35 payload objectives, one unmodified baseline, and 12 attack methods. The experiment preserves DSH's agent loop, tool registry, model adapter, and session-event path; source tools and sensitive sinks are local fixtures, so attempted actions are recorded without external side effects. We evaluate each trace with a deterministic rule-based judge, \JudgeR{} (RuleJudge), and a semantic LLM-based judge, \JudgeL{} (LLMJudge). The strongest observed attack success rates are 17.0% under \JudgeL{} for fake-completion attack in text mode, 25.5% under \JudgeR{} for hidden Unicode in file mode, and 16.0% under \JudgeR{} for the skills channel in file mode. \JudgeL{} also assigns partial compliance more often than \JudgeR{} (7.3% versus 2.0%). We relate these results to DSH's treatment of tool results, additional contexts, and tool-call policy hooks, then identify controls that should sit between untrusted content and sensitive actions. Our code is available at this https URL .

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.16393
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.16393 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.16393 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.16393 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers