Hugging Face Daily Papers · · 6 min read

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate under white-box assumptions, relying on text encoder optimization, weight editing, or inference-time intervention, and fundamentally cannot scale to proprietary models. Black-box alternatives based on LLM prompt rewriting offer broader applicability, yet fail in a critical regime we identify as the \\textit{benign adversarial} problem: prompts that are linguistically safe but still trigger harmful generation due to the model's learned data distribution. We propose DiSCO, a zero-shot, strictly black-box defense that operates entirely at the prompt level as a plug-and-play module, requiring no model retraining, fine-tuning, or access to model internals. DiSCO performs distribution-guided suffix expansion via beam search, optimized through contrastive scoring over safe and unsafe image pools generated by the target model itself, with iterative adaptive feedback until safe content is produced. We demonstrate that DiSCO consistently enhances the safety of both undefended and defended models on the I2P benchmark under multiple red-teaming attacks, achieving 37.7% and 25.13% ASR reduction, respectively, while maintaining semantic fidelity and improving image coherence. As a black-box, architecture-agnostic module, DiSCO can be readily applied to any text-to-image system without necessitating any changes to the model itself.</p>\n","updatedAt":"2026-08-19T08:49:53.482Z","author":{"_id":"65c898197faf326d067e2c0d","avatarUrl":"/avatars/dde3c047e627bacd69e0f063a4a833b2.svg","fullname":"Tong Zhang","name":"Tong98Zhang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8974401950836182},"editors":["Tong98Zhang"],"editorAvatarUrls":["/avatars/dde3c047e627bacd69e0f063a4a833b2.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.17067","authors":[{"_id":"6a854ed1536bdd3bdd48f8ac","name":"Tong Zhang","hidden":false},{"_id":"6a854ed1536bdd3bdd48f8ad","name":"Motasem Alfarra","hidden":false},{"_id":"6a854ed1536bdd3bdd48f8ae","name":"Carlos Hinojosa","hidden":false},{"_id":"6a854ed1536bdd3bdd48f8af","name":"Christos Louizos","hidden":false},{"_id":"6a854ed1536bdd3bdd48f8b0","name":"Bernard Ghanem","hidden":false}],"publishedAt":"2026-08-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-19T00:00:00.000Z","title":"DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization","submittedOnDailyBy":{"_id":"65c898197faf326d067e2c0d","avatarUrl":"/avatars/dde3c047e627bacd69e0f063a4a833b2.svg","isPro":false,"fullname":"Tong Zhang","user":"Tong98Zhang","type":"user","name":"Tong98Zhang"},"summary":"As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate under white-box assumptions, relying on text encoder optimization, weight editing, or inference-time intervention, and fundamentally cannot scale to proprietary models. Black-box alternatives based on LLM prompt rewriting offer broader applicability, yet fail in a critical regime we identify as the benign adversarial problem: prompts that are linguistically safe but still trigger harmful generation due to the model's learned data distribution. We propose DiSCO, a zero-shot, strictly black-box defense that operates entirely at the prompt level as a plug-and-play module, requiring no model retraining, fine-tuning, or access to model internals. DiSCO performs distribution-guided suffix expansion via beam search, optimized through contrastive scoring over safe and unsafe image pools generated by the target model itself, with iterative adaptive feedback until safe content is produced. We demonstrate that DiSCO consistently enhances the safety of both undefended and defended models on the I2P benchmark under multiple red-teaming attacks, achieving 37.7% and 25.13% ASR reduction, respectively, while maintaining semantic fidelity and improving image coherence. As a black-box, architecture-agnostic module, DiSCO can be readily applied to any text-to-image system without necessitating any changes to the model itself.","upvotes":7,"discussionId":"6a854ed1536bdd3bdd48f8b1","ai_summary":"DiSCO is a black-box, zero-shot prompt-level defense that uses distribution-guided suffix expansion and contrastive scoring to reduce harmful image generation without altering the model.","ai_keywords":["text-to-image generative models","red-teaming adversarial attacks","white-box defenses","text encoder optimization","weight editing","inference-time intervention","black-box defense","LLM prompt rewriting","benign adversarial problem","DiSCO","zero-shot defense","distribution-guided suffix expansion","beam search","contrastive scoring","I2P benchmark","ASR reduction"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"667dda1c1fc985946ecf1d78","name":"kaust-generative-ai","fullname":"KAUST Center of Excellence in Generative AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/643651b41adb261e94e11ca4/sYnzHTtTrFK2h-nSrIL7A.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65c898197faf326d067e2c0d","avatarUrl":"/avatars/dde3c047e627bacd69e0f063a4a833b2.svg","isPro":false,"fullname":"Tong Zhang","user":"Tong98Zhang","type":"user"},{"_id":"65f6eff66396309f02a18dab","avatarUrl":"/avatars/3b315ecbe2815859ca7fb16d277740c8.svg","isPro":false,"fullname":"Harethah Abu Shairah","user":"HarethahMo","type":"user"},{"_id":"65803d64defc9c0d30db4f88","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65803d64defc9c0d30db4f88/gVxJaB8vSUFj4r0vnTVbP.jpeg","isPro":true,"fullname":"Shuming Liu","user":"sming256","type":"user"},{"_id":"665ee96b9f29909d03abfa37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/665ee96b9f29909d03abfa37/M1iPuO3LzRu8s0apAgjkf.jpeg","isPro":false,"fullname":"Mohamad Zbib","user":"zbeeb","type":"user"},{"_id":"6620e21d11561bf979229d9f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6620e21d11561bf979229d9f/qzMMJI4PJfEkJDYK4mWZM.jpeg","isPro":false,"fullname":"Carlos Hinojosa","user":"carlosh93","type":"user"},{"_id":"672b17efcb9a8b2f215a8330","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ROlQr69Tn-Q7aDG3c3gdB.png","isPro":false,"fullname":"Karen Sanchez","user":"ksanchez84","type":"user"},{"_id":"690a2a000aeb0271a8b63d86","avatarUrl":"/avatars/dc85b673618f82824137bd867e970fcb.svg","isPro":false,"fullname":"Motasem Alfarra","user":"malfarra","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"667dda1c1fc985946ecf1d78","name":"kaust-generative-ai","fullname":"KAUST Center of Excellence in Generative AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/643651b41adb261e94e11ca4/sYnzHTtTrFK2h-nSrIL7A.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.17067.md","query":{}}">
Papers
arxiv:2608.17067

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

Published on Aug 17
· Submitted by
Tong Zhang
on Aug 19
Authors:
,

Abstract

DiSCO is a black-box, zero-shot prompt-level defense that uses distribution-guided suffix expansion and contrastive scoring to reduce harmful image generation without altering the model.

As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate under white-box assumptions, relying on text encoder optimization, weight editing, or inference-time intervention, and fundamentally cannot scale to proprietary models. Black-box alternatives based on LLM prompt rewriting offer broader applicability, yet fail in a critical regime we identify as the benign adversarial problem: prompts that are linguistically safe but still trigger harmful generation due to the model's learned data distribution. We propose DiSCO, a zero-shot, strictly black-box defense that operates entirely at the prompt level as a plug-and-play module, requiring no model retraining, fine-tuning, or access to model internals. DiSCO performs distribution-guided suffix expansion via beam search, optimized through contrastive scoring over safe and unsafe image pools generated by the target model itself, with iterative adaptive feedback until safe content is produced. We demonstrate that DiSCO consistently enhances the safety of both undefended and defended models on the I2P benchmark under multiple red-teaming attacks, achieving 37.7% and 25.13% ASR reduction, respectively, while maintaining semantic fidelity and improving image coherence. As a black-box, architecture-agnostic module, DiSCO can be readily applied to any text-to-image system without necessitating any changes to the model itself.

Community

Paper submitter about 1 hour ago

As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate under white-box assumptions, relying on text encoder optimization, weight editing, or inference-time intervention, and fundamentally cannot scale to proprietary models. Black-box alternatives based on LLM prompt rewriting offer broader applicability, yet fail in a critical regime we identify as the \textit{benign adversarial} problem: prompts that are linguistically safe but still trigger harmful generation due to the model's learned data distribution. We propose DiSCO, a zero-shot, strictly black-box defense that operates entirely at the prompt level as a plug-and-play module, requiring no model retraining, fine-tuning, or access to model internals. DiSCO performs distribution-guided suffix expansion via beam search, optimized through contrastive scoring over safe and unsafe image pools generated by the target model itself, with iterative adaptive feedback until safe content is produced. We demonstrate that DiSCO consistently enhances the safety of both undefended and defended models on the I2P benchmark under multiple red-teaming attacks, achieving 37.7% and 25.13% ASR reduction, respectively, while maintaining semantic fidelity and improving image coherence. As a black-box, architecture-agnostic module, DiSCO can be readily applied to any text-to-image system without necessitating any changes to the model itself.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.17067
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.17067 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.17067 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.17067 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers