Hugging Face Daily Papers · · 4 min read

One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Accepted to EMNLP 2026 Findings. A single polluted web page is enough to make production LLMs recommend a brand that does not exist — 27% at rank 1, 73.8% when the top three pages are swapped. Turning on reasoning makes it worse.</p>\n","updatedAt":"2026-08-25T06:12:11.727Z","author":{"_id":"69f0d85eedc6cb78ca435342","avatarUrl":"/avatars/f8a2404943ddda3470ae87c4b75a42ac.svg","fullname":"Minghao Luo","name":"leoluo25933","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9487366080284119},"editors":["leoluo25933"],"editorAvatarUrls":["/avatars/f8a2404943ddda3470ae87c4b75a42ac.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2606.13610","authors":[{"_id":"6a433ed2763f63ca3757ea24","user":{"_id":"69f0d85eedc6cb78ca435342","avatarUrl":"/avatars/f8a2404943ddda3470ae87c4b75a42ac.svg","isPro":false,"fullname":"Minghao Luo","user":"leoluo25933","type":"user","name":"leoluo25933"},"name":"Minghao Luo","status":"claimed_verified","statusLastChangedAt":"2026-07-01T08:46:43.798Z","hidden":false},{"_id":"6a433ed2763f63ca3757ea25","name":"Liang Chen","hidden":false}],"publishedAt":"2026-08-24T00:00:00.000Z","submittedOnDailyAt":"2026-08-25T00:00:00.000Z","title":"One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders","submittedOnDailyBy":{"_id":"69f0d85eedc6cb78ca435342","avatarUrl":"/avatars/f8a2404943ddda3470ae87c4b75a42ac.svg","isPro":false,"fullname":"Minghao Luo","user":"leoluo25933","type":"user","name":"leoluo25933"},"summary":"Search-augmented LLMs increasingly mediate everyday consumer recommendations by retrieving live web content. This creates a new risk: LLM recommenders may consume web content that Generative Engine Optimization (GEO) operators have polluted to mislead them. We ask: to what extent do they become unwitting promoters of fake products? We introduce FORGE (Fake Online Recommendations in Generative Environments), which locally rewrites real products in a frozen set of retrieved web pages into fake ones and measures how often the LLM recommends the fake product, across 225 real products in 15 categories and 5 consumer scenarios. Across 12 commercial and open-weights LLMs, all models are vulnerable: a single polluted page yields fooled rates of up to 27%, while the full top-3 replacement raises this to 73.8%. Vulnerability varies across categories, increasing when models lack stable prior knowledge of the products. Reasoning does not mitigate this vulnerability; instead, it often generates spurious social proof to justify false recommendations. None of the four defenses is adequate: the skepticism prompt can exacerbate vulnerability much like reasoning, the two consensus filters risk suppressing legitimate products, and credibility re-ranking helps every model but removes only a sixth of the fakes. We release the FORGE benchmark and the evaluation code at https://github.com/leoluolol/forge-benchmark.","upvotes":2,"discussionId":"6a433ed2763f63ca3757ea26","ai_summary":"Search-augmented LLM recommenders are highly vulnerable to web content polluted by generative engine optimization, frequently promoting fake products despite reasoning and defenses.","ai_keywords":["search-augmented LLMs","Generative Engine Optimization","FORGE benchmark","fake product recommendations","reasoning","consensus filters","credibility re-ranking"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69f0d85eedc6cb78ca435342","avatarUrl":"/avatars/f8a2404943ddda3470ae87c4b75a42ac.svg","isPro":false,"fullname":"Minghao Luo","user":"leoluo25933","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2606/2606.13610.md","query":{}}">
Papers
arxiv:2606.13610

One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders

Published on Aug 24
· Submitted by
Minghao Luo
on Aug 25
Authors:

Abstract

Search-augmented LLM recommenders are highly vulnerable to web content polluted by generative engine optimization, frequently promoting fake products despite reasoning and defenses.

Search-augmented LLMs increasingly mediate everyday consumer recommendations by retrieving live web content. This creates a new risk: LLM recommenders may consume web content that Generative Engine Optimization (GEO) operators have polluted to mislead them. We ask: to what extent do they become unwitting promoters of fake products? We introduce FORGE (Fake Online Recommendations in Generative Environments), which locally rewrites real products in a frozen set of retrieved web pages into fake ones and measures how often the LLM recommends the fake product, across 225 real products in 15 categories and 5 consumer scenarios. Across 12 commercial and open-weights LLMs, all models are vulnerable: a single polluted page yields fooled rates of up to 27%, while the full top-3 replacement raises this to 73.8%. Vulnerability varies across categories, increasing when models lack stable prior knowledge of the products. Reasoning does not mitigate this vulnerability; instead, it often generates spurious social proof to justify false recommendations. None of the four defenses is adequate: the skepticism prompt can exacerbate vulnerability much like reasoning, the two consensus filters risk suppressing legitimate products, and credibility re-ranking helps every model but removes only a sixth of the fakes. We release the FORGE benchmark and the evaluation code at https://github.com/leoluolol/forge-benchmark.

Community

Paper author Paper submitter about 3 hours ago

Accepted to EMNLP 2026 Findings. A single polluted web page is enough to make production LLMs recommend a brand that does not exist — 27% at rank 1, 73.8% when the top three pages are swapped. Turning on reasoning makes it worse.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.13610
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2606.13610 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2606.13610 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2606.13610 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers