Hugging Face Daily Papers · · 4 min read

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Multi-agent medical AI fails socially, not visually: agents that resist shortcuts alone adopt wrong answers 38% of the time under peer agreement. Independent referees detect it at 77–88% precision.</p>\n","updatedAt":"2026-08-17T11:23:53.977Z","author":{"_id":"628ddf04986ae70e823298f7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/628ddf04986ae70e823298f7/P6GyCswDo3dDMd59DEkWC.png","fullname":"Sebastián Andres Cajas Ordóñez","name":"sebasmos","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8881661295890808},"editors":["sebasmos"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/628ddf04986ae70e823298f7/P6GyCswDo3dDMd59DEkWC.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.03744","authors":[{"_id":"6a82ef38b25f624fb96ac382","name":"Sebastián Andrés Cajas Ordóñez","hidden":false},{"_id":"6a82ef38b25f624fb96ac383","name":"Agastya Munnangi","hidden":false},{"_id":"6a82ef38b25f624fb96ac384","name":"Aldo Marzullo","hidden":false},{"_id":"6a82ef38b25f624fb96ac385","name":"Felipe Ocampo Osorio","hidden":false},{"_id":"6a82ef38b25f624fb96ac386","name":"Quang Bui","hidden":false},{"_id":"6a82ef38b25f624fb96ac387","name":"Mohammad Shahin","hidden":false},{"_id":"6a82ef38b25f624fb96ac388","name":"Armaan Grewal","hidden":false},{"_id":"6a82ef38b25f624fb96ac389","name":"Emmanuel Paul Kwesiga","hidden":false},{"_id":"6a82ef38b25f624fb96ac38a","name":"Anqi Peter Li","hidden":false},{"_id":"6a82ef38b25f624fb96ac38b","name":"Josephine Nanyonjo","hidden":false},{"_id":"6a82ef38b25f624fb96ac38c","name":"Aaditya Panchal","hidden":false},{"_id":"6a82ef38b25f624fb96ac38d","name":"Arshnoor Bhutani","hidden":false},{"_id":"6a82ef38b25f624fb96ac38e","name":"Nikhil Jaiswal","hidden":false},{"_id":"6a82ef38b25f624fb96ac38f","name":"Milit S. Patel","hidden":false},{"_id":"6a82ef38b25f624fb96ac390","name":"Maximin Lange","hidden":false},{"_id":"6a82ef38b25f624fb96ac391","name":"Leo Anthony Celi","hidden":false}],"publishedAt":"2026-08-04T14:37:16.000Z","submittedOnDailyAt":"2026-08-17T00:00:00.000Z","title":"Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems","submittedOnDailyBy":{"_id":"628ddf04986ae70e823298f7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/628ddf04986ae70e823298f7/P6GyCswDo3dDMd59DEkWC.png","isPro":false,"fullname":"Sebastián Andres Cajas Ordóñez","user":"sebasmos","type":"user","name":"sebasmos"},"summary":"Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore. Across seven cohorts on six public datasets spanning text (MedQA-USMLE, MedMCQA, MIMIC-CXR reports), imaging (NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert) and tabular ICU records (SUPPORT2), Gemini committees resist these cues in isolation (flip 5-16%), yet a socially plausible shortcut spreads: when two peers assert the same wrong answer, the holdout under test adopts it in 38% of cases, as does a false \"pre-screen\" system flag, on both capability tiers. Of three oversight agents, a gate cannot separate adoption from honest agreement (false-positive rate 100%); a same-lineage judge reading only the transcript flags adoption on text (precision 100%, recall 93%) but collapses onto the gate in imaging; a referee that privately re-queries the holdout transfers to imaging (77-88% precision, 13-21% false-positive rate). Tripling a cue's visual salience does not move contagion, whereas a second peer voice raises it by half again. Gaming a hidden rubric is near-silent: only 1/10 text and 1/134 imaging drifters name the rubric they moved toward. What games a committee is social plausibility, and only a referee independent of self-report catches it. Code: https://github.com/criticaldata/benchmaxxing","upvotes":1,"discussionId":"6a82ef39b25f624fb96ac392","githubRepo":"https://github.com/criticaldata/benchmaxxing","githubRepoAddedBy":"user","ai_summary":"Multi-agent clinical committees are vulnerable to socially plausible shortcuts rather than isolated cues, and only independent referee oversight reliably detects adoption.","ai_keywords":["language-model agents","clinical decision support","multi-agent deliberation","shortcut learning","social plausibility","oversight agents","referee","cross-modal evaluation"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"63728bde14d543d507ae970d","name":"MIT","fullname":"Massachusetts Institute of Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/S90qoeEJeEYaYf-c7Zs8g.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6270324ebecab9e2dcf245de","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6270324ebecab9e2dcf245de/cMbtWSasyNlYc9hvsEEzt.jpeg","isPro":false,"fullname":"Kye Gomez","user":"kye","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63728bde14d543d507ae970d","name":"MIT","fullname":"Massachusetts Institute of Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/S90qoeEJeEYaYf-c7Zs8g.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.03744.md","query":{}}">
Papers
arxiv:2608.03744

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems

Published on Aug 4
· Submitted by
Sebastián Andres Cajas Ordóñez
on Aug 17
Authors:
,

Abstract

Multi-agent clinical committees are vulnerable to socially plausible shortcuts rather than isolated cues, and only independent referee oversight reliably detects adoption.

Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore. Across seven cohorts on six public datasets spanning text (MedQA-USMLE, MedMCQA, MIMIC-CXR reports), imaging (NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert) and tabular ICU records (SUPPORT2), Gemini committees resist these cues in isolation (flip 5-16%), yet a socially plausible shortcut spreads: when two peers assert the same wrong answer, the holdout under test adopts it in 38% of cases, as does a false "pre-screen" system flag, on both capability tiers. Of three oversight agents, a gate cannot separate adoption from honest agreement (false-positive rate 100%); a same-lineage judge reading only the transcript flags adoption on text (precision 100%, recall 93%) but collapses onto the gate in imaging; a referee that privately re-queries the holdout transfers to imaging (77-88% precision, 13-21% false-positive rate). Tripling a cue's visual salience does not move contagion, whereas a second peer voice raises it by half again. Gaming a hidden rubric is near-silent: only 1/10 text and 1/134 imaging drifters name the rubric they moved toward. What games a committee is social plausibility, and only a referee independent of self-report catches it. Code: https://github.com/criticaldata/benchmaxxing

Community

Paper submitter about 4 hours ago

Multi-agent medical AI fails socially, not visually: agents that resist shortcuts alone adopt wrong answers 38% of the time under peer agreement. Independent referees detect it at 77–88% precision.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.03744
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.03744 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.03744 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.03744 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers