= 0.8\n\nis a hypothesis you can run on the corpus - statistically test it, rank it against every other candidate, and reuse the exact same function downstream. Stage 2 repairs them: an evaluator LLM checks each predicate against its actual execution trace and broken ones are rewritten. Survivors run over every training example, and only features passing Fisher's exact test under BH control - replicated on a held-out split get through. But a correlation in the data isn't a shortcut the model uses. Correlation doesn't mean causation right? So we intervene. Generate a minimal edit that removes the surface pattern while preserving the label, then measure the paired shift in predicted probability as causal signal. On MNLI: BERT causally exploits 9 of 10 validated features. RoBERTa only 6. And the strongest correlation in the entire dataset - always/every in premise + never/no in hypothesis, OR 10.01 is used by neither. A high odds ratio is a hypothesis, not a finding. And because the predicates are executable, they hand you group labels for free: 71.84% worst-group accuracy on CivilComments-WILDS, matching hand-labeled DFR, with no additional demographic annotation. For NLI tasks, HANS accuracy improves by up to 12.58 pp. We tried the discovery half at RewardBench2. No debiasing, just which surface predicates separate chosen from rejected and Math's strongest signal is the phrase \"let me help you solve this step by step.\" OR 22.7. This forces to learn string, not a reasoning style.","html":"<p>Language Models are very good at exploiting shortcuts. Existing approaches either require manual specification of the spurious feature or automate discovery only partially. The gap between dataset-level correlation and model-level exploitation is unaddressed. UNMASK closes that. The pipeline generates candidate shortcuts as executable boolean predicates.</p>\n<p>has_neg(h) AND overlap(p,h) >= 0.8</p>\n<p>is a hypothesis you can run on the corpus - statistically test it, rank it against every other candidate, and reuse the exact same function downstream. Stage 2 repairs them: an evaluator LLM checks each predicate against its actual execution trace and broken ones are rewritten. Survivors run over every training example, and only features passing Fisher's exact test under BH control - replicated on a held-out split get through. But a correlation in the data isn't a shortcut the model uses. Correlation doesn't mean causation right? So we intervene. Generate a minimal edit that removes the surface pattern while preserving the label, then measure the paired shift in predicted probability as causal signal. On MNLI: BERT causally exploits 9 of 10 validated features. RoBERTa only 6. And the strongest correlation in the entire dataset - always/every in premise + never/no in hypothesis, OR 10.01 is used by neither. A high odds ratio is a hypothesis, not a finding. And because the predicates are executable, they hand you group labels for free: 71.84% worst-group accuracy on CivilComments-WILDS, matching hand-labeled DFR, with no additional demographic annotation. For NLI tasks, HANS accuracy improves by up to 12.58 pp. We tried the discovery half at RewardBench2. No debiasing, just which surface predicates separate chosen from rejected and Math's strongest signal is the phrase \"let me help you solve this step by step.\" OR 22.7. This forces to learn string, not a reasoning style.</p>\n","updatedAt":"2026-08-17T07:54:34.297Z","author":{"_id":"63a9a5ff769a10efc4fc7c06","avatarUrl":"/avatars/6330221c8ae50a1be2ffc882485df603.svg","fullname":"Chidaksh Ravuru","name":"Chidaksh","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8811398148536682},"editors":["Chidaksh"],"editorAvatarUrls":["/avatars/6330221c8ae50a1be2ffc882485df603.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.09209","authors":[{"_id":"6a80e82fb601d59c652811ba","user":{"_id":"63a9a5ff769a10efc4fc7c06","avatarUrl":"/avatars/6330221c8ae50a1be2ffc882485df603.svg","isPro":false,"fullname":"Chidaksh Ravuru","user":"Chidaksh","type":"user","name":"Chidaksh"},"name":"Chidaksh Ravuru","status":"claimed_verified","statusLastChangedAt":"2026-08-16T00:45:04.720Z","hidden":false},{"_id":"6a80e82fb601d59c652811bb","name":"Shashank Srivastava","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/63a9a5ff769a10efc4fc7c06/aul6IPjeWPPL80ksua6ZT.png"],"publishedAt":"2026-08-10T00:00:00.000Z","submittedOnDailyAt":"2026-08-17T00:00:00.000Z","title":"UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers","submittedOnDailyBy":{"_id":"63a9a5ff769a10efc4fc7c06","avatarUrl":"/avatars/6330221c8ae50a1be2ffc882485df603.svg","isPro":false,"fullname":"Chidaksh Ravuru","user":"Chidaksh","type":"user","name":"Chidaksh"},"summary":"Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic or causal relevance, boosting benchmark performance while failing on adversarial or out-of-distribution inputs. Existing approaches either require manual specification of the feature vocabulary or automate discovery only partially, leaving the gap between dataset-level correlation and model-level exploitation unaddressed. We present U N M ASK, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation. Given unlabeled training examples, U N M ASK generates candidate surface patterns as executable boolean expressions, filters them through a statistical validation protocol with independent replication, and establishes causal model dependence via verified counterfactual interventions. Causally confirmed features then serve as annotation-free group definitions for Deep Feature Reweighting, eliminating the group labels that standard DFR requires. Applied to BERT and RoBERTa trained on MNLI, our pipeline independently rediscovers established lexical-overlap and negation biases, verifying 9 of 10 features on BERT and 6 on RoBERTa, and improving HANS accuracy by up to 12.58 pp. On CivilComments-WILDS, programmatic groups match the 70.1% worst- group accuracy of hand-labeled DFR (Kirichenko et al., 2023) without demographic annotation. We further demonstrate that the discovery and validation stages generalize to reward model preference data, surfacing interpretable spurious correlations in RewardBench2.","upvotes":1,"discussionId":"6a80e830b601d59c652811bc","ai_summary":"UNMASK automatically discovers and mitigates spurious correlations in text classifiers via causal verification and group-based reweighting without manual annotations.","ai_keywords":["spurious correlations","text classifiers","boolean expressions","counterfactual interventions","Deep Feature Reweighting","BERT","RoBERTa","MNLI","HANS","CivilComments-WILDS","worst-group accuracy","reward model","RewardBench2"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"669f9d1fec8789263c0e355a","name":"UNC-ChapelHill","fullname":"University of North Carolina at Chapel Hill","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/669f9c85bd649dba3b88e581/H5uB8_MCewnMtxEUnAvTL.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63a9a5ff769a10efc4fc7c06","avatarUrl":"/avatars/6330221c8ae50a1be2ffc882485df603.svg","isPro":false,"fullname":"Chidaksh Ravuru","user":"Chidaksh","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"669f9d1fec8789263c0e355a","name":"UNC-ChapelHill","fullname":"University of North Carolina at Chapel Hill","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/669f9c85bd649dba3b88e581/H5uB8_MCewnMtxEUnAvTL.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.09209.md","query":{}}">
UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers
Abstract
UNMASK automatically discovers and mitigates spurious correlations in text classifiers via causal verification and group-based reweighting without manual annotations.
Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic or causal relevance, boosting benchmark performance while failing on adversarial or out-of-distribution inputs. Existing approaches either require manual specification of the feature vocabulary or automate discovery only partially, leaving the gap between dataset-level correlation and model-level exploitation unaddressed. We present U N M ASK, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation. Given unlabeled training examples, U N M ASK generates candidate surface patterns as executable boolean expressions, filters them through a statistical validation protocol with independent replication, and establishes causal model dependence via verified counterfactual interventions. Causally confirmed features then serve as annotation-free group definitions for Deep Feature Reweighting, eliminating the group labels that standard DFR requires. Applied to BERT and RoBERTa trained on MNLI, our pipeline independently rediscovers established lexical-overlap and negation biases, verifying 9 of 10 features on BERT and 6 on RoBERTa, and improving HANS accuracy by up to 12.58 pp. On CivilComments-WILDS, programmatic groups match the 70.1% worst- group accuracy of hand-labeled DFR (Kirichenko et al., 2023) without demographic annotation. We further demonstrate that the discovery and validation stages generalize to reward model preference data, surfacing interpretable spurious correlations in RewardBench2.
Community
This comment has been hidden (marked as Resolved) Language Models are very good at exploiting shortcuts. Existing approaches either require manual specification of the spurious feature or automate discovery only partially. The gap between dataset-level correlation and model-level exploitation is unaddressed. UNMASK closes that. The pipeline generates candidate shortcuts as executable boolean predicates.
has_neg(h) AND overlap(p,h) >= 0.8
is a hypothesis you can run on the corpus - statistically test it, rank it against every other candidate, and reuse the exact same function downstream. Stage 2 repairs them: an evaluator LLM checks each predicate against its actual execution trace and broken ones are rewritten. Survivors run over every training example, and only features passing Fisher's exact test under BH control - replicated on a held-out split get through. But a correlation in the data isn't a shortcut the model uses. Correlation doesn't mean causation right? So we intervene. Generate a minimal edit that removes the surface pattern while preserving the label, then measure the paired shift in predicted probability as causal signal. On MNLI: BERT causally exploits 9 of 10 validated features. RoBERTa only 6. And the strongest correlation in the entire dataset - always/every in premise + never/no in hypothesis, OR 10.01 is used by neither. A high odds ratio is a hypothesis, not a finding. And because the predicates are executable, they hand you group labels for free: 71.84% worst-group accuracy on CivilComments-WILDS, matching hand-labeled DFR, with no additional demographic annotation. For NLI tasks, HANS accuracy improves by up to 12.58 pp. We tried the discovery half at RewardBench2. No debiasing, just which surface predicates separate chosen from rejected and Math's strongest signal is the phrase "let me help you solve this step by step." OR 22.7. This forces to learn string, not a reasoning style.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.09209 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.09209 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.09209 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.