repair episode already in context lowers false alarms by 2.8 to 11.5 pp against a length-matched non-audit control, in 15 of 15 model x wording combinations, with the present task held byte-identical.\n- **Discrimination does not follow.** The criterion moves in 15 of 15 and survives correction in 13. d' survives in 0 of 15, though its test is half as sensitive by construction.\n- **Polarity does not explain it.** An episode whose audit reported an error is more lenient still - the opposite sign to what the accumulated-message literature predicts.\n\nThe caution is about placement: be wary of putting a verifier in a context where it has already carried out the kind of repair it is about to judge.\n\nCode: https://github.com/parsa-mz/crtitxer","html":"<p>A verifier that flags fewer errors has not necessarily become more accurate. It may only have moved its threshold.</p>\n<p>We looked for the accuracy gain on three open-weight models, at the quantity the improvement account predicts - discrimination - and it is not there. What we find instead is leniency:</p>\n<ul>\n<li><strong>The episode moves the threshold.</strong> A completed audit --> repair episode already in context lowers false alarms by 2.8 to 11.5 pp against a length-matched non-audit control, in 15 of 15 model x wording combinations, with the present task held byte-identical.</li>\n<li><strong>Discrimination does not follow.</strong> The criterion moves in 15 of 15 and survives correction in 13. d' survives in 0 of 15, though its test is half as sensitive by construction.</li>\n<li><strong>Polarity does not explain it.</strong> An episode whose audit reported an error is more lenient still - the opposite sign to what the accumulated-message literature predicts.</li>\n</ul>\n<p>The caution is about placement: be wary of putting a verifier in a context where it has already carried out the kind of repair it is about to judge.</p>\n<p>Code: <a href=\"https://github.com/parsa-mz/crtitxer\" rel=\"nofollow\">https://github.com/parsa-mz/crtitxer</a></p>\n","updatedAt":"2026-08-18T04:17:17.886Z","author":{"_id":"63f9a22120589ee7cd649a33","avatarUrl":"/avatars/85e233f34930b5ff50a8cc1ec0a9d72f.svg","fullname":"Parsa Mazaheri","name":"parsa-mz","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9481157660484314},"editors":["parsa-mz"],"editorAvatarUrls":["/avatars/85e233f34930b5ff50a8cc1ec0a9d72f.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.16003","authors":[{"_id":"6a83db0e675db694db8cd5bd","name":"Parsa Mazaheri","hidden":false},{"_id":"6a83db0e675db694db8cd5be","name":"Kasra Mazaheri","hidden":false}],"publishedAt":"2026-08-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency","submittedOnDailyBy":{"_id":"63f9a22120589ee7cd649a33","avatarUrl":"/avatars/85e233f34930b5ff50a8cc1ec0a9d72f.svg","isPro":false,"fullname":"Parsa Mazaheri","user":"parsa-mz","type":"user","name":"parsa-mz"},"summary":"Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human-verified-correct ProcessBench traces with the present task held byte-identical, we find that a completed audit -> repair episode already in the model's context lowers false alarms in 15 of 15 model x wording combinations, by 2.8 to 11.5 percentage points against a length-matched non-audit control, a 9 to 25% reduction relative to that control. The direction contradicts what the accumulated-message literature predicts: an episode whose audit reported an error lowers false alarms further still, at all five wordings on the model where that manipulation lands cleanly, though a negativity asymmetry predicts more flagging. Decomposing the episode finds repair content and audit verdict complementary: different components carry the effect on different model families. Signal-detection analysis locates the change in the threshold rather than in discrimination -- the criterion moves in 15 of 15 combinations and survives correction in 13 while d' survives in none, though the d' test is half as sensitive by construction -- and a hand audit of 50 false alarms finds 82% simply wrong, so at this operating point the shift need not be harmful. With reasoning enabled the effect keeps its relative size on both models tested, and the threshold reading holds there too.","upvotes":2,"discussionId":"6a83db0e675db694db8cd5bf","githubRepo":"https://github.com/parsa-mz/crtitxer","githubRepoAddedBy":"user","ai_summary":"Prior audit and repair episodes in context reduce false alarms by shifting decision thresholds rather than discrimination, with repair content and audit verdict complementarily affecting different model families.","ai_keywords":["signal-detection analysis","d'","false alarms","audit","repair","reasoning","threshold","discrimination"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63f9a22120589ee7cd649a33","avatarUrl":"/avatars/85e233f34930b5ff50a8cc1ec0a9d72f.svg","isPro":false,"fullname":"Parsa Mazaheri","user":"parsa-mz","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.16003.md","query":{}}">
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
Abstract
Prior audit and repair episodes in context reduce false alarms by shifting decision thresholds rather than discrimination, with repair content and audit verdict complementarily affecting different model families.
Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human-verified-correct ProcessBench traces with the present task held byte-identical, we find that a completed audit -> repair episode already in the model's context lowers false alarms in 15 of 15 model x wording combinations, by 2.8 to 11.5 percentage points against a length-matched non-audit control, a 9 to 25% reduction relative to that control. The direction contradicts what the accumulated-message literature predicts: an episode whose audit reported an error lowers false alarms further still, at all five wordings on the model where that manipulation lands cleanly, though a negativity asymmetry predicts more flagging. Decomposing the episode finds repair content and audit verdict complementary: different components carry the effect on different model families. Signal-detection analysis locates the change in the threshold rather than in discrimination -- the criterion moves in 15 of 15 combinations and survives correction in 13 while d' survives in none, though the d' test is half as sensitive by construction -- and a hand audit of 50 false alarms finds 82% simply wrong, so at this operating point the shift need not be harmful. With reasoning enabled the effect keeps its relative size on both models tested, and the threshold reading holds there too.
Community
A verifier that flags fewer errors has not necessarily become more accurate. It may only have moved its threshold.
We looked for the accuracy gain on three open-weight models, at the quantity the improvement account predicts - discrimination - and it is not there. What we find instead is leniency:
- The episode moves the threshold. A completed audit --> repair episode already in context lowers false alarms by 2.8 to 11.5 pp against a length-matched non-audit control, in 15 of 15 model x wording combinations, with the present task held byte-identical.
- Discrimination does not follow. The criterion moves in 15 of 15 and survives correction in 13. d' survives in 0 of 15, though its test is half as sensitive by construction.
- Polarity does not explain it. An episode whose audit reported an error is more lenient still - the opposite sign to what the accumulated-message literature predicts.
The caution is about placement: be wary of putting a verifier in a context where it has already carried out the kind of repair it is about to judge.
Code: https://github.com/parsa-mz/crtitxer
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.16003 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.16003 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.16003 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.