New work from Hume AI on quantifying how an ASR model reproduces a benchmark's reference text rather than transcribing the audio</p>\n","updatedAt":"2026-08-21T05:55:08.144Z","author":{"_id":"6384db7fb2906edaf835a91d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6384db7fb2906edaf835a91d/MOTXxaOmjlTZ8wONYifnD.jpeg","fullname":"Eric Bezzam","name":"bezzam","type":"user","isPro":true,"isHf":true,"isHfAdmin":false,"isMod":false,"followerCount":126,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1583856921041-5dd96eb166059660ed1ee413.png","fullname":"Hugging Face","name":"huggingface","type":"org","isHf":true,"details":"The AI community building the future.","plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8191012144088745},"editors":["bezzam"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6384db7fb2906edaf835a91d/MOTXxaOmjlTZ8wONYifnD.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.19936","authors":[{"_id":"6a87e79a89e517cbfd75dcc7","name":"Theo Lebryk","hidden":false},{"_id":"6a87e79a89e517cbfd75dcc8","name":"David Ayllon","hidden":false},{"_id":"6a87e79a89e517cbfd75dcc9","name":"Alice Baird","hidden":false},{"_id":"6a87e79a89e517cbfd75dcca","name":"Jakub Piotr Cłapa","hidden":false},{"_id":"6a87e79a89e517cbfd75dccb","name":"Jens Madsen","hidden":false},{"_id":"6a87e79a89e517cbfd75dccc","name":"Panagiotis Tzirakis","hidden":false}],"publishedAt":"2026-08-20T00:00:00.000Z","submittedOnDailyAt":"2026-08-21T00:00:00.000Z","title":"Towards Quantifying Benchmark Optimization in ASR Models","submittedOnDailyBy":{"_id":"6384db7fb2906edaf835a91d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6384db7fb2906edaf835a91d/MOTXxaOmjlTZ8wONYifnD.jpeg","isPro":true,"fullname":"Eric Bezzam","user":"bezzam","type":"user","name":"bezzam"},"summary":"Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.","upvotes":2,"discussionId":"6a87e79b89e517cbfd75dccd","githubRepo":"https://github.com/HumeAI/asr-benchmark-optimization","githubRepoAddedBy":"user","ai_summary":"High-performing speech recognition models reproduce benchmark transcripts despite contradictory audio, revealing benchmark-optimized behaviors that inflate scores without improving real-world transcription.","ai_keywords":["Automatic Speech Recognition","benchmark optimization","reference disagreement","masked-number recovery","orthographic switching","mechanistic probes","low-rank linear steering"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"6776f7c98fcda5708ecf384c","name":"HumeAI","fullname":"Hume AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62a4e6b36ee324b134c5a6f1/dLbgpDAbPEirXkgH66tlc.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6384db7fb2906edaf835a91d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6384db7fb2906edaf835a91d/MOTXxaOmjlTZ8wONYifnD.jpeg","isPro":true,"fullname":"Eric Bezzam","user":"bezzam","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6776f7c98fcda5708ecf384c","name":"HumeAI","fullname":"Hume AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62a4e6b36ee324b134c5a6f1/dLbgpDAbPEirXkgH66tlc.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.19936.md","query":{}}">
Towards Quantifying Benchmark Optimization in ASR Models
Abstract
High-performing speech recognition models reproduce benchmark transcripts despite contradictory audio, revealing benchmark-optimized behaviors that inflate scores without improving real-world transcription.
Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.
Community
New work from Hume AI on quantifying how an ASR model reproduces a benchmark's reference text rather than transcribing the audio
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.19936 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.19936 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.19936 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.