Hugging Face Daily Papers · · 4 min read

Length-Adaptive Decoding for Masked Diffusion Machine Translation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Masked diffusion LMs must pick the target canvas length before denoising starts — there is no autoregressive EOS to stop generation, so a wrong length drops source content or pads the output before a single token is revealed. Entropy-Valley takes that decision out of the hyperparameters and puts it in the model: probe each candidate canvas with a single all-mask forward pass and decode the one with the lowest mean predictive entropy — the canvas the frozen backbone is most prepared to fill. Training-free, no length head, no reference lengths.<br>On WMT22 with LLaDA-8B it closes 64.9% / 65.3% / 33.0% of the gap to a reference-length oracle on En→Zh / Zh→En / En→De, three expert translators confirm the gain is adequacy rather than fluency, and it beats DAEDAL in both En↔Zh directions and CAL on En→Zh at lower measured inference cost. The pattern holds on Dream-Base and DiffuLLaMA.</p>\n","updatedAt":"2026-08-26T02:58:57.698Z","author":{"_id":"65f974fc55d5ad897260fc67","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Sa3flUvGuzwnMvkw4jvt2.png","fullname":"Yan Zhan","name":"YanZhanPKU","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8501120805740356},"editors":["YanZhanPKU"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Sa3flUvGuzwnMvkw4jvt2.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.22274","authors":[{"_id":"6a8d20065add2537c32e97b7","user":{"_id":"65f974fc55d5ad897260fc67","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Sa3flUvGuzwnMvkw4jvt2.png","isPro":false,"fullname":"Yan Zhan","user":"YanZhanPKU","type":"user","name":"YanZhanPKU"},"name":"Yan Zhan","status":"claimed_verified","statusLastChangedAt":"2026-08-25T08:13:40.460Z","hidden":false},{"_id":"6a8d20065add2537c32e97b8","name":"Mengkai Hou","hidden":false},{"_id":"6a8d20065add2537c32e97b9","name":"Wanting Zhang","hidden":false},{"_id":"6a8d20065add2537c32e97ba","name":"Zhijun Gao","hidden":false}],"publishedAt":"2026-08-23T00:00:00.000Z","submittedOnDailyAt":"2026-08-26T00:00:00.000Z","title":"Length-Adaptive Decoding for Masked Diffusion Machine Translation","submittedOnDailyBy":{"_id":"65f974fc55d5ad897260fc67","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Sa3flUvGuzwnMvkw4jvt2.png","isPro":false,"fullname":"Yan Zhan","user":"YanZhanPKU","type":"user","name":"YanZhanPKU"},"summary":"Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithfully, while fixed canvas decoding must choose target length before denoising. Existing masked diffusion decoding work mainly studies token unmasking order, leaving this length decision under-explored despite its direct effect on coverage and redundancy. We introduce Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entropy from all-mask forward passes and selects the canvas the backbone is most prepared to fill. Relative to a baseline using training corpus length statistics, EV recovers 64.9%, 65.3%, and 33.0% of the COMET-22 gain from reference target lengths on EntoZh, ZhtoEn, and EntoDe. Our diagnostics show that denoising-friendly lengths need not match reference lengths. Evaluation by three translation experts supports the EnleftrightarrowZh adequacy gains, with stronger evidence on ZhtoEn. Compared with a LLaMA-3-8B autoregressive (AR) model trained on the same fine-tuning data, the EV system ties on EntoZh and leads on ZhtoEn; an oracle-length diagnostic further shows that, in this masked diffusion MT setting, deciding which tokens to reveal first matters less than how the target length is supplied.","upvotes":3,"discussionId":"6a8d20065add2537c32e97bb","projectPage":"https://huggingface.co/collections/YanZhanPKU/entropy-valley","githubRepo":"https://github.com/Entropy-Valley/Entropy-Valley","githubRepoAddedBy":"user","ai_summary":"Entropy-Valley selects target lengths for masked diffusion translation by scoring predictive entropy, improving adequacy and showing length choice matters more than unmasking order.","ai_keywords":["masked diffusion language models","dLLMs","entropy-valley","predictive entropy","all-mask forward passes","COMET-22","denoising-friendly lengths","autoregressive","masked diffusion MT"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"61dcd8e344f59573371b5cb6","name":"PekingUniversity","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vavgrBsnkSejriUF4lXDE.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65f974fc55d5ad897260fc67","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Sa3flUvGuzwnMvkw4jvt2.png","isPro":false,"fullname":"Yan Zhan","user":"YanZhanPKU","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"67243cde84b36676c782a5c3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/E5f7_kchaXVh0GDzOrgLR.png","isPro":false,"fullname":"Mengkai Hou","user":"HMonKY","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61dcd8e344f59573371b5cb6","name":"PekingUniversity","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vavgrBsnkSejriUF4lXDE.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.22274.md","query":{}}">
Papers
arxiv:2608.22274

Length-Adaptive Decoding for Masked Diffusion Machine Translation

Published on Aug 23
· Submitted by
Yan Zhan
on Aug 26
Authors:

Abstract

Entropy-Valley selects target lengths for masked diffusion translation by scoring predictive entropy, improving adequacy and showing length choice matters more than unmasking order.

Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithfully, while fixed canvas decoding must choose target length before denoising. Existing masked diffusion decoding work mainly studies token unmasking order, leaving this length decision under-explored despite its direct effect on coverage and redundancy. We introduce Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entropy from all-mask forward passes and selects the canvas the backbone is most prepared to fill. Relative to a baseline using training corpus length statistics, EV recovers 64.9%, 65.3%, and 33.0% of the COMET-22 gain from reference target lengths on EntoZh, ZhtoEn, and EntoDe. Our diagnostics show that denoising-friendly lengths need not match reference lengths. Evaluation by three translation experts supports the EnleftrightarrowZh adequacy gains, with stronger evidence on ZhtoEn. Compared with a LLaMA-3-8B autoregressive (AR) model trained on the same fine-tuning data, the EV system ties on EntoZh and leads on ZhtoEn; an oracle-length diagnostic further shows that, in this masked diffusion MT setting, deciding which tokens to reveal first matters less than how the target length is supplied.

Community

Paper author Paper submitter about 5 hours ago

Masked diffusion LMs must pick the target canvas length before denoising starts — there is no autoregressive EOS to stop generation, so a wrong length drops source content or pads the output before a single token is revealed. Entropy-Valley takes that decision out of the hyperparameters and puts it in the model: probe each candidate canvas with a single all-mask forward pass and decode the one with the lowest mean predictive entropy — the canvas the frozen backbone is most prepared to fill. Training-free, no length head, no reference lengths.
On WMT22 with LLaDA-8B it closes 64.9% / 65.3% / 33.0% of the gap to a reference-length oracle on En→Zh / Zh→En / En→De, three expert translators confirm the gain is adequacy rather than fluency, and it beats DAEDAL in both En↔Zh directions and CAL on En→Zh at lower measured inference cost. The pattern holds on Dream-Base and DiffuLLaMA.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.22274
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.22274 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers