Hugging Face Daily Papers · · 4 min read

D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Multi-teacher on-policy distillation trains a single student from several domain-expert teachers, but different domains converge at very different rates, so the fixed data mixtures used in prior work keep spending rollouts on domains that have already saturated while starving the ones that still have headroom. D³-MOPD reuses the per-domain reverse-KL that MOPD already computes as a progress signal: an off-process watcher tracks each domain's KL trajectory to estimate its remaining headroom and current improvement rate, and reallocates sampling ratios on the fly — no extra probes, no change to the training loop.</p>\n","updatedAt":"2026-08-27T02:37:49.227Z","author":{"_id":"65328aa39326d6da5ff19b52","avatarUrl":"/avatars/5c3de984cd6eba69616bb608796865c5.svg","fullname":"Fei Zhao","name":"Hiiamein","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9225901365280151},"editors":["Hiiamein"],"editorAvatarUrls":["/avatars/5c3de984cd6eba69616bb608796865c5.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.24987","authors":[{"_id":"6a8fa16a2c24e8c5fab32940","user":{"_id":"671b8777ac4168f79848b282","avatarUrl":"/avatars/b4b8062a8fe890fb6bac8917630bfb5a.svg","isPro":false,"fullname":"Zechen Sun","user":"Mintszc","type":"user","name":"Mintszc"},"name":"Zechen Sun","status":"claimed_verified","statusLastChangedAt":"2026-08-27T08:45:04.933Z","hidden":false},{"_id":"6a8fa16a2c24e8c5fab32941","name":"Zhiwei Zhang","hidden":false},{"_id":"6a8fa16a2c24e8c5fab32942","name":"Fei Zhao","hidden":false},{"_id":"6a8fa16a2c24e8c5fab32943","name":"Juntao Li","hidden":false},{"_id":"6a8fa16a2c24e8c5fab32944","name":"Mu Chuan","hidden":false},{"_id":"6a8fa16a2c24e8c5fab32945","name":"Huayu Deng","hidden":false},{"_id":"6a8fa16a2c24e8c5fab32946","name":"Guojian Zhan","hidden":false},{"_id":"6a8fa16a2c24e8c5fab32947","name":"Wenliang Chen","hidden":false},{"_id":"6a8fa16a2c24e8c5fab32948","name":"Yao Hu","hidden":false},{"_id":"6a8fa16a2c24e8c5fab32949","name":"Min Zhang","hidden":false}],"publishedAt":"2026-08-25T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation","submittedOnDailyBy":{"_id":"65328aa39326d6da5ff19b52","avatarUrl":"/avatars/5c3de984cd6eba69616bb608796865c5.svg","isPro":false,"fullname":"Fei Zhao","user":"Hiiamein","type":"user","name":"Hiiamein"},"summary":"Multi-teacher on-policy distillation (MOPD) distills several domain-expert teachers into a single student by minimizing per-domain reverse-KL divergence on the student's own rollouts. Existing approaches typically fix the per-domain data mixture before training, overlooking the fact that different domains converge at substantially different rates: some plateau early while others continue to improve throughout the training budget. A fixed mixture therefore wastes compute on fast-converging domains and undertrains slower-converging ones. To address this, we propose D^3-MOPD (Dynamic Domain ScheDuling for MOPD), a zero-overhead scheduler that repurposes the per-domain reverse-KL signal already produced during training to adapt the domain mixture online. Running asynchronously outside the training process, an off-process watcher periodically tracks each domain's KL trajectory, estimates remaining headroom and current improvement rate, and accordingly adjusts the domain sampling ratios without altering the core training loop. Our D^3-MOPD scales naturally to arbitrary numbers of domains, and the expected benefit grows as more domains introduce more diverse convergence patterns for the scheduler to exploit. On a Qwen3.6-35B-A3B student distilled from four domain-expert teachers, D^3-MOPD closes 97% of the average student-to-teacher performance gap, compared with 63% for vanilla MOPD, reaches the same peak performance with an approximately 3times reduction in rollout steps, and surpasses the specialist teachers on three of seven benchmarks.","upvotes":18,"discussionId":"6a8fa16b2c24e8c5fab3294a","ai_summary":"D³-MOPD dynamically adjusts domain sampling ratios during multi-teacher distillation by monitoring per-domain reverse-KL trajectories, improving convergence efficiency and closing most of the student-to-teacher performance gap.","ai_keywords":["multi-teacher on-policy distillation","reverse-KL divergence","domain mixture","D³-MOPD","off-process watcher","KL trajectory"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65328aa39326d6da5ff19b52","avatarUrl":"/avatars/5c3de984cd6eba69616bb608796865c5.svg","isPro":false,"fullname":"Fei Zhao","user":"Hiiamein","type":"user"},{"_id":"678b35ef0d44299225201c71","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/O8DWzE62cSFclPXpoE_vT.png","isPro":false,"fullname":"zhangzhiwei","user":"zhangzhiwei666","type":"user"},{"_id":"644e3e5f030210812f413073","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/uW8TKV2sds97lFBwnt6JK.jpeg","isPro":true,"fullname":"Zilong Chen","user":"heheyas","type":"user"},{"_id":"671b8777ac4168f79848b282","avatarUrl":"/avatars/b4b8062a8fe890fb6bac8917630bfb5a.svg","isPro":false,"fullname":"Zechen Sun","user":"Mintszc","type":"user"},{"_id":"611e25751f0dcb7bec13d0d4","avatarUrl":"/avatars/e019a455351e402d8b349878deb94192.svg","isPro":false,"fullname":"Jing Ye","user":"1245244103","type":"user"},{"_id":"64c1ec9db005aab93d622ba3","avatarUrl":"/avatars/907b806f5ffa0d9b7bee541e36e23a8b.svg","isPro":false,"fullname":"zzz","user":"zhengzzz","type":"user"},{"_id":"687771f9c38b08df75a56410","avatarUrl":"/avatars/2cb7e6b0db97097cd749a4c4285bb770.svg","isPro":false,"fullname":"syy","user":"mirrorball799","type":"user"},{"_id":"6732fb1d24b316be87acaafe","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6732fb1d24b316be87acaafe/BzD8HL4vhh3mfeSF3rm_1.jpeg","isPro":false,"fullname":"Quantong Qiu","user":"QQTang1223","type":"user"},{"_id":"671b4b355a5fadd960a2f7b6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/671b4b355a5fadd960a2f7b6/ZlZU2o43XAvTQb7Bdx-Io.jpeg","isPro":false,"fullname":"Ruoxi Sun","user":"xii0929","type":"user"},{"_id":"655ca632e3aff30e60ca4529","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/655ca632e3aff30e60ca4529/O1M64LqhHMT61rrIqKCWN.jpeg","isPro":false,"fullname":"sora","user":"AmamiSora","type":"user"},{"_id":"68144b68a706ff01feda036c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68144b68a706ff01feda036c/vlQHsP2dG3ha_nrP9_WNr.jpeg","isPro":false,"fullname":"ymrl","user":"ymrl","type":"user"},{"_id":"6794fb4346f22e87c856732a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/nZAwcm_jqwOKk7oYcRQk6.png","isPro":false,"fullname":"Zhiyi Hong","user":"ACEEE1222","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.24987.md","query":{}}">
Papers
arxiv:2608.24987

D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

Published on Aug 25
· Submitted by
Fei Zhao
on Aug 27
Authors:

Abstract

D³-MOPD dynamically adjusts domain sampling ratios during multi-teacher distillation by monitoring per-domain reverse-KL trajectories, improving convergence efficiency and closing most of the student-to-teacher performance gap.

Multi-teacher on-policy distillation (MOPD) distills several domain-expert teachers into a single student by minimizing per-domain reverse-KL divergence on the student's own rollouts. Existing approaches typically fix the per-domain data mixture before training, overlooking the fact that different domains converge at substantially different rates: some plateau early while others continue to improve throughout the training budget. A fixed mixture therefore wastes compute on fast-converging domains and undertrains slower-converging ones. To address this, we propose D^3-MOPD (Dynamic Domain ScheDuling for MOPD), a zero-overhead scheduler that repurposes the per-domain reverse-KL signal already produced during training to adapt the domain mixture online. Running asynchronously outside the training process, an off-process watcher periodically tracks each domain's KL trajectory, estimates remaining headroom and current improvement rate, and accordingly adjusts the domain sampling ratios without altering the core training loop. Our D^3-MOPD scales naturally to arbitrary numbers of domains, and the expected benefit grows as more domains introduce more diverse convergence patterns for the scheduler to exploit. On a Qwen3.6-35B-A3B student distilled from four domain-expert teachers, D^3-MOPD closes 97% of the average student-to-teacher performance gap, compared with 63% for vanilla MOPD, reaches the same peak performance with an approximately 3times reduction in rollout steps, and surpasses the specialist teachers on three of seven benchmarks.

Community

Paper submitter about 6 hours ago

Multi-teacher on-policy distillation trains a single student from several domain-expert teachers, but different domains converge at very different rates, so the fixed data mixtures used in prior work keep spending rollouts on domains that have already saturated while starving the ones that still have headroom. D³-MOPD reuses the per-domain reverse-KL that MOPD already computes as a progress signal: an off-process watcher tracks each domain's KL trajectory to estimate its remaining headroom and current improvement rate, and reallocates sampling ratios on the fly — no extra probes, no change to the training loop.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.24987
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.24987 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.24987 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.24987 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers