Author here 👋</p>\n<p>We fine-tuned four MoE models (Qwen / gpt-oss / Nemotron families, 3.6–4.0B active params) to reason in Greek, and the headline is a null: accuracy barely moves. Worse, the benchmark can't see it anyway — changing <strong>only the random seed</strong> swings the score by 7.7 points, larger than every data and recipe effect we measured. If you're evaluating language-adaptation at this scale on accuracy alone, you're reading noise.</p>\n<p>The real change is elsewhere. Base models never think in Greek: <strong>0 of 1,000</strong> reasoning traces, even when the question is Greek. So the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After SFT, every checkpoint reasons in the language of the question on ~98% of items — one family at 3× fewer tokens — with grammaticality up on all four and general ability within a few points of base. Nothing forgotten, fluency gained.</p>\n<p>What SFT can't do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit \"think in English\" is obeyed under half the time. RLVR, pre-registered before training, fixes the first two outright (fallback 24% → 2.5%, leak 3.5% → 0.0%, both against a flat random-reward control) and moves the third +9.1pp — while the Greek reasoning habit survives an accuracy-only gradient untouched.</p>\n<p>We also propose six behavioural dimensions to make this measurable, each gated to reject any metric correlating with output length, and we report <strong>how our own instruments lied</strong>: six failures, each caught by a control. That section is the one we'd most like feedback on.</p>\n<p>Five checkpoints released. The instruments, controls and pre-registration travel to any low-resource language — Greek is just the case that let us measure them.</p>\n<p>Happy to answer anything here.</p>\n","updatedAt":"2026-08-21T12:22:30.530Z","author":{"_id":"6338c06c107c4835a05699f9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6338c06c107c4835a05699f9/uQbrMgQySY2UW7z3R9gT5.jpeg","fullname":"Ayoub Kirouane","name":"ayoubkirouane","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":70,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9396817088127136},"editors":["ayoubkirouane"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6338c06c107c4835a05699f9/uQbrMgQySY2UW7z3R9gT5.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.17744","authors":[{"_id":"6a85ace215847476b097d259","user":{"_id":"6338c06c107c4835a05699f9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6338c06c107c4835a05699f9/uQbrMgQySY2UW7z3R9gT5.jpeg","isPro":false,"fullname":"Ayoub Kirouane","user":"ayoubkirouane","type":"user","name":"ayoubkirouane"},"name":"Ayoub Kirouane","status":"claimed_verified","statusLastChangedAt":"2026-08-19T16:45:04.700Z","hidden":false},{"_id":"6a85ace215847476b097d25a","name":"Christos Petrocheilos","hidden":false}],"publishedAt":"2026-08-18T13:09:03.000Z","submittedOnDailyAt":"2026-08-21T00:00:00.000Z","title":"Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See","submittedOnDailyBy":{"_id":"6338c06c107c4835a05699f9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6338c06c107c4835a05699f9/uQbrMgQySY2UW7z3R9gT5.jpeg","isPro":false,"fullname":"Ayoub Kirouane","user":"ayoubkirouane","type":"user","name":"ayoubkirouane"},"summary":"Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After supervised fine-tuning (SFT), every released checkpoint reasons in the language of the question on ~98% of items, one family at 3x fewer tokens, with judged grammaticality improving on all four models and general ability within a few points of each base: nothing was forgotten, and fluency was gained. We propose six behavioural dimensions that make such changes measurable, each gated to reject any metric that correlates with output length, and we report how our own instruments lied: six failures, each caught by a control. What SFT cannot do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit \"think in English\" is obeyed under half the time. Reinforcement learning with verifiable rewards, pre-registered before training, fixes the first two outright (fallback 24% to 2.5%, leak 3.5% to 0.0%, both against a flat random-reward control) and moves the third (+9.1pp), while the Greek reasoning habit survives an accuracy-only gradient untouched. We release five checkpoints. The instruments, the controls and the pre-registration travel to any low-resource language; Greek is the case that let us measure them.","upvotes":3,"discussionId":"6a85ace215847476b097d25b","ai_summary":"Fine-tuning large mixture-of-experts models on a low-resource language shifts reasoning into that language without harming accuracy, while reinforcement learning with verifiable rewards fixes formatting and leakage defects.","ai_keywords":["mixture-of-experts","supervised fine-tuning","reinforcement learning with verifiable rewards","reasoning traces","low-resource language","behavioral dimensions","output-length gating"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"68f14d43f5266b706cf99e17","name":"KIEFERSA","fullname":"KIEFER","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68f14c3c6ede47be458ff051/7iTmA80GKtq3CNQDhvUoc.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65d082ca6130ef7be09f0422","avatarUrl":"/avatars/87d8feb75319ab57707791ce835531c2.svg","isPro":false,"fullname":"karim","user":"chiefkarim","type":"user"},{"_id":"6338c06c107c4835a05699f9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6338c06c107c4835a05699f9/uQbrMgQySY2UW7z3R9gT5.jpeg","isPro":false,"fullname":"Ayoub Kirouane","user":"ayoubkirouane","type":"user"},{"_id":"6799d2770a1d27acbcc635bb","avatarUrl":"/avatars/66af80ba8bba84f9105551fc41314b42.svg","isPro":false,"fullname":"Christos Petrocheilos","user":"cpetos","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68f14d43f5266b706cf99e17","name":"KIEFERSA","fullname":"KIEFER","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68f14c3c6ede47be458ff051/7iTmA80GKtq3CNQDhvUoc.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.17744.md","query":{}}">
Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See
Abstract
Fine-tuning large mixture-of-experts models on a low-resource language shifts reasoning into that language without harming accuracy, while reinforcement learning with verifiable rewards fixes formatting and leakage defects.
Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After supervised fine-tuning (SFT), every released checkpoint reasons in the language of the question on ~98% of items, one family at 3x fewer tokens, with judged grammaticality improving on all four models and general ability within a few points of each base: nothing was forgotten, and fluency was gained. We propose six behavioural dimensions that make such changes measurable, each gated to reject any metric that correlates with output length, and we report how our own instruments lied: six failures, each caught by a control. What SFT cannot do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit "think in English" is obeyed under half the time. Reinforcement learning with verifiable rewards, pre-registered before training, fixes the first two outright (fallback 24% to 2.5%, leak 3.5% to 0.0%, both against a flat random-reward control) and moves the third (+9.1pp), while the Greek reasoning habit survives an accuracy-only gradient untouched. We release five checkpoints. The instruments, the controls and the pre-registration travel to any low-resource language; Greek is the case that let us measure them.
Community
Author here 👋
We fine-tuned four MoE models (Qwen / gpt-oss / Nemotron families, 3.6–4.0B active params) to reason in Greek, and the headline is a null: accuracy barely moves. Worse, the benchmark can't see it anyway — changing only the random seed swings the score by 7.7 points, larger than every data and recipe effect we measured. If you're evaluating language-adaptation at this scale on accuracy alone, you're reading noise.
The real change is elsewhere. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek. So the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After SFT, every checkpoint reasons in the language of the question on ~98% of items — one family at 3× fewer tokens — with grammaticality up on all four and general ability within a few points of base. Nothing forgotten, fluency gained.
What SFT can't do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit "think in English" is obeyed under half the time. RLVR, pre-registered before training, fixes the first two outright (fallback 24% → 2.5%, leak 3.5% → 0.0%, both against a flat random-reward control) and moves the third +9.1pp — while the Greek reasoning habit survives an accuracy-only gradient untouched.
We also propose six behavioural dimensions to make this measurable, each gated to reject any metric correlating with output length, and we report how our own instruments lied: six failures, each caught by a control. That section is the one we'd most like feedback on.
Five checkpoints released. The instruments, controls and pre-registration travel to any low-resource language — Greek is just the case that let us measure them.
Happy to answer anything here.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.17744 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.17744 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.