What happens when on-policy RLVR improves the objective in front of it, but makes successful behavior for the next objective harder to sample?</p>\n<p>In this paper, we study this effect across mathematical reasoning and constrained instruction following. We call it <strong>verifier-induced support reshaping</strong>.</p>\n<p>On IFEval with Qwen3-8B-Base, Math-RLVR raises pass@1 by 6.5 percentage points relative to Base while lowering best@32 by 9.8 points: a random sample succeeds more often, yet fewer prompts remain recoverable across 32 samples. We further find that the largest measured policy shifts concentrate near response openings. In controlled opening interventions, route selection has a causal role in math searchability within the tested settings.</p>\n<p>We would love to hear whether others have observed similar effects in multi-stage or multi-objective RLVR.</p>\n<p>Paper:<br><a href=\"https://arxiv.org/pdf/2608.00220\" rel=\"nofollow\">https://arxiv.org/pdf/2608.00220</a></p>\n<p>Code:<br><a href=\"https://github.com/sylvain-wei/verifier-induced-support-reshaping\" rel=\"nofollow\">https://github.com/sylvain-wei/verifier-induced-support-reshaping</a></p>\n<p>Project page:<br><a href=\"https://sylvain-wei.github.io/verifier-induced-support-reshaping/\" rel=\"nofollow\">https://sylvain-wei.github.io/verifier-induced-support-reshaping/</a></p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/67244a81aa8556c561925ab6/YtIZI6b9Dp0ANYjoO6rvu.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/67244a81aa8556c561925ab6/YtIZI6b9Dp0ANYjoO6rvu.png\" alt=\"image\"></a><br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/67244a81aa8556c561925ab6/v89nFCORbhPTtMYwCHeDs.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/67244a81aa8556c561925ab6/v89nFCORbhPTtMYwCHeDs.png\" alt=\"image2\"></a></p>\n","updatedAt":"2026-08-17T03:29:54.360Z","author":{"_id":"67244a81aa8556c561925ab6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/w-vZ0uwYACagrNq-H1oyO.jpeg","fullname":"Shaohang Wei","name":"SylvainWei","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8557078242301941},"editors":["SylvainWei"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/w-vZ0uwYACagrNq-H1oyO.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.00220","authors":[{"_id":"6a827a1ab601d59c65281439","name":"Shaohang Wei","hidden":false},{"_id":"6a827a1ab601d59c6528143a","name":"Zikun Su","hidden":false},{"_id":"6a827a1ab601d59c6528143b","name":"Feifan Song","hidden":false},{"_id":"6a827a1ab601d59c6528143c","name":"Wen Luo","hidden":false},{"_id":"6a827a1ab601d59c6528143d","name":"Wei Li","hidden":false},{"_id":"6a827a1ab601d59c6528143e","name":"Guangyue Peng","hidden":false},{"_id":"6a827a1ab601d59c6528143f","name":"Houfeng Wang","hidden":false}],"publishedAt":"2026-07-31T00:00:00.000Z","submittedOnDailyAt":"2026-08-17T00:00:00.000Z","title":"Verifier-Induced Support Reshaping in On-Policy Optimization","submittedOnDailyBy":{"_id":"67244a81aa8556c561925ab6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/w-vZ0uwYACagrNq-H1oyO.jpeg","isPro":false,"fullname":"Shaohang Wei","user":"SylvainWei","type":"user","name":"SylvainWei"},"summary":"We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sample and reinforce. We call this verifier-induced support reshaping and define effective rewardable support as successful trajectories reachable within a fixed rollout budget. Across two model families, we study this effect through repeated verifier-scored sampling and bidirectional training on mathematical reasoning and constrained instruction following, including sequential training with the opposite verifier. Math-RLVR raises average instruction-following success but reduces the number of prompts with any successful response under repeated sampling. On IFEval with Qwen3-8B-Base, pass@1 rises by 6.5 percentage points while best@32 falls by 9.8 percentage points, and the same divergence appears across both models and IF benchmarks. Conversely, IF-RLVR shifts math responses from step-by-step openings toward direct answers, lowers best@k across sampling budgets, and reduces reward variation for later Math-RLVR. Token-distribution analyses and controlled opening interventions show that these changes concentrate in the first few response tokens. RLVR mainly reranks openings already available in the base policy, and the selected opening causally affects math searchability. The tested reference-policy constraints, routing priors, and on-policy distillation preserve cross-task support only partially; MathIF and ReasonIF show that marginal gains translate only partly into responses that are both correct and constraint-following. Therefore, endpoint improvements do not guarantee future trainability or joint capability under on-policy optimization. Code is available at https://github.com/sylvain-wei/verifier-induced-support-reshaping","upvotes":3,"discussionId":"6a827a1ab601d59c65281440","projectPage":"https://sylvain-wei.github.io/verifier-induced-support-reshaping/","githubRepo":"https://github.com/sylvain-wei/verifier-induced-support-reshaping","githubRepoAddedBy":"user","ai_summary":"On-policy reinforcement learning with verifiable rewards can improve immediate task performance while reducing the diversity of successful responses needed for future training, a phenomenon called verifier-induced support reshaping.","ai_keywords":["on-policy reinforcement learning","verifiable rewards","RLVR","verifier-induced support reshaping","effective rewardable support","best@k","pass@1","cross-task support","Math-RLVR","IF-RLVR"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"61dcd8e344f59573371b5cb6","name":"PekingUniversity","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vavgrBsnkSejriUF4lXDE.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"661ab1f1fa3b144a381fa454","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661ab1f1fa3b144a381fa454/IlpZBb9NCjo7ntFwMIH53.png","isPro":false,"fullname":"Urro","user":"urroxyz","type":"user"},{"_id":"67244a81aa8556c561925ab6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/w-vZ0uwYACagrNq-H1oyO.jpeg","isPro":false,"fullname":"Shaohang Wei","user":"SylvainWei","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61dcd8e344f59573371b5cb6","name":"PekingUniversity","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vavgrBsnkSejriUF4lXDE.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.00220.md","query":{}}">
Verifier-Induced Support Reshaping in On-Policy Optimization
Abstract
On-policy reinforcement learning with verifiable rewards can improve immediate task performance while reducing the diversity of successful responses needed for future training, a phenomenon called verifier-induced support reshaping.
We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sample and reinforce. We call this verifier-induced support reshaping and define effective rewardable support as successful trajectories reachable within a fixed rollout budget. Across two model families, we study this effect through repeated verifier-scored sampling and bidirectional training on mathematical reasoning and constrained instruction following, including sequential training with the opposite verifier. Math-RLVR raises average instruction-following success but reduces the number of prompts with any successful response under repeated sampling. On IFEval with Qwen3-8B-Base, pass@1 rises by 6.5 percentage points while best@32 falls by 9.8 percentage points, and the same divergence appears across both models and IF benchmarks. Conversely, IF-RLVR shifts math responses from step-by-step openings toward direct answers, lowers best@k across sampling budgets, and reduces reward variation for later Math-RLVR. Token-distribution analyses and controlled opening interventions show that these changes concentrate in the first few response tokens. RLVR mainly reranks openings already available in the base policy, and the selected opening causally affects math searchability. The tested reference-policy constraints, routing priors, and on-policy distillation preserve cross-task support only partially; MathIF and ReasonIF show that marginal gains translate only partly into responses that are both correct and constraint-following. Therefore, endpoint improvements do not guarantee future trainability or joint capability under on-policy optimization. Code is available at https://github.com/sylvain-wei/verifier-induced-support-reshaping
Community
What happens when on-policy RLVR improves the objective in front of it, but makes successful behavior for the next objective harder to sample?
In this paper, we study this effect across mathematical reasoning and constrained instruction following. We call it verifier-induced support reshaping.
On IFEval with Qwen3-8B-Base, Math-RLVR raises pass@1 by 6.5 percentage points relative to Base while lowering best@32 by 9.8 points: a random sample succeeds more often, yet fewer prompts remain recoverable across 32 samples. We further find that the largest measured policy shifts concentrate near response openings. In controlled opening interventions, route selection has a causal role in math searchability within the tested settings.
We would love to hear whether others have observed similar effects in multi-stage or multi-objective RLVR.
Paper:
https://arxiv.org/pdf/2608.00220
Code:
https://github.com/sylvain-wei/verifier-induced-support-reshaping
Project page:
https://sylvain-wei.github.io/verifier-induced-support-reshaping/


Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.00220 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.00220 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.00220 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.