Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed weighted sum before group-wise standardization. We show that this design leads to two fundamental problems: rollouts with distinct reward profiles can receive identical advantages, and all objectives are optimized with fixed relative weights regardless of their current level of saturation. As a result, training continues to allocate gradient budget to already-solved objectives instead of focusing on those with greater remaining headroom.</p>\n<p>We introduce \\textbf{Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization} (SA-MRPO), which standardizes each reward objective independently and adaptively discounts its contribution according to a batch-level estimate of objective saturation. This dynamically reallocates optimization effort toward under-optimized objectives while empirically maintaining performance on those that are already well satisfied. We further show that saturation-aware reweighting can reverse the sign of an update, rather than merely rescale its magnitude.</p>\n<p> Across mathematical reasoning with two- and three-objective reward combinations, SA-MRPO improves the harder correctness objective over GDPO in 12 of 15 benchmark comparisons, with gains of up to $5%$ on AIME24. On adaptive reasoning, it improves accuracy on all five benchmarks by 3.8% on average and up to 9.2% on AMC23, and on coding benchmarks it improves the pass rate by up to 2.3%, while in all settings maintaining the easier objectives near their already satisfied levels.</p>\n","updatedAt":"2026-08-18T06:18:26.737Z","author":{"_id":"6419309f22270b3ccf177c77","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6419309f22270b3ccf177c77/KQa1586iBBKqucUlfpuPp.jpeg","fullname":"William Li","name":"williamium","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9234110713005066},"editors":["williamium"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6419309f22270b3ccf177c77/KQa1586iBBKqucUlfpuPp.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.16072","authors":[{"_id":"6a83f8a0675db694db8cd635","name":"Yixuan Wang","hidden":false},{"_id":"6a83f8a0675db694db8cd636","name":"Yifei Chen","hidden":false},{"_id":"6a83f8a0675db694db8cd637","name":"Haichao Zhang","hidden":false},{"_id":"6a83f8a0675db694db8cd638","name":"Haozheng Luo","hidden":false},{"_id":"6a83f8a0675db694db8cd639","name":"Xander Wu","hidden":false},{"_id":"6a83f8a0675db694db8cd63a","name":"Jie Ni","hidden":false},{"_id":"6a83f8a0675db694db8cd63b","name":"Yun Fu","hidden":false},{"_id":"6a83f8a0675db694db8cd63c","name":"Nuno Vasconcelos","hidden":false},{"_id":"6a83f8a0675db694db8cd63d","name":"Yijiang Li","hidden":false}],"publishedAt":"2026-08-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization","submittedOnDailyBy":{"_id":"6419309f22270b3ccf177c77","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6419309f22270b3ccf177c77/KQa1586iBBKqucUlfpuPp.jpeg","isPro":true,"fullname":"William Li","user":"williamium","type":"user","name":"williamium"},"summary":"Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed weighted sum before group-wise standardization. We show that this design leads to two fundamental problems: rollouts with distinct reward profiles can receive identical advantages, and all objectives are optimized with fixed relative weights regardless of their current level of saturation. As a result, training continues to allocate gradient budget to already-solved objectives instead of focusing on those with greater remaining headroom. We introduce Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization (SA-MRPO), which standardizes each reward objective independently and adaptively discounts its contribution according to a batch-level estimate of objective saturation. This dynamically reallocates optimization effort toward under-optimized objectives while empirically maintaining performance on those that are already well satisfied. We further show that saturation-aware reweighting can reverse the sign of an update, rather than merely rescale its magnitude. Across mathematical reasoning with two- and three-objective reward combinations, SA-MRPO improves the harder correctness objective over GDPO in 12 of 15 benchmark comparisons, with gains of up to 5% on AIME24. On adaptive reasoning it improves accuracy on all five benchmarks, by 3.8% on average and up to 9.2 % on AMC23, and on coding benchmarks it improves pass rate by up to 2.3%, while in all settings maintaining the easier objectives near their already satisfied levels.","upvotes":14,"discussionId":"6a83f8a0675db694db8cd63e","ai_summary":"SA-MRPO independently standardizes multi-objective rewards and adaptively discounts saturated objectives to redirect optimization toward under-optimized goals.","ai_keywords":["group-relative advantages","multi-reward policy optimization","saturation-aware advantage reweighting","SA-MRPO","objective saturation","gradient budget reallocation"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"697e87d12cc19315a8497001","name":"UCSanDiego","fullname":"University of California at San Diego","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/697e8687c00f332cf492d29e/KUQpvngxP4r9oBSDZwIwZ.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6419309f22270b3ccf177c77","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6419309f22270b3ccf177c77/KQa1586iBBKqucUlfpuPp.jpeg","isPro":true,"fullname":"William Li","user":"williamium","type":"user"},{"_id":"63ef0af2bfe4ead22ca8f69a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676610243576-noauth.jpeg","isPro":false,"fullname":"Haozheng Luo","user":"robinzixuan","type":"user"},{"_id":"646c866497819a8be939e882","avatarUrl":"/avatars/505a5c1ab24dad4ccd90efceaf36deff.svg","isPro":false,"fullname":"Yixuan Wang","user":"noMushroom","type":"user"},{"_id":"662119a94f00769a14f1a665","avatarUrl":"/avatars/660f75b2aee2213d967caa289f03503b.svg","isPro":false,"fullname":"Yifei Chen","user":"YifeiCnaranja","type":"user"},{"_id":"674d3f23671ffd5aefadc16c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/QJ1ClKhrk5PaAERc9-d0y.png","isPro":false,"fullname":"Guan","user":"Jerry-PigeonG","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6a1590059224246011aba2b6","avatarUrl":"/avatars/d9e158378a6f9dbdbd7819bd4fef2ce4.svg","isPro":false,"fullname":"Dong Yutian","user":"doyu0d","type":"user"},{"_id":"6a1470b7b28ec6a2ad92193c","avatarUrl":"/avatars/6243b0a5d5d3c8bec740d07bdbb12950.svg","isPro":false,"fullname":"Grayson Scott","user":"grascott99","type":"user"},{"_id":"6a14682131430965c63f1bc3","avatarUrl":"/avatars/06ff73863b78bc69f313c8d2015ff241.svg","isPro":false,"fullname":"Hu Linxi","user":"hulinxi","type":"user"},{"_id":"6a159979cf287713e52336ad","avatarUrl":"/avatars/ac05039b5d575768350567bdb7397a4f.svg","isPro":false,"fullname":"Lily Torres","user":"lilytorres6","type":"user"},{"_id":"6a15e69a8642d018701cf529","avatarUrl":"/avatars/d4239043b35c3c8552b4dce80d19084b.svg","isPro":false,"fullname":"Thomas ROBINSON","user":"thomasmk30","type":"user"},{"_id":"6a15dabccfff5937535b56f1","avatarUrl":"/avatars/c673889a37f80cc19bf6bef0f60b2172.svg","isPro":false,"fullname":"Mateo Smith","user":"msmith25","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"697e87d12cc19315a8497001","name":"UCSanDiego","fullname":"University of California at San Diego","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/697e8687c00f332cf492d29e/KUQpvngxP4r9oBSDZwIwZ.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.16072.md","query":{}}">
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
Abstract
SA-MRPO independently standardizes multi-objective rewards and adaptively discounts saturated objectives to redirect optimization toward under-optimized goals.
Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed weighted sum before group-wise standardization. We show that this design leads to two fundamental problems: rollouts with distinct reward profiles can receive identical advantages, and all objectives are optimized with fixed relative weights regardless of their current level of saturation. As a result, training continues to allocate gradient budget to already-solved objectives instead of focusing on those with greater remaining headroom. We introduce Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization (SA-MRPO), which standardizes each reward objective independently and adaptively discounts its contribution according to a batch-level estimate of objective saturation. This dynamically reallocates optimization effort toward under-optimized objectives while empirically maintaining performance on those that are already well satisfied. We further show that saturation-aware reweighting can reverse the sign of an update, rather than merely rescale its magnitude. Across mathematical reasoning with two- and three-objective reward combinations, SA-MRPO improves the harder correctness objective over GDPO in 12 of 15 benchmark comparisons, with gains of up to 5% on AIME24. On adaptive reasoning it improves accuracy on all five benchmarks, by 3.8% on average and up to 9.2 % on AMC23, and on coding benchmarks it improves pass rate by up to 2.3%, while in all settings maintaining the easier objectives near their already satisfied levels.
Community
Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed weighted sum before group-wise standardization. We show that this design leads to two fundamental problems: rollouts with distinct reward profiles can receive identical advantages, and all objectives are optimized with fixed relative weights regardless of their current level of saturation. As a result, training continues to allocate gradient budget to already-solved objectives instead of focusing on those with greater remaining headroom.
We introduce \textbf{Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization} (SA-MRPO), which standardizes each reward objective independently and adaptively discounts its contribution according to a batch-level estimate of objective saturation. This dynamically reallocates optimization effort toward under-optimized objectives while empirically maintaining performance on those that are already well satisfied. We further show that saturation-aware reweighting can reverse the sign of an update, rather than merely rescale its magnitude.
Across mathematical reasoning with two- and three-objective reward combinations, SA-MRPO improves the harder correctness objective over GDPO in 12 of 15 benchmark comparisons, with gains of up to $5%$ on AIME24. On adaptive reasoning, it improves accuracy on all five benchmarks by 3.8% on average and up to 9.2% on AMC23, and on coding benchmarks it improves the pass rate by up to 2.3%, while in all settings maintaining the easier objectives near their already satisfied levels.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.16072 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.16072 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.16072 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.