Hugging Face Daily Papers · · 4 min read

WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

The project page provides an overview of WarpSAC, demonstrations, benchmark results, and sim-to-real evaluations. The GitHub repository contains the source code, installation instructions, training scripts, environment integrations, and experiment configurations.</p>\n","updatedAt":"2026-08-27T04:17:01.729Z","author":{"_id":"6695eb6c0a6ba12a31c3a1be","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6695eb6c0a6ba12a31c3a1be/g5t82nhTIXYe92mKL535o.jpeg","fullname":"Yifu Yuan","name":"IffYuan","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7802889943122864},"editors":["IffYuan"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6695eb6c0a6ba12a31c3a1be/g5t82nhTIXYe92mKL535o.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.24479","authors":[{"_id":"6a8e524a7bc881afa25f30c1","name":"Zihao Wu","hidden":false},{"_id":"6a8e524a7bc881afa25f30c2","name":"Hongyao Tang","hidden":false},{"_id":"6a8e524a7bc881afa25f30c3","name":"Yi Ma","hidden":false},{"_id":"6a8e524a7bc881afa25f30c4","name":"Huizhong Song","hidden":false},{"_id":"6a8e524a7bc881afa25f30c5","name":"Pengyi Li","hidden":false},{"_id":"6a8e524a7bc881afa25f30c6","name":"Yifu Yuan","hidden":false},{"_id":"6a8e524a7bc881afa25f30c7","name":"Fei Ni","hidden":false},{"_id":"6a8e524a7bc881afa25f30c8","name":"Jinyi Liu","hidden":false},{"_id":"6a8e524a7bc881afa25f30c9","name":"Wei Wei","hidden":false},{"_id":"6a8e524a7bc881afa25f30ca","name":"Jianrong Wang","hidden":false},{"_id":"6a8e524a7bc881afa25f30cb","name":"Yan Zheng","hidden":false},{"_id":"6a8e524a7bc881afa25f30cc","name":"Jianye Hao","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6695eb6c0a6ba12a31c3a1be/EEPZUSWBKIO_CfnKEqF1K.mp4"],"publishedAt":"2026-08-25T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation","submittedOnDailyBy":{"_id":"6695eb6c0a6ba12a31c3a1be","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6695eb6c0a6ba12a31c3a1be/g5t82nhTIXYe92mKL535o.jpeg","isPro":false,"fullname":"Yifu Yuan","user":"IffYuan","type":"user","name":"IffYuan"},"summary":"Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for data-limited replay. Through controlled experiments across eight benchmark families, we show that these stabilizers are data-regime-dependent: parameter normalization helps with narrow replay coverage but restricts value fitting when data are abundant, while clipped double-Q can be relaxed in high-throughput manipulation. Age-biased replay weighting improves learning efficiency across regimes, especially with limited network capacity.\n Based on these findings, we propose WarpSAC, a regime-aware family of off-policy RL algorithms. WarpSAC uses Sample Weight Decay for efficient exploitation and provides two variants: WarpSAC-L (Norm ON, clipped double-Q) for data-limited CPU-scale training, and WarpSAC-A (Norm OFF, single-Q) for data-abundant GPU-parallel training. WarpSAC improves normalized score--step AUC over FlashSAC by 4.5% across nine CPU-scale environments and 23.1% across fourteen GPU-parallel environments. It increases UnitreeG1TransportBox-v1 success rate from 19.8% to 96.4%, improves mean normalized wall-time AUC on MuJoCo Playground by 19.1%, and achieves 36.4% faster sim-to-real deployment on Unitree G1 than FlashSAC. These results show that scalable off-policy RL should adapt its stabilizers to the available data regime.","upvotes":77,"discussionId":"6a8e524a7bc881afa25f30cd","projectPage":"https://wzhhasadream.github.io/WarpSAC","githubRepo":"https://github.com/wzhhasadream/warprl","githubRepoAddedBy":"user","ai_summary":"Off-policy reinforcement learning stabilizers vary with data availability, motivating regime-aware algorithms that adapt normalization and Q-function clipping to improve efficiency across CPU and GPU-parallel training.","ai_keywords":["off-policy reinforcement learning","replay","parameter normalization","clipped double-Q","age-biased replay weighting","WarpSAC","Sample Weight Decay","sim-to-real"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"64b74d6c9f5987572c9be2b7","name":"TianjinUniversity","fullname":"Tianjin University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b74c979ebb7e6c7dd3c59f/OWmVv_2a2SIsp1MjkERFn.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69be5202e3f99173bbe7c7a0","avatarUrl":"/avatars/a47bcaf2fc2f3fcbd56992e5b22642b8.svg","isPro":false,"fullname":"wzh","user":"wzhhasadream","type":"user"},{"_id":"664cae9d57e83e254b5f7511","avatarUrl":"/avatars/6d1aa56adc3b85866db19ede43099ca9.svg","isPro":false,"fullname":"Hongyao Tang","user":"thy-1now","type":"user"},{"_id":"6a15c99948e6f96e34025eb1","avatarUrl":"/avatars/1273c8b65a9dd06b0004a33b30c72294.svg","isPro":false,"fullname":"LUO Yiran","user":"LUYI2026","type":"user"},{"_id":"6a14686491d4af20d4030a76","avatarUrl":"/avatars/6c36c9a3d564174094e7ddbfc5260ae7.svg","isPro":false,"fullname":"Huang Siyu","user":"huangsiy3","type":"user"},{"_id":"6a147d3193b9759a7f062980","avatarUrl":"/avatars/167f360f649939b767745a7b30e2ac7a.svg","isPro":false,"fullname":"Aurora Hill","user":"aurorahi8","type":"user"},{"_id":"6a14564a3789c92742679ffd","avatarUrl":"/avatars/43623d2e0331834f07b7193b28be494c.svg","isPro":false,"fullname":"Charles Lopez","user":"charleslopez81","type":"user"},{"_id":"6a14692f2a9759cfbdfa9fc6","avatarUrl":"/avatars/cf44586b2b90564c128edcb6fb8cc1d0.svg","isPro":false,"fullname":"Yu Ziyi","user":"yuziyimh","type":"user"},{"_id":"6a15e69a8642d018701cf529","avatarUrl":"/avatars/d4239043b35c3c8552b4dce80d19084b.svg","isPro":false,"fullname":"Thomas ROBINSON","user":"thomasmk30","type":"user"},{"_id":"6a146d3b486a5aab39d36238","avatarUrl":"/avatars/4dc1f205e97b39a2951be5300a9f3f37.svg","isPro":false,"fullname":"Oliver Lopez","user":"oliver-lopez","type":"user"},{"_id":"6a14682131430965c63f1bc3","avatarUrl":"/avatars/06ff73863b78bc69f313c8d2015ff241.svg","isPro":false,"fullname":"Hu Linxi","user":"hulinxi","type":"user"},{"_id":"6a1470b7b28ec6a2ad92193c","avatarUrl":"/avatars/6243b0a5d5d3c8bec740d07bdbb12950.svg","isPro":false,"fullname":"Grayson Scott","user":"grascott99","type":"user"},{"_id":"6a146c2131430965c63f5ecb","avatarUrl":"/avatars/28e5680cbea1978e4cc6a005614e66a6.svg","isPro":false,"fullname":"Lin Wenxuan","user":"linwenxuan6","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"64b74d6c9f5987572c9be2b7","name":"TianjinUniversity","fullname":"Tianjin University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b74c979ebb7e6c7dd3c59f/OWmVv_2a2SIsp1MjkERFn.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.24479.md","query":{}}">
Papers
arxiv:2608.24479

WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

Published on Aug 25
· Submitted by
Yifu Yuan
on Aug 27
Authors:
,

Abstract

Off-policy reinforcement learning stabilizers vary with data availability, motivating regime-aware algorithms that adapt normalization and Q-function clipping to improve efficiency across CPU and GPU-parallel training.

Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for data-limited replay. Through controlled experiments across eight benchmark families, we show that these stabilizers are data-regime-dependent: parameter normalization helps with narrow replay coverage but restricts value fitting when data are abundant, while clipped double-Q can be relaxed in high-throughput manipulation. Age-biased replay weighting improves learning efficiency across regimes, especially with limited network capacity. Based on these findings, we propose WarpSAC, a regime-aware family of off-policy RL algorithms. WarpSAC uses Sample Weight Decay for efficient exploitation and provides two variants: WarpSAC-L (Norm ON, clipped double-Q) for data-limited CPU-scale training, and WarpSAC-A (Norm OFF, single-Q) for data-abundant GPU-parallel training. WarpSAC improves normalized score--step AUC over FlashSAC by 4.5% across nine CPU-scale environments and 23.1% across fourteen GPU-parallel environments. It increases UnitreeG1TransportBox-v1 success rate from 19.8% to 96.4%, improves mean normalized wall-time AUC on MuJoCo Playground by 19.1%, and achieves 36.4% faster sim-to-real deployment on Unitree G1 than FlashSAC. These results show that scalable off-policy RL should adapt its stabilizers to the available data regime.

Community

This comment has been hidden (marked as Resolved)
Paper submitter about 5 hours ago

The project page provides an overview of WarpSAC, demonstrations, benchmark results, and sim-to-real evaluations. The GitHub repository contains the source code, installation instructions, training scripts, environment integrations, and experiment configurations.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.24479
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.24479 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.24479 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.24479 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers