Hi, I really liked this paper and all the ablation studies you did, congrats!<br>Do you plan to share the repo you used to train this model and the checkpoints? I wanted to check how does EBT (Energy based transformer) compared with the RecurrentGPT.</p>\n","updatedAt":"2026-08-25T13:44:27.996Z","author":{"_id":"60eeedbf50b60c406afc1291","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1649111275459-60eeedbf50b60c406afc1291.png","fullname":"Samuel Arcadinho","name":"SSamDav","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9572992324829102},"editors":["SSamDav"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1649111275459-60eeedbf50b60c406afc1291.png"],"reactions":[{"reaction":"❤️","users":["melhoushi"],"count":1}],"isReport":false},"replies":[{"id":"6a8ea464caad9959079c91db","author":{"_id":"64ea6965b96ff0e1753c128a","avatarUrl":"/avatars/4763d3e76b10cfb82e09fb0c26875efb.svg","fullname":"Amr Hegazy","name":"Amr-Hegazy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false},"createdAt":"2026-08-26T08:31:32.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Hi, Thank you for your interest in our paper! We actually just released the training code: https://github.com/Amr-Hegazy1/gated-recurrent-transformer","html":"<p>Hi, Thank you for your interest in our paper! We actually just released the training code: <a href=\"https://github.com/Amr-Hegazy1/gated-recurrent-transformer\" rel=\"nofollow\">https://github.com/Amr-Hegazy1/gated-recurrent-transformer</a></p>\n","updatedAt":"2026-08-26T08:31:32.083Z","author":{"_id":"64ea6965b96ff0e1753c128a","avatarUrl":"/avatars/4763d3e76b10cfb82e09fb0c26875efb.svg","fullname":"Amr Hegazy","name":"Amr-Hegazy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8800825476646423},"editors":["Amr-Hegazy"],"editorAvatarUrls":["/avatars/4763d3e76b10cfb82e09fb0c26875efb.svg"],"reactions":[],"isReport":false,"parentCommentId":"6a8d9c3b2f645e17210a7acd"}}]},{"id":"6a8ea4501063d9f0feb9d857","author":{"_id":"64ea6965b96ff0e1753c128a","avatarUrl":"/avatars/4763d3e76b10cfb82e09fb0c26875efb.svg","fullname":"Amr Hegazy","name":"Amr-Hegazy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false},"createdAt":"2026-08-26T08:31:12.000Z","type":"comment","data":{"edited":true,"hidden":true,"hiddenBy":"","hiddenReason":"Off-Topic","latest":{"raw":"This comment has been hidden","html":"This comment has been hidden","updatedAt":"2026-08-26T08:31:46.632Z","author":{"_id":"64ea6965b96ff0e1753c128a","avatarUrl":"/avatars/4763d3e76b10cfb82e09fb0c26875efb.svg","fullname":"Amr Hegazy","name":"Amr-Hegazy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"editors":[],"editorAvatarUrls":[],"reactions":[]}},{"id":"6a8fbb0c0e2e23c320365ccd","author":{"_id":"64ea6965b96ff0e1753c128a","avatarUrl":"/avatars/4763d3e76b10cfb82e09fb0c26875efb.svg","fullname":"Amr Hegazy","name":"Amr-Hegazy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false},"createdAt":"2026-08-27T04:20:28.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Can a transformer gain expressive depth without adding unique layers? We introduce **Gated Recurrent Transformers**, which repeatedly apply a small shared core while using a lightweight elementwise update gate to modulate how representations evolve across recurrences. This lets the same parameters specialize across computation steps rather than performing an identical transformation each time. Under isoFLOPs, a 3-layer model matches a 12-layer GPT-2 Small baseline, while at larger scale we retain competitive quality with substantially fewer parameters and lower decoding memory.","html":"<p>Can a transformer gain expressive depth without adding unique layers? We introduce <strong>Gated Recurrent Transformers</strong>, which repeatedly apply a small shared core while using a lightweight elementwise update gate to modulate how representations evolve across recurrences. This lets the same parameters specialize across computation steps rather than performing an identical transformation each time. Under isoFLOPs, a 3-layer model matches a 12-layer GPT-2 Small baseline, while at larger scale we retain competitive quality with substantially fewer parameters and lower decoding memory.</p>\n","updatedAt":"2026-08-27T04:20:28.615Z","author":{"_id":"64ea6965b96ff0e1753c128a","avatarUrl":"/avatars/4763d3e76b10cfb82e09fb0c26875efb.svg","fullname":"Amr Hegazy","name":"Amr-Hegazy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8243993520736694},"editors":["Amr-Hegazy"],"editorAvatarUrls":["/avatars/4763d3e76b10cfb82e09fb0c26875efb.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.15062","authors":[{"_id":"6a88ea503d26296ea30913d1","user":{"_id":"64ea6965b96ff0e1753c128a","avatarUrl":"/avatars/4763d3e76b10cfb82e09fb0c26875efb.svg","isPro":false,"fullname":"Amr Hegazy","user":"Amr-Hegazy","type":"user","name":"Amr-Hegazy"},"name":"Amr Hegazy","status":"claimed_verified","statusLastChangedAt":"2026-08-22T00:45:04.266Z","hidden":false},{"_id":"6a88ea503d26296ea30913d2","name":"Amr Alanwar","hidden":false},{"_id":"6a88ea503d26296ea30913d3","user":{"_id":"63c9725ebedad7e2bf160bdc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c9725ebedad7e2bf160bdc/wzPuyhOXCYBNGwZDshbnL.jpeg","isPro":false,"fullname":"Mostafa Elhoushi","user":"melhoushi","type":"user","name":"melhoushi"},"name":"Mostafa Elhoushi","status":"claimed_verified","statusLastChangedAt":"2026-08-27T00:45:04.180Z","hidden":false}],"publishedAt":"2026-08-25T00:00:00.000Z","submittedOnDailyAt":"2026-08-27T00:00:00.000Z","title":"Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation in Transformers","submittedOnDailyBy":{"_id":"64ea6965b96ff0e1753c128a","avatarUrl":"/avatars/4763d3e76b10cfb82e09fb0c26875efb.svg","isPro":false,"fullname":"Amr Hegazy","user":"Amr-Hegazy","type":"user","name":"Amr-Hegazy"},"summary":"Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While unique weights across layers preserve functional specialization---from input-grounding to abstract refinement---they incur a substantial memory footprint. Conversely, standard depth-sharing enforces uniform transformations that collapse representational diversity and degrade modeling quality. We introduce Gated Recurrent Transformer, a recurrent depth transformer where fixed-depth prelude and coda blocks bracket a single shared core iterated R times. Inspired by gated recurrent neural networks, we employ a lightweight projection and an elementwise update gate---conditioned on the hidden state, the fixed prelude output, and noise resampled at every step---to modulate the recurrent update. This allows the model to specialize the input to the same few layers across recurrences, rather than requiring many unique layers to achieve functional diversity. Under an isoFLOPS constraint, a 3-layer Gated Recurrent Transformer matches the accuracy of a 12-layer GPT-2 Small baseline with similar training and inference FLOPs, and leads MoR and heavy-tail depth sampling in all nine scale-by-budget cells; at medium and large scale it approaches dense quality at the standard token budget and overtakes it at medium scale once that budget is doubled. Under an isoPARAMS constraint, deeper recurrence achieves a 2.76 validation loss versus 2.84 for a non-recurrent counterpart at matched parameter and data budget. Our results demonstrate that adaptive depth reuse is a principled strategy for trading parameters for quality: at large scale, 63% fewer parameters and 59% less peak decoding memory for a 10% increase in compiled generation latency.","upvotes":3,"discussionId":"6a88ea503d26296ea30913d4","githubRepo":"https://github.com/Amr-Hegazy1/gated-recurrent-transformer","githubRepoAddedBy":"user","ai_summary":"A gated recurrent transformer reuses a shared core across depth with adaptive update gates, achieving comparable or better quality than deeper models with far fewer parameters and lower memory.","ai_keywords":["Gated Recurrent Transformer","recurrent depth transformer","depth-sharing","update gate","isoFLOPS","isoPARAMS","adaptive depth reuse"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"61fae781e68759322b9767be","name":"TUM","fullname":"Technical University of Munich","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1661167219960-629521a0f937190946e15d7f.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"60eeedbf50b60c406afc1291","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1649111275459-60eeedbf50b60c406afc1291.png","isPro":false,"fullname":"Samuel Arcadinho","user":"SSamDav","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"68b2a4157f881fc640ba7d80","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/lMTgr3pe7pOHtMe7bVF7F.png","isPro":false,"fullname":"khtsly","user":"khtsly","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61fae781e68759322b9767be","name":"TUM","fullname":"Technical University of Munich","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1661167219960-629521a0f937190946e15d7f.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.15062.md","query":{}}">
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation in Transformers
Abstract
A gated recurrent transformer reuses a shared core across depth with adaptive update gates, achieving comparable or better quality than deeper models with far fewer parameters and lower memory.
Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While unique weights across layers preserve functional specialization---from input-grounding to abstract refinement---they incur a substantial memory footprint. Conversely, standard depth-sharing enforces uniform transformations that collapse representational diversity and degrade modeling quality. We introduce Gated Recurrent Transformer, a recurrent depth transformer where fixed-depth prelude and coda blocks bracket a single shared core iterated R times. Inspired by gated recurrent neural networks, we employ a lightweight projection and an elementwise update gate---conditioned on the hidden state, the fixed prelude output, and noise resampled at every step---to modulate the recurrent update. This allows the model to specialize the input to the same few layers across recurrences, rather than requiring many unique layers to achieve functional diversity. Under an isoFLOPS constraint, a 3-layer Gated Recurrent Transformer matches the accuracy of a 12-layer GPT-2 Small baseline with similar training and inference FLOPs, and leads MoR and heavy-tail depth sampling in all nine scale-by-budget cells; at medium and large scale it approaches dense quality at the standard token budget and overtakes it at medium scale once that budget is doubled. Under an isoPARAMS constraint, deeper recurrence achieves a 2.76 validation loss versus 2.84 for a non-recurrent counterpart at matched parameter and data budget. Our results demonstrate that adaptive depth reuse is a principled strategy for trading parameters for quality: at large scale, 63% fewer parameters and 59% less peak decoding memory for a 10% increase in compiled generation latency.
Community
Hi, I really liked this paper and all the ablation studies you did, congrats!
Do you plan to share the repo you used to train this model and the checkpoints? I wanted to check how does EBT (Energy based transformer) compared with the RecurrentGPT.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.