Serve DeepSeek-V4-Flash on your gaming PC</p>\n","updatedAt":"2026-08-19T03:40:57.891Z","author":{"_id":"642b970ceb31218a5f204a29","avatarUrl":"/avatars/582287f477bbb1a0842787145e375fd3.svg","fullname":"andy-yang","name":"andy-yang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6347628235816956},"editors":["andy-yang"],"editorAvatarUrls":["/avatars/582287f477bbb1a0842787145e375fd3.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.16157","authors":[{"_id":"6a83ccb0675db694db8cd51b","name":"Shuo Yang","hidden":false},{"_id":"6a83ccb0675db694db8cd51c","name":"Xiaoze Fan","hidden":false},{"_id":"6a83ccb0675db694db8cd51d","name":"Melissa Pan","hidden":false},{"_id":"6a83ccb0675db694db8cd51e","name":"Haocheng Xi","hidden":false},{"_id":"6a83ccb0675db694db8cd51f","name":"Zhe Wang","hidden":false},{"_id":"6a83ccb0675db694db8cd520","name":"Shanlin Sun","hidden":false},{"_id":"6a83ccb0675db694db8cd521","name":"Kurt Keutzer","hidden":false},{"_id":"6a83ccb0675db694db8cd522","name":"Song Han","hidden":false},{"_id":"6a83ccb0675db694db8cd523","name":"Matei Zaharia","hidden":false},{"_id":"6a83ccb0675db694db8cd524","name":"Chenfeng Xu","hidden":false},{"_id":"6a83ccb0675db694db8cd525","name":"Ion Stoica","hidden":false}],"publishedAt":"2026-08-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-19T00:00:00.000Z","title":"FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution","submittedOnDailyBy":{"_id":"642b970ceb31218a5f204a29","avatarUrl":"/avatars/582287f477bbb1a0842787145e375fd3.svg","isPro":false,"fullname":"andy-yang","user":"andy-yang","type":"user","name":"andy-yang"},"summary":"Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure. We present FreeToken, an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform. FreeToken co-designs the full serving stack, including model layout and loading, expert residency, CPU--GPU execution, agentic state reuse, and runtime memory management, around two realities of local AI: agent workloads continuously change their execution pattern, and edge hardware exposes heterogeneous resources whose balance differs from machine to machine. Rather than committing to a fixed offloading strategy, FreeToken continuously maps computation and model state onto the resources actually available. FreeToken supports more than 20 MoE models and real coding and tool-using agents across hardware ranging from an 8GB laptop GPU to a single workstation GPU. More importantly, it changes what these machines can practically serve, from a 35B model on a laptop to a 284B model on a gaming desktop and the 753B GLM-5.2 on a single workstation GPU. FreeToken turns open weights into deployable local software, making the machines users already own a practical platform for frontier-scale intelligence. We release the system at flashml.ai.","upvotes":37,"discussionId":"6a83ccb0675db694db8cd526","projectPage":"https://www.flashml.ai/","githubRepo":"https://github.com/FlashML-org/FreeToken","githubRepoAddedBy":"user","ai_summary":"FreeToken is an edge-native Mixture-of-Experts serving system that dynamically maps computation and model state onto heterogeneous local hardware to run large open-weight models on personal machines.","ai_keywords":["MoE serving","expert residency","CPU-GPU execution","agentic state reuse","runtime memory management","offloading strategy","open-weight models"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":18,"organization":{"_id":"66b1baeff10262fc4fa61961","name":"UCBerkeley","fullname":"University of California, Berkeley","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f425c3a096536aeab42dea/bxNKEkprdm5JI1wkjmNAL.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"670d34451de84e5d8870570e","avatarUrl":"/avatars/21f6b758f1a6f60f8bcc4a8ce9cdb96e.svg","isPro":false,"fullname":"jasonfan","user":"jasonfxz","type":"user"},{"_id":"654ca71d5255ee86711b52c5","avatarUrl":"/avatars/52bf00fd74c8db5643c4daa185c678e6.svg","isPro":false,"fullname":"Chenfeng Xu","user":"chenfengx","type":"user"},{"_id":"642b970ceb31218a5f204a29","avatarUrl":"/avatars/582287f477bbb1a0842787145e375fd3.svg","isPro":false,"fullname":"andy-yang","user":"andy-yang","type":"user"},{"_id":"66ff8aa43a31c499dc48fdd6","avatarUrl":"/avatars/060dc90fb13991bd013ce8173f12ae3e.svg","isPro":false,"fullname":"Jiarong Xing","user":"JerryPotter","type":"user"},{"_id":"64ebbae6895a36ab28de811a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ebbae6895a36ab28de811a/gBiaQP4paS4L13eu-yRm7.jpeg","isPro":false,"fullname":"Shiyi Cao","user":"eva98","type":"user"},{"_id":"67e1e67dddee0d80c39597ff","avatarUrl":"/avatars/b7554fe58510a3c0d11ccc7156210169.svg","isPro":false,"fullname":"Fangzhou","user":"Fzz1","type":"user"},{"_id":"677da7a639aa9532a65cac96","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/677da7a639aa9532a65cac96/pGFj13iTWurEhhLd_Jfor.jpeg","isPro":false,"fullname":"NovaSky","user":"NovaSkyAI","type":"user"},{"_id":"69446b15835f00df604cbc7a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/203QH9qkSLhigPnPA2SZW.webp","isPro":false,"fullname":"Qiuyang Mang","user":"qmang","type":"user"},{"_id":"63f30d28f4e30ffd2bda9aff","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676872959244-noauth.jpeg","isPro":false,"fullname":"Jiaming Tang","user":"Sakits","type":"user"},{"_id":"65b04425c11502dcf173d716","avatarUrl":"/avatars/f0a3c41059b75f569a419536e0372e66.svg","isPro":false,"fullname":"Hongyi Liu","user":"aladinggit","type":"user"},{"_id":"68145099418194b36493e11c","avatarUrl":"/avatars/c95e2e424c03b01b51dcf61eec51acf4.svg","isPro":false,"fullname":"Shubham Agarwal","user":"f20180301","type":"user"},{"_id":"6a83e39d289f02397d8fad3a","avatarUrl":"/avatars/df9e8296f5024c4ec72e74afa19751b0.svg","isPro":false,"fullname":"William Zheng","user":"wxzhengberkeley","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66b1baeff10262fc4fa61961","name":"UCBerkeley","fullname":"University of California, Berkeley","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f425c3a096536aeab42dea/bxNKEkprdm5JI1wkjmNAL.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.16157.md","query":{}}">
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
Abstract
FreeToken is an edge-native Mixture-of-Experts serving system that dynamically maps computation and model state onto heterogeneous local hardware to run large open-weight models on personal machines.
Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure. We present FreeToken, an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform. FreeToken co-designs the full serving stack, including model layout and loading, expert residency, CPU--GPU execution, agentic state reuse, and runtime memory management, around two realities of local AI: agent workloads continuously change their execution pattern, and edge hardware exposes heterogeneous resources whose balance differs from machine to machine. Rather than committing to a fixed offloading strategy, FreeToken continuously maps computation and model state onto the resources actually available. FreeToken supports more than 20 MoE models and real coding and tool-using agents across hardware ranging from an 8GB laptop GPU to a single workstation GPU. More importantly, it changes what these machines can practically serve, from a 35B model on a laptop to a 284B model on a gaming desktop and the 753B GLM-5.2 on a single workstation GPU. FreeToken turns open weights into deployable local software, making the machines users already own a practical platform for frontier-scale intelligence. We release the system at flashml.ai.
Community
Serve DeepSeek-V4-Flash on your gaming PC
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.16157 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.16157 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.16157 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.