[Paper] ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning
Somebody please create MOE models of recent Dense models like Qwen3.8-27B, Muse-Glimmer-30B, etc., Thanks u/KSAM-The-Randomizer for sharing this on my old thread. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.