CUDA: extend MOE fusion to specdec, earlier MOE glu fusion and topk-router fusion were restricted to 1 token by ynankani · Pull Request #27621 · ggml-org/llama.cpp
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| I haven't had a chance to test it yet, but it looks very promising. It seems to speed up MTP for MoE models across different draft widths (especially greater than 1). Check the benchmarks. [link] [comments] |
More from r/LocalLLaMA
-
How bad do you think models like Qwen3.8-27B or GLM-5.3-Flash would be with H-Neurons disabled?
Aug 31
-
vote for the Qwen 3.8
Aug 31
-
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp · Hugging Face
Aug 31
-
pipecat-ai/phonellm-alpha-1: GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost
Aug 31
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.