r/LocalLLaMA · · 1 min read

[Open PR] llama : add --n-cpu-ffn option by John-194 · Pull Request #26622 · ggml-org/llama.cpp

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

[Open PR] llama : add --n-cpu-ffn option by John-194 · Pull Request #26622 · ggml-org/llama.cpp

PR by u/Stainless-Bacon 👍

It would be handy & awesome to have options --n-cpu-ffn / --cpu-ffn for Dense models like how we have --n-cpu-moe / --cpu-moe for MOE models.

Also check his threads:

(

Awesome to see the big comment by u/Pablo_the_brave there, filled with so much stuff. Quoting a line from there

Do not touch block 64 (MTP) if you are using speculative decoding — its FFN should remain on the GPU.

)

We should've got this option long time back actually. This PR instantly reminded me of last year thread. (I literally used his -ot command for sometime with Qwen3-14B. I'm just happy that I was able to recall a last year thread.)

Anyway .... Better late than never. Waiting for this merge.

submitted by /u/pmttyji
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA