AMD Users: Have you tried the llamma.cpp AMD-Ecosystem branch? Up to 2x PP Speed
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
AMD has it's own llama.cpp branch: https://github.com/AMD-Ecosystem/llama.cpp
And despite the Deprecation warning it's actively maintained (things are later upstreamed to the normal llama.cpp).
What i noticed with my Strix Halo:
It has some interesting new patches (if you use ROCm/Hip)
The Prompt Processing speed with dense model is sometimes over 2 times faster ! I get around 550 tokens/s with a 14B dense compared to 230 with the normal llama.cpp. However TG is around 15% slower than with Vulkan.
MoE speed is the same.
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.