r/LocalLLaMA · · 1 min read

Your own GGUF

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Hello, I have a few questions that I can't seem to find a clear answer to.

Does it make sense to make your own GGUF?

I noticed that when I compile llamacpp (vulkan or rocm), the processing and generation is a bit better, does it work similarly with doing GGUF yourself?

If I use Vulkan, is it worth doing GGUF using llama-quantize vulkan version (not rocm version)?

To what extent does it make sense to place certain model elements at higher precision (conversation, document analysis)?

I use gemma 4 31B the most.

submitted by /u/Daniokenon
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA