Has there been any recent new development on which quant is considered optimal?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I recall in earlier days, q4 was said to be optimal.
That is to say, if you have a:
small q8 model
medium q4 model
large q2
Assuming they use the same amount of GPU VRAM, medium q4 would be the best-performing model.
I also know that Apple (crazy that I am citing Apple here, given how secretive they tend to be) was quite public about using q4 quant models for thier on device.
[link] [comments]
More from r/LocalLLaMA
-
6x P40 running Minimax M2.7_Q3_XL
Jul 2
-
Fine-tuned Gemma-4-31B specifically for Copywriting & Creative Writing Tasks (Scored +290 Elo over base using EqBench3)
Jul 2
-
Gemma 4 WebGPU Kernels 255 tok/s by x/@xenovacom
Jul 2
-
openlumara, my manually coded super-token-efficient harness, now works across any UI that can connect to an openAI endpoint! koboldlite, openwebui, you name it. basically, openAI bridge. yay!
Jul 2
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.