AtomicChat/Qwen3.8-Flash-Next-GGUF is Really Good
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| specs
AtomicChat/Qwen3.8-Flash-Next-GGUF
u/erikdhoward suggested to try the Atomic Chat quant which I did not know anything about. I tried it, and it is... really good. AtomicChat's quant uses llama.cpp mmap (through GGUF shard layout vs. in the runtime) and keeps the PLE table (n-grams) pageable backed by a file. Because of this the same model that took 106GB, now takes 65GB (starts from 55GB) in RAM. And since PLE is pageable the prefill is actually not that bad, cold start is about 500 t/s. oMLX "right behind you!"
[link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.