Qwen3.8 Flash Quants
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
~20–30GB smaller than Unsloth/AesSedai Q4 at similar PPL.
After several days of testing I released a set of mainline-compatible imatrix quants for Qwen3.8-Flash-Next.
Goal: same quality band as the popular Unsloth / AesSedai Q4 builds, less disk and RAM. Savings are roughly 20–30GB depending on the file you compare against. PPL is in the model card and is competitive with both.
- Repo: https://huggingface.co/agentionai/Qwen3.8-Flash-Next-AP-GGUF
- Q4 quants are the ones I would start with. Q3 and Q5 are coming.
- Recipe is per-layer / tailored, not a blanket lower bpw.
AMD / Strix Halo: separate ROCmFP4 build that is a bit better and faster than the Q4_XS on that hardware. (https://huggingface.co/agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF)
If you try it, post your quant, RAM/VRAM, tok/s, and whether quality felt on par with Unsloth IQ4_XS / Q4_K. That is the comparison I care about.
[link] [comments]
More from r/LocalLLaMA
-
an unscientific qwen 3.8 flash next and glm 5.3 flash comparison
Aug 30
-
Ran Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max — speed vs context depth, 100 turns, one graph
Aug 30
-
Nemotron-3.5-Lightning at 11.77 GiB, a 16 GB option for a model that didn't have one
Aug 29
-
Humaneval benchmark for Deepseek V4 Flash 0731 vs GLM5.3 Flash on 2x DGX Spark setup
Aug 29
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.