r/LocalLLaMA · · 1 min read

Qwen3.8 Flash Quants

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

~20–30GB smaller than Unsloth/AesSedai Q4 at similar PPL.

After several days of testing I released a set of mainline-compatible imatrix quants for Qwen3.8-Flash-Next.

Goal: same quality band as the popular Unsloth / AesSedai Q4 builds, less disk and RAM. Savings are roughly 20–30GB depending on the file you compare against. PPL is in the model card and is competitive with both.

AMD / Strix Halo: separate ROCmFP4 build that is a bit better and faster than the Q4_XS on that hardware. (https://huggingface.co/agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-imatrix-GGUF)

If you try it, post your quant, RAM/VRAM, tok/s, and whether quality felt on par with Unsloth IQ4_XS / Q4_K. That is the comparison I care about.

submitted by /u/Dutchnamn
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA