r/LocalLLaMA · · 1 min read

Fastest NVFP4 quant of Qwen3.8 27B out there

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Fastest NVFP4 quant of Qwen3.8 27B out there

Here's a brand new Blackwell-native, prefill-optimized 4-bit quant that runs 50% faster on compatible hardware than a Q4 quant of the same memory footprint.

And it runs 4-7% faster than other NVFP4 quants as benchmarked on RTX 5090 32GB.

Quant Benchmark Speed
NVFP4 pp2048 6250 t/s
unsloth NVFP4 pp2048 6010 t/s
Q4_0 pp2048 4130 t/s
Q6_K pp2048 3210 t/s

This GGUF also includes a quantized MTP draft head for a good measure.

Check it out for all details and specifically recommended settings for 15% faster MTP.

submitted by /u/ionsago
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA