Fastest NVFP4 quant of Qwen3.8 27B out there
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Here's a brand new Blackwell-native, prefill-optimized 4-bit quant that runs 50% faster on compatible hardware than a Q4 quant of the same memory footprint. And it runs 4-7% faster than other NVFP4 quants as benchmarked on RTX 5090 32GB.
This GGUF also includes a quantized MTP draft head for a good measure. Check it out for all details and specifically recommended settings for 15% faster MTP. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.