Qwen3.8-Flash-Next optimised for Macs
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Running on a M1 Max 64 GB: How is it possible? * Custom Q4 quant: benchmarked all metal tensors then picked and spliced tensors from multiple Unsloth and AtomicChat quants to achieve best performance/bit. https://github.com/mihailescu2m/llama.cpp Special thanks to Claude - three weeks worth of tokens and some extra out of pocket usage credits made it all possible. Feedback appreciated. Note: enabling MTP uses more RAM, which means less cache for tensors, leading to prefill going from 180 tps to g170 tps (at 4K). For 256K context, more RAM is needed for KV cache, prefill goes down to 150 tps. But with MTP, decode gains +70%, going up to 22 btps. So if you need highest prefill, disable MTP. A [link] [comments] |
More from r/LocalLLaMA
-
Qwen3.8-Flash-Next NVFP4 Day-3 support for 4xV100
Aug 30
-
It's official! 192GB Framework
Aug 30
-
an unscientific qwen 3.8 flash next and glm 5.3 flash comparison
Aug 30
-
Ran Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max — speed vs context depth, 100 turns, one graph
Aug 30
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.