DeepSeek V4 Flash on an M2 Ultra: repacked to 141 GiB losslessly, smaller than the Q4 GGUF, at 25.8 t/s (42 t/s peak)
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
This is one more vibe slopped custom optimization for, in this case, my hardware (m2 ultra 60 cores, 192gb). It is just a fork from llama.cpp with a few changes, it achieves:
- DeepSeek V4 Flash, no kv cache quant
- 141GiB model, byte-identical lossless, smaller than the public GGUFs (more room for context!)
- Faster than even the M3 Ultra (16 t/s vs 25 t/s)
- SSD KV cache and dynamic lanes, 1M context total, 8 lanes
- PP is a bit low at ~350 t/s at 8k-32k, but SSD cache compensates for it a lot... but we could probably push this number higher, lot of compute being left on the table
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.