Deepseek-V4-Flash-0731 Dwarfstar on Mac
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Here is the prefill performance in an M2 Ultra with 192GB of RAM. For decode, at the following depth: 45k: 23.5 t/s 192k: 18 t/s That speed is maintained with 8k token output at those depths. [link] [comments] |
More from r/LocalLLaMA
-
Qwen 3.8 Flash Next locally on simple mobile phone at 3.5 tok/s
Aug 30
-
Unpopular opinion Qwen 3.8 is hard to understand
Aug 30
-
Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM
Aug 30
-
Oh so that's where my PCIe lanes went...
Aug 30
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.