DeepSeek-V4-Flash-0731 UD-Q8_K_XL 17.20~ t/s on A6000 + 256GB DDR4
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Hello everyone I want to join the hype of posting specs.
CPU: AMD EPYC 74F3 24-Core
RAM: 8 Channel 3200 DDR4
GPU: RTX A6000 48GB
Prompt processing is in the high 70t/s (got down to mid 30t/s at 300k context). Inference is a steady 17.20t/s~ and the 48GB VRAM is enough to have the full 1mil context but PP will be so bad. Sadly not as cool like those M5 Macs.
Anyone else having similar specs?
Edit: I was informed about batch size and set mine to 8096 and my Prompt processing jumped to almost 400t/s at the start. it got to around 300t/s at 20k context. Better than my 70t/s stock lol.
[link] [comments]
More from r/LocalLLaMA
-
Demo of local document extraction (52 pages) using Arctic Embed and Bonsai 8B on an Iphone 16 (KernelAI app)
Aug 30
-
Will apple still release devices with mobile HbM in 2027 ?
Aug 30
-
Whatever happened to OpenClaw and its derivatives?
Aug 30
-
Qwen 3.8 Flash Next locally on simple mobile phone at 3.5 tok/s
Aug 30
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.