Today I hit 181 toks/s (aggregate) on Qwen3.8-Flash-Next on 2x DGX Sparks
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Hey all, and hello fellow DGX Spark-ers! Today I managed some pretty crazy numbers: 181 tok/s aggregate on 2× DGX Spark on Qwen3.8-Flash-Next at 512Kcontext (2.8M kvc) I hit 181 tok/s aggregate today across a multi-agent fleet on a 2-node DGX Spark cluster. Single-stream decode is 30–50 tok/s — the 181 is total throughput with ~9 concurrent agent sessions sharing the engine. I actually peaked to 195 while writing this. Quick rundown of how it's served: Hardware
Model
The trick: PLE table on NVMe
vLLM config (official day-0 image, vllm/vllm-openai)
Serving stack
Happy to answer questions about any of it. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.