r/LocalLLaMA · · 1 min read

Qwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

https://developer.nvidia.com/blog/serve-qwen3-8-2-4t-a95b-a-2-4t-parameter-model-with-configurable-reasoning-on-nvidia-gb300-nvl72/

4k tokens per second per GPU of which there are 72. 350 tokens per second per user

"Without additional model tuning, the model achieves a throughput of over 4K tokens per second per GPU and over 350 tokens per second per user on NVIDIA GB300 NVL72 in FP8 precision on Day 0. Further optimizations, including NVFP4 precision, are expected to deliver enhanced performance gains over time. "

submitted by /u/RhubarbSimilar1683
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA