r/LocalLLaMA · · 1 min read

SGLang support for Qwen3.8-27B: 200+ tok/s on 5090, 38 tok/s on DGX Spark (NVFP4 + DSpark)

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

SGLang support for Qwen3.8-27B: 200+ tok/s on 5090, 38 tok/s on DGX Spark (NVFP4 + DSpark)

Hey r/LocalLLaMA 👋

This is Kai from SGLang.

We just shipped day-0 support for Qwen3.8-27B.

To push performance for running this model locally, we combined NVFP4 + DSpark and got:

  • 200+ tok/s decode on a single RTX 5090 and RTX Pro 6000
  • 38 tok/s decode on DGX Spark

Here's the cookbook: https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8-27B

We're committed to making SGLang great for local AI 🫡

Would love any feedback and thoughts on what we could do better to grow with the local AI community.

NVFP4 checkpoint: huggingface.co/RadixArk/Qwen3.8-27B-NVFP4
DSpark checkpoint: huggingface.co/RadixArk/Qwen3.8-27B-DSpark
Our x post: x.com/sgl_project/status/2088281320422322413

submitted by /u/unseenmarscai
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA