SGLang support for Qwen3.8-27B: 200+ tok/s on 5090, 38 tok/s on DGX Spark (NVFP4 + DSpark)
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Hey r/LocalLLaMA 👋 This is Kai from SGLang. We just shipped day-0 support for Qwen3.8-27B. To push performance for running this model locally, we combined NVFP4 + DSpark and got:
Here's the cookbook: https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8-27B We're committed to making SGLang great for local AI 🫡 Would love any feedback and thoughts on what we could do better to grow with the local AI community. NVFP4 checkpoint: huggingface.co/RadixArk/Qwen3.8-27B-NVFP4 [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.