Qwen3.8-27B on 2x 3090 + vLLM + DFlash2: 218 tok/s single request
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| I hacked this together so there's probably more on the table in terms of performance. Measured with the Club-3090 canonical bench suite (bench.sh, 3 warmups + 5 measured runs, temp 0.6 / top_p 0.95 / top_k 20).
Stack
[link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.