How many tokens/second output are you getting with Qwen3.8-27B?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Trying to get a feel for where I stand. If you can list your relevant hardware and model used, that would be awesome.
Here's mine:
Model: Qwen3.8-27B-heretic-ara, Q5_K_M GGUF
T/s: ~30-32 t/sec (I think, I'll verify in a bit)
Hardware: 3090 GPU | 64 GBs DDR4 RAM | AMD 7950x CPU
Harness: Pi
Inference: llama.ccp
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.