r/LocalLLaMA · · 1 min read

How many tokens/second output are you getting with Qwen3.8-27B?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Trying to get a feel for where I stand. If you can list your relevant hardware and model used, that would be awesome.

Here's mine:

Model: Qwen3.8-27B-heretic-ara, Q5_K_M GGUF

T/s: ~30-32 t/sec (I think, I'll verify in a bit)

Hardware: 3090 GPU | 64 GBs DDR4 RAM | AMD 7950x CPU

Harness: Pi

Inference: llama.ccp

submitted by /u/CooLittleFonzies
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA