r/LocalLLaMA · · 1 min read

Qwen3.8-27B Q8_0 on Strix Halo is seriously impressive

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Qwen3.8-27B Q8_0 on Strix Halo is seriously impressive

Sorry for the slop, but I was impressed by this model as I have been testing Qwen3.8-27B Q8_0 locally on my ROG Flow Z13 (Ryzen AI Max+ 395, 128 GB unified memory) and this model was the only one who could made this short simulator (and I have tested a lot of models).

Prompt:
"Create a beautiful, relaxing flight simulator in a single HTML page."

It generated the whole thing through an agent using file/bash tools.

Setup:

Lemonade Server + llama.cpp ROCm

Q8_0 weights + Q8 KV cache

Native MTP speculative decoding

64 GB VRAM / 64 GB RAM split

~142k context

I'm seeing roughly 9-19 tok/s depending on the agent step, with some generations sustaining 16-19 tok/s and MTP acceptance reaching 97-99% and it took ~20min to generate this simulator.

submitted by /u/seti_at_home
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA