Qwen3.8-27B Q6_K at 128K on a single 32GB GPU
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Fresh download today. Really quick: my impression of Qwen3.8-27B: I ran the Q6_K GGUF locally on a 32GB R9700 through llama.cpp + OpenCode, with 128K context. Q8 would not fit with enough left over for context. I gave it my real, years-old swimming pool-controller repository and asked for a simple reliability audit. It spent about 21 minutes tracing the architecture, startup/cron behavior, serial input stream handling, GPIO control, CGI functions, Git history, and lots of years of logs. It used about 52K of context to produce a 30KB report that, overall, showed that it grasped the project well.
Too early to claim it’s better than 3.6-27B Q8, which fit better on 32GB VRAM than this 3.8 at Q8 (weights took up 91% at 32K context, and 97% at 64K context—ouch)…the harness, prompt, and available context were better controlled this time. It also made at least one wrong inference where my own operator use mattered. That was a quirk in the way I use the system, which it couldn’t have known anyway…it gets a pass. What impressed me was the way the agentic behavior held up making sense of my hot mess of really old relay control code. 3.6 27B Q8 did almost as well; Devstral and some others…let’s just say it was ugly. This was straightforward.
So far, so this was just a quick look at Qwen3.8-27B Q6_K at 128K on a single 32GB card system, llama.cpp/Vulkan. Very nice at this point
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.