r/LocalLLaMA · · 1 min read

Over 200k context on 16GB VRAM with Qwen 3.8 27B UD-IQ3_XXS

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Over 200k context on 16GB VRAM with Qwen 3.8 27B UD-IQ3_XXS

I was using UD-Q3_K_XL until now with more than 140000 context. Quality wise it's very good, very few erroneous tool calls. Then I saw many others here reporting good results with IQ3_XXS, so I gave it a try.

The downside is prompt processing speed went down from 700-800 tk/s to 400 tk/s. Quality difference is yet to be tested.

KV cache were both quantized to q5_1

(llama.CPP compiled with DGGML_CUDA_FA_ALL_QUANTS=ON)

Served without MTP and mmproj.

My setup is a measly laptop with TB4 and Aorus 5060ti AI Box eGPU.

Windows 11, cuz Nvidia.

Apologies for any mistake in the post.

submitted by /u/abskvrm
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA