Over 200k context on 16GB VRAM with Qwen 3.8 27B UD-IQ3_XXS
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| I was using UD-Q3_K_XL until now with more than 140000 context. Quality wise it's very good, very few erroneous tool calls. Then I saw many others here reporting good results with IQ3_XXS, so I gave it a try. The downside is prompt processing speed went down from 700-800 tk/s to 400 tk/s. Quality difference is yet to be tested. KV cache were both quantized to q5_1 (llama.CPP compiled with DGGML_CUDA_FA_ALL_QUANTS=ON) Served without MTP and mmproj. My setup is a measly laptop with TB4 and Aorus 5060ti AI Box eGPU. Windows 11, cuz Nvidia. Apologies for any mistake in the post. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.