Qwen 3.6 27B MTP speed on 3080ti (getting 4.5 t/s)
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Using LM Studio with 3080ti (12gb of VRAM) and 128gb of ddr4.
Model version: Qwen 3.6 27B MTP UD q4_k_xl
Is this my hardware limit?
Is there anyway to speed this up using the current hardware?
[link] [comments]
More from r/LocalLLaMA
-
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp · Hugging Face
Aug 31
-
pipecat-ai/phonellm-alpha-1: GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost
Aug 31
-
Could this affect M5 Ultra price/availability?
Aug 31
-
GLM 5.3, GLM 5.3 Flash or 3.8 Qwen Flash for Replacing Kimi k3 IQ2_XXS
Aug 31
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.