GLM 5.3, GLM 5.3 Flash or 3.8 Qwen Flash for Replacing Kimi k3 IQ2_XXS
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Kimi seem fine at IQ2_xxs but is slow 4tks (passable) but it can drop to 2tks (well, not great) doesn't seem so viable. Would the new GLM(s), or Qwen Flash a good substitute, especially to have a better interactive experience while having high intelligence.
[link] [comments]
More from r/LocalLLaMA
-
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp · Hugging Face
Aug 31
-
pipecat-ai/phonellm-alpha-1: GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost
Aug 31
-
Could this affect M5 Ultra price/availability?
Aug 31
-
snkii/Sori-1B: Audio-Grounded LM Trained From Scratch (No Text-Only Pretraining)
Aug 31
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.