Qwen 3.6 27B absolutely fails at agentic work
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I have been running Qwen 3.5 122B at 4 bit for quite a while, and have started running it at 5 bit recently now that Llama.cpp has comparable performance to VLLM.
I have also tried, several times, to use Qwen 3.6 27B at 8 bit & 16 bit, as numerous people have claimed that 27B is better than 122B.
And it is, on single prompts. It will output very impressive demo HTML pages. It has the ability to generate much longer content than any of the 3.5 series models.
However, on agentic work, it absolutely falls apart. It makes mistakes continuously and does not follow directions. I cannot get the model to not screw up. Every 4 turns or so it does something completely braindead.
Am I the only one who has noticed this? I am back to using 122B again after trying, yet again, to make 27B work.
Llama.cpp, nightly compiled from Git, on RTX 6000
[link] [comments]
More from r/LocalLLaMA
-
Unpopular opinion Qwen 3.8 is hard to understand
Aug 30
-
Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM
Aug 30
-
Oh so that's where my PCIe lanes went...
Aug 30
-
Uncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and…
Aug 30
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.