Unpopular opinion : Qwen 3.8 27b is not an overthinker
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Yes it uses a ton more reasoning tokens than 3.6 did
But test in on the same tasks with the other chinese models, glm 5.3, deepseek v4 flash and pro, etc it's really similar, and they are needed
The reality is, we're just frustrated because our hardware do not allow most of us to have 1M context (I know that it's not supported yet) with 150 tps decode
Furthermore, if you don't mind the quality drop, you can just add a reasoning budget, it will still be better than 3.6
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.