Has anyone else found vLLM outputs noticeably worse than llama.cpp for the same model?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I'm wondering if anyone else has come across this.
I've tested the same model on llama.cpp and vLLM with similar settings and quantizations. The performance and concurrency in vLLM are much noticeably better, but sometimes the model feels less reliable.
Some things I've noticed:
* More mistakes with formatting and tool calls
* Forgetting context suddenly
* Sometimes acting like messages didn't exist
* Lower quality code even with similar parameters
I'm not trying to start a comparison. I just want to know if others have seen differences in quality between inference backends... Is it usually because of quantization, chat templates, parser problems or configuration errors.
What has your experience been, like?
[link] [comments]
More from r/LocalLLaMA
-
Qwen 3.8 Flash Next locally on simple mobile phone at 3.5 tok/s
Aug 30
-
Unpopular opinion Qwen 3.8 is hard to understand
Aug 30
-
Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM
Aug 30
-
Oh so that's where my PCIe lanes went...
Aug 30
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.