Ornith-1.5-35B-A3B on 8 GB VRAM: I think I've found my sweet spot
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| A few days ago I posted asking what people considered the best local model for an 8 GB VRAM GPU. At the time, my personal sweet spot was Qwen3.6-35B-A3B, for agentic coding with Pi.dev. Well… Thanks to the suggestions in that thread, I think I've found something even better. I've been testing Ornith-1.5-35B-A3B Q4_K_M, and on my system the results have been genuinely impressive. My setup:
After testing dozens of different models, architectures and quantizations, Ornith has currently become my model of choice for agentic coding, without much hesitation. The really interesting part is the combination of speed and actual results. Just tonight I gave it a fairly complex code-analysis project. It went through the codebase, performed the analysis and completed the task in a relatively short amount of time, averaging around 32 tok/s. And the final result? Honestly, I'd call it near flawless. That's a pretty significant improvement over the speed I was getting with Qwen3.6, but the bigger difference for me isn't even the raw generation speed. It's how effectively Ornith handles the whole agentic workflow. And all of this while maintaining a 128K context window. I've tested a lot of models at this point - different parameter counts, MoE models, dense models, quantizations, coding fine-tunes, etc. Of course, this is very much a "right now" statement. 😄 There will probably be another model release tomorrow that makes me eat these words. That's how quickly things are moving. But as of today, for my particular hardware and my particular use case, Ornith-1.5 is my clear winner. The combination of quality + agentic coding ability + context length + speed + relatively modest hardware requirements is just ridiculously good. I'm curious whether other people are getting similar results with Ornith, especially on 8 GB GPUs or other relatively constrained systems. If you have questions, feel free to ask. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.