I added MTP to local SoTA Agentic Coding Model Ornith 35B FP8 E4M3
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Just wanted to share that I was looking for an optimal way to run Ornith 35B in FP8 with E4M3 and MTP with vLLM but there was no out-of-the-box model with MTP drafter support. So I grafted this new model! It's 18% faster than without MTP and the drafter acceptance rate is not bad (70% on avg). It should run on any RTX based setup > 80GB VRAM with full context window 256k. Might also do well on Unified Memory Systems like GB10 (for this - use my script and graft the MTP model into a target NVFP4 model!). I work with Hopper and Ada gen hardware, so this is the Pareto optimal for me. Have fun! Grafter script and vLLM high performance inference container: https://github.com/kyr0/Ornith-35B-FP8-E4M3-MTP [link] [comments] |
More from r/LocalLLaMA
-
Unpopular opinion Qwen 3.8 is hard to understand
Aug 30
-
Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM
Aug 30
-
Oh so that's where my PCIe lanes went...
Aug 30
-
Uncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and…
Aug 30
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.