r/LocalLLaMA · · 1 min read

We quantized the new Ornith 1.5 9B and 35B-A3B

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

We quantized the new Ornith 1.5 9B and 35B-A3B

ornith lab dropped new ornith 1.5 today, a 9B dense with vision and a 35B-A3B MoE, both MIT, trained on a loop that generates its own tasks. in addition there was giant 397b model, but we didn't quantize it (but if you want to try - we will do it)

we made our AD (Atomic Dynamic) quants for both, 9B (14 builds) and 35B-A3B (13 builds), and measured them against stock llama.cpp quants (on the same imatrix) their mean KLD and top-1 against our own BF16 conversion

Ornith-1.5-9B

file size mean KLD top-1
Q8_0 9.53 GB 0.0022 97.94%
AD-Q8_0-Q6_K 8.55 GB 0.0035 97.46%
Q5_K_M 6.47 GB 0.0299 92.80%
AD-Q5_K-Q4_K 5.93 GB 0.0255 93.10%
AD-Q4_K-IQ4_XS 5.61 GB 0.0344 91.93%
AD-IQ3_S-IQ3_XXS 4.29 GB 0.1441 83.44%

Ornith-1.5-35B-A3B

file size mean KLD top-1
Q6_K 28.51 GB 0.0167 94.63%
AD-Q6_K-Q5_K 26.25 GB 0.0158 94.85%
Q5_K_M 24.73 GB 0.0269 93.31%
AD-Q5_K-Q4_K 22.14 GB 0.0251 93.52%
Q4_K_M 21.17 GB 0.0477 91.01%
AD-Q4_K-IQ4_XS 20.13 GB 0.0315 92.71%

Collections on HF with the imatrix, the per-tensor layouts and everything else: https://huggingface.co/collections/AtomicChat/ornith-15-9b
https://huggingface.co/collections/AtomicChat/ornith-15-35b-a3b

Our local ai open source app https://atomic.chat (I'm cofounder). Feel free to ask any questions and share your feedback!

submitted by /u/Fun-Meaning-6474
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA