PR for running Ternary-Bonsai-8B-Q2_0.gguf in llama.cpp with CUDA support just got merged
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Time to see what it's capable of
Join the most real place on the internet
By continuing, you agree to our User Agreement and acknowledge that you understand the Privacy Policy.
Comments Section
Why just 8B? 27B should also be supported right?
Mb, I meant 27B, idk what happened to my brain when posting. Doesn't matter which one though, it's Q2_0 being supported rn
More from r/LocalLLaMA
-
Demo of local document extraction (52 pages) using Arctic Embed and Bonsai 8B on an Iphone 16 (KernelAI app)
Aug 30
-
Will apple still release devices with mobile HbM in 2027 ?
Aug 30
-
Whatever happened to OpenClaw and its derivatives?
Aug 30
-
Qwen 3.8 Flash Next locally on simple mobile phone at 3.5 tok/s
Aug 30
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.