r/LocalLLaMA · · 1 min read

Is ternary (1.58-bit) LLMs making a come back?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I'm just thinking, ever since microsoft announced bitnet, this sub (and myself) has been hoping for massive ternary models. In the last month alone, prismML dropped 27B ternary (though I've read community experience suggested it sometimes didn't hold up to it's benchmarks), Deepgrove dropped their ternary maple-20b-a1b which from my experience works really well and clocks like 100 tok/s on an iphone, and Doses AI dropped pestle-27b-ternary medical specialised which beats medgemma-27b nearly across the board.

The common problem across all of them is long-horizon agentic coding/work, but i really think that's because all of these are new small labs that haven't prioritised RL-maxxing yet - they have indicated this is their next step though.

I'm hopeful, and it seems like we could be very close to a massive ternary model MoE that's actually competitive with qwen3.8 at coding and agentic work. Or have most folks lost faith in ternary architecture?

submitted by /u/Individual-Dot5488
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA