Are models with N-Gram tables going to completely change the AI race?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
The news about Qwen 3.8 Flash Next is the first I'm reading about n-gram tables. I may be completely misunderstanding how they work but it seems they could open the door for 1T+ parameter models to be run on a single server with modest GPUs and a ton of system RAM rather than needing a rack of GPU servers connected with something like NVlink.
Could we be looking at shrinking the capability gap between self hosted and flagship models faster than we thought, or am I way off base?
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.