I just built a mini Kimi-K3 from Scratch under 250$. Already beats GPT-2 (124M)!
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| I pre-trained a 1.02-billion-parameter on Kimi K3 replica trained on 5.00 billion decontaminated tokens for $250. This model has 1.02 billion parameters, of which 145 million are active per token. It is roughly one two-thousandth of K3 by total size. It saw 5,000,003,584 tokens, which is a rounding error against the corpora frontier models are trained on. It has never been instruction-tuned, and it has only ever done one thing: predict the next token. What it does have is K3's architecture: - Kimi Delta Attention, Gated MLA, Attention Residuals - LatentMoE with the same aux-loss-free balancer - Same activation function with the same two constants - K3's own 163,840-token tokenizer, unmodified. I report a 33.4% HellaSwag which beats the GPT-2 124M score of 28% Read the entire tutorial here: https://books.vizuara.ai/book/pretraining-a-mini-k3 [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.