I hosted Kimi K3 (2.8T parameters) using 8 B300s. 92 tok/s, $190 per million tokens
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| What I ran:
I also ran Unsloth's Dynamic GGUF. Their 1-bit UD-IQ1_S (594 GB) fits 8x A100-80GB via llama.cpp. $19.99 per hour, 2.8x cheaper. Result: ~9 tok/s, TTFT 7 to 60 s, ~$620 per million tokens, so 3.3x more expensive per token. Quality at 1-bit was fine (correct arithmetic, coherent prose). Full write-up with every flag, the Modal deployment file, and the raw benchmark JSON: https://books.vizuara.ai/book/kimi-k3-hosting [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.