Freetokens project is impressive
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| A new project was released yesterday and I have the opportunity to test it today. Papper: https://arxiv.org/abs/2608.16157 My initial tests with the following setup: I got 100tok/s on QWEN3.6-35B-A3B NVFP4 (20GB - does not fit in my VRAM). (Example bellow with a 1028 token prompt - ~110 tok/s) [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.