r/LocalLLaMA · · 1 min read

Freetokens project is impressive

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Freetokens project is impressive

A new project was released yesterday and I have the opportunity to test it today.

Papper: https://arxiv.org/abs/2608.16157
Github: https://github.com/FlashML-org/FreeToken

My initial tests with the following setup:
RTX 5080 (16 GB)
DDR6 64GB
AMD Ryzen 9 9950X3D

I got 100tok/s on QWEN3.6-35B-A3B NVFP4 (20GB - does not fit in my VRAM).
Have you already tried it?

(Example bellow with a 1028 token prompt - ~110 tok/s)

https://preview.redd.it/x21sl7oo2wkh1.png?width=833&format=png&auto=webp&s=7372da17978de714ac95d464f340bb24017cab80

submitted by /u/ViRROOO
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA