r/LocalLLaMA · · 1 min read

Qwen 30b MoE - 30tps - 6GB vram - Done!

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Qwen 30b MoE - 30tps - 6GB vram - Done!

So,
I have been dreaming of getting 17 tokens per second using my RTX 3050 6GB version on a decent context window for Hermes needed above 60k. The hope is that has was a 22GB of DDR 4, hoping they can take some of those experts and give me room for context. What did I get 10 or less tokens per second. 😄

Not today!!

Today I could run it with 90k context with Hermes I had 20-25 tps. And when I changed harness I got even 30-35tps 🥳🥳🥳
NOT benchmarks- but actual session generation with context and actual work being done 😄😄😄

I will come to edit the post and add details. Just wanted to share the joy with anyone out there with a peasant rig like mine 😅

May be someone who does better can also share the positive vibe.

Cheers for now 🙋🏾‍♂️

submitted by /u/Bakkario
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA