r/LocalLLaMA · · 1 min read

Useful > Fast

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I run 2 x 3090 on a ryzen 7 with 32gb ddr4 6000.

I see a lot of posts about maxing speed / context pool. I’ve done this myself with 3.8 27b.

What I’m interested in though is once the dust settles and we look at utility, what balance people are striking in terms of similar setups and models. As a for instance, some of my requirements involve producing images and video clips. Some of them involve creating scripts for things.

This has led to an overnight CPU/Ram run on flux2 and minimax h3, a reduction in maximum context for 3.8 27b so I can run a Gemma MOE in offload for improved writing prose / checking over Qwen and a smaller vision tower to automatically checking flux and H3 outputs mid run so I don’t lose a night.

I need concurrency so have a set up that gives me that when I’ve got open code (I know not everyone’s go to but I like it), open science, and Hermes all going side by side.

It’s fun to optimise, but I’d love to hear how people are setting up to from a flexibility and utility viewpoint once that’s done for their real use cases.

submitted by /u/jbro1985
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA