r/LocalLLaMA · · 1 min read

A happy user of llama-swap

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I recently considered purchasing another GPU, not so much because I need more bandwidth or vram, but because I have several local GPU-based scheduled workloads that use different models/config. After shopping around looking at prices, I decided to use a different approach.

After having issues with llama.cpp model router concerning 127.0.0.1 vs 0.0.0.0 (it seems to assume that that it and the client are on the same server) I tried llama-swap https://github.com/mostlygeek/llama-swap. I scheduled my different GPU workloads a bit more creatively, and swapped models as the first step in the crontab.

It has worked really really well. No-Statement-0001 even quickly answered a question that I had. Great software and highly recommended.

cheers

submitted by /u/our_sole
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA