A happy user of llama-swap
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I recently considered purchasing another GPU, not so much because I need more bandwidth or vram, but because I have several local GPU-based scheduled workloads that use different models/config. After shopping around looking at prices, I decided to use a different approach.
After having issues with llama.cpp model router concerning 127.0.0.1 vs 0.0.0.0 (it seems to assume that that it and the client are on the same server) I tried llama-swap https://github.com/mostlygeek/llama-swap. I scheduled my different GPU workloads a bit more creatively, and swapped models as the first step in the crontab.
It has worked really really well. No-Statement-0001 even quickly answered a question that I had. Great software and highly recommended.
cheers
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.