Llamacpp server : How do the -np and -c flags interact?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I've been using lm studio for a few months. I want to try hermes agents with Qwen 3.6 MoE, so I'm switching to llama.cpp and I don't understand well how the server slots -np and the context size -c interact.
The context for each parallel client appears to be equally distributed across server slots (so each client is allowed c / np context).
I have some questions:
- What are the consequences of launching a server with a greater context -c than what the model allows?
- What if c / np is greater than the model max context? Are there any negative to that regarding model performance?
- If a rig allows to allocate twice the context max size in vram, is it twice energy and time efficient to serve two agents in parallel rather than sequentially?
[link] [comments]
More from r/LocalLLaMA
-
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp · Hugging Face
Aug 31
-
pipecat-ai/phonellm-alpha-1: GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost
Aug 31
-
Could this affect M5 Ultra price/availability?
Aug 31
-
GLM 5.3, GLM 5.3 Flash or 3.8 Qwen Flash for Replacing Kimi k3 IQ2_XXS
Aug 31
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.