Qwen3.8-Flash-Next (UD-IQ4_XS) on 2x RTX 3060 + 7800X3D, from initial 36 tps prefill to 400 tps and other benchmarks (-sm tensor trap) + VRAM/RAM usage
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
TL;DR:
llama.cpp with --load-mode mmap used 21-32 GB of RAM, with ik_llama.cpp using 106-108 GB. -sm tensor killed my prefill, changing to -sm layer went from 36 tps to 135 tps. After that, -ubatch 2048 pushed it up to 400 tps at -c 131072. I can't fit -ubatch 2048 at -c 262144, which is why it caps out for me at ~180-200 tps prefill and 12 tps decode at 262144 context.
After testing locally and getting 14-15 tps decode initially, I hit a brick wall with 36 tps prefill which isn't usable to me. After failing to improve that value after tweaking a lot of settings, I decided to run a bunch of benchmarks locally and the results surprised me, so I decided to post them so others can benefit from the results. All tests were from locally compiled llama.cpp and ik_llama.cpp binaries on my CachyOS system. CUDA 0 is on 16x PCI lanes and has all 12 GB available, while CUDA 1 runs my desktop UI and idles at 1.5/12 GB normally on 4x PCI lanes.
Everything below this line is an AI-generated report of the benchmark findings, followed by the raw data. It's AI-generated because fuck you; I'm not writing that all out. Have a nice day.
Report:
Benchmark data for Qwen3.8-Flash-Next UD-IQ4_XS (93.7 GB, 125B MoE / 6B active) on a consumer dual-GPU box, comparing llama.cpp and ik_llama.cpp.
Everything below was measured on one machine over a single session. All prefill figures come from an 8k-token synthetic prompt with cache_prompt: false.
Findings
1. -sm tensor costs 7x prefill on llama.cpp. 41.7 t/s with -sm tensor, 303 t/s with -sm layer, same hardware and same everything else. llama.cpp op-offloads CPU-resident MoE weights to the GPU for batches of 32 or more, copying only the experts a batch actually uses (ggml-backend.cpp:1643). That code path targets a single device, so when tensor-parallel splitting shards the weights it stops applying and every expert matmul falls back to the CPU. Measured proof: -sm tensor (35 t/s) and GGML_OP_OFFLOAD_MIN_BATCH=999999, which disables op-offload outright, (38 t/s) give the same number.
Instrumentation during prefill:
| Split mode | llama-server CPU | GPU0 util | GPU1 util | Prefill |
|---|---|---|---|---|
-sm tensor | 1191% | ~0% | ~0% | 41.7 t/s |
-sm layer | 139% | 49% | 38% | 135 t/s |
2. -ub is the second big knob, and it does nothing until finding 1 is fixed. At -sm tensor every ubatch value tested returned the same speed. At -sm layer the same sweep gives 135 / 191 / 303 t/s for 512 / 1024 / 2048. The two fixes multiply. Neither alone gets you close.
3. ik_llama.cpp has no split-mode cliff. -sm layer and -sm graph measure the same on ik (407 vs 401 t/s). The trap is specific to llama.cpp.
4. Which engine wins depends entirely on whether -ub 2048 fits. The engines tie at -ub 512 and -ub 1024. ik has a fast path at -ub 2048.
5. ik_llama.cpp needs roughly 75 GB more system RAM than llama.cpp. Same model, same machine, same flags. llama.cpp under --load-mode mmap reports 21 to 32 GB used with 93 GB available, because the model pages sit in page cache and are reclaimable. ik reports 106 to 108 GB used with only 16 to 18 GB available. Both fit in 128 GB, but on a 96 GB box this is the difference between ik running and not running, and it leaves very little headroom for anything else on the machine.
| Engine | RAM used | RAM available | Measured on |
|---|---|---|---|
llama.cpp --load-mode mmap | 21 to 32 GB | 93 GB | rows L, O, V3, LL |
| ik_llama.cpp (default mmap) | 106 to 108 GB | 16 to 18 GB | rows IKM1, IKM2, IKM3 |
6. Keeping expert layers in VRAM is worth less than the ubatch buffer. -ncmoe 48 / 46 / 44 give 135 / 139 / 141 t/s at -ub 512. Pulling experts into VRAM buys almost nothing, and it costs enough VRAM that -ub 2048 then OOMs. Setting -ncmoe 48 (all experts on CPU) and spending the VRAM on the ubatch compute buffer is the better trade.
7. Decode is memory-bandwidth bound and no flag fixes it. 12 t/s at shallow context on 128 GB DDR5 running at 3200 MT/s (about 40 GB/s usable). For reference, the PR thread reports 28 t/s for the same model on a 12-channel DDR5 EPYC 9555.
8. Extra parallel slots cost single-stream speed and add little aggregate. Aggregate throughput stayed roughly flat from 1 to 4 concurrent slots, and running the server with --parallel 8 dropped single-stream decode from about 12 t/s to about 5 t/s.
Recommendations
| Context | Engine | Key flags | Prefill | Decode | VRAM (GPU0/GPU1) | System RAM |
|---|---|---|---|---|---|---|
| up to 131K | ik_llama.cpp | -sm layer -ncmoe 48 -ub 2048 -b 4096 -fa on | 407 t/s | 13 t/s | 9.2 / 9.1 GB | ~108 GB |
| 196K to 262K | llama.cpp | -sm layer --n-cpu-moe 48 -ub 1024 -b 4096 --flash-attn on | 215 t/s | 12 t/s | 9.5 / 8.9 GB | ~32 GB |
If you have less than about 128 GB of RAM, use llama.cpp regardless of context size. See finding 5.
Additional settings used in both: -t 8 -tb 16, -ctk q8_0 -ctv q8_0, -ts 60,40, --parallel 1.
Things that did not help, all measured:
| Setting | Result |
|---|---|
-rtr (ik runtime repack) | 159 t/s vs 407 t/s. Optimises for CPU kernels and loses GPU op-offload |
-ictk q8_0 (ik indexer cache) | No change to speed or VRAM. Still OOMs at 262K + -ub 2048 |
-no-fmoe (ik) | 399 vs 407 t/s, so -fmoe is worth about 2% |
-ub 4096 | OOM at every context tested |
--threads-batch 8 vs 16 | 300 vs 303 t/s, no meaningful difference once the GPU does the work |
KV cache q5_1 instead of q8_0 | Did not free enough VRAM to change any outcome |
Lowering -ncmoe to 46 or 44 | +4 to +6 t/s at -ub 512, and OOM at -ub 2048 |
Avoid --threads-batch 12 on an 8-core/16-thread CPU. ggml puts a barrier after every op, so 4 cores end up running 2 threads while 4 run 1, and every op waits on the doubled cores.
Test system
| Device | Device info |
|---|---|
| CPU | AMD Ryzen 7 7800X3D, 8C/16T, AVX-512 |
| RAM | 128 GB DDR5 at 3200 MT/s (4x32 GB; the board will not POST at rated speed with 4 DIMMs) |
| GPU0 | RTX 3060 12 GB, PCIe 4.0 x16, direct to CPU |
| GPU1 | RTX 3060 12 GB, PCIe 4.0 x4, behind the chipset, also drives the desktop (~1.3 GB) |
| OS | CachyOS, Linux 7.2.0 |
| Model | unsloth Qwen3.8-Flash-Next-GGUF UD-IQ4_XS, 93.7 GB, 3 shards |
| Architecture | 48 layers: 36 Gated DeltaNet + 12 Qwen Sparse Attention, 512 experts, 262144 native context |
| llama.cpp | build 4e97ac86e, CUDA on, GGML_NATIVE=ON, arch 86 |
| ik_llama.cpp | build 7cff686d (includes PR #2365 and the #2367 grid-overflow fix) |
The x4 link on GPU1 was investigated and ruled out as the bottleneck. During slow prefill both GPUs sat near 0% utilisation, so the link was never saturated.
llama.cpp results
Context 131072, 8k prompt, -b 4096 -t 8 -tb 16 -ctk q8_0 -ctv q8_0, --load-mode mmap. VRAM is nvidia-smi used, sampled with the server loaded and the benchmark just finished. RAM is total system used, which includes about 6 GB of desktop.
| # | -sm | -ncmoe | -ub | -tb | Prefill t/s | Decode t/s | GPU0 MiB | GPU1 MiB | RAM |
|---|---|---|---|---|---|---|---|---|---|
| A | tensor | 40 | 512 | 16 | 41.7 | n/a | 11541 | 11267 | 32G |
| G | tensor | 48 | 512 | 16 | 35 | n/a | 6458 | 6260 | 31G |
| P | layer | 48 | 512 | 16 | 38 | 9 | 4620 | 5958 | 31G |
| C | layer | 48 | 512 | 16 | 133 | n/a | 5654 | 5983 | 31G |
| L | layer | 48 | 512 | 16 | 135 | 10 | 5656 | 5974 | 31G |
| M | layer | 46 | 512 | 16 | 139 | 12 | 5656 | 9052 | 31G |
| N | layer | 44 | 512 | 16 | 141 | 12 | 5654 | 11328 | 31G |
| LL | layer | 48 | 1024 | 16 | 191 | 8 | 6587 | 6464 | 21G |
| U | layer | 48 | 2048 | 8 | 300 | 12 | 7778 | 8900 | 31G |
| O | layer | 48 | 2048 | 16 | 303 | 12 | 7778 | 8902 | 31G |
Row P is row L with GGML_OP_OFFLOAD_MIN_BATCH=999999 set, which disables op-offload. Rows A and G VRAM were sampled live mid-run rather than at test end.
Context 262144
| # | -sm | -ncmoe | -ub | -ts | Prefill t/s | Decode t/s | GPU0 MiB | GPU1 MiB | RAM |
|---|---|---|---|---|---|---|---|---|---|
| V3 | layer | 48 | 1024 | 60,40 | 215 | 12 | 9478 | 8925 | 32G |
llama.cpp configurations that failed to load
| # | ctx | -sm | -ncmoe | -ub | -ts | Failure |
|---|---|---|---|---|---|---|
| B/D/E | 131072 | layer | 40 | 512 | 51,49 | OOM, 12281 MiB on device 1 |
| S | 131072 | layer | 46 | 2048 | 51,49 | OOM, 3888 MiB on device 1 |
| T | 131072 | layer | 44 | 2048 | 51,49 | OOM, 3888 MiB on device 1 |
| R | 131072 | layer | 48 | 4096 | 51,49 | OOM, 7776 MiB on device 1 |
| V | 262144 | layer | 48 | 2048 | 51,49 | OOM, 7216 MiB on device 1 |
| V1 | 262144 | layer | 48 | 2048 | 70,30 | OOM, 6920 MiB on device 0 |
| V2 | 262144 | layer | 48 | 2048 | 60,40 | OOM, 7200 MiB on device 1 (KV at q5_1) |
ik_llama.cpp results
Context as noted, 8k prompt, -ncmoe 48 -b 4096 -t 8 -tb 16 -ctk q8_0 -ctv q8_0 -fa on -ts 60,40.
| # | ctx | -sm | -ub | Extra | Prefill t/s | Decode t/s | GPU0 MiB | GPU1 MiB | RAM |
|---|---|---|---|---|---|---|---|---|---|
| IK8 | 131072 | layer | 512 | 136 | 8 | ||||
| IKM3 | 131072 | layer | 512 | rerun of IK8 | 136 | 10 | 5659 | 5817 | 106G |
| IK7 | 131072 | graph | 512 | 137 | 9 | ||||
| IK9 | 131072 | layer | 1024 | 178 | 10 | ||||
| IK4 | 131072 | layer | 2048 | --no-mmap -rtr | 159 | 11 | |||
| IK5 | 131072 | layer | 2048 | -no-fmoe | 399 | 12 | |||
| IK6 | 131072 | graph | 2048 | 401 | 12 | ||||
| IK1 | 131072 | layer | 2048 | 407 | 13 | ||||
| IKM1 | 131072 | layer | 2048 | rerun of IK1 | 405 | 12 | 9223 | 9142 | 108G |
| IK3 | 262144 | layer | 1024 | 179 | 10 | ||||
| IKM2 | 262144 | layer | 1024 | rerun of IK3 | 178 | 9 | 9872 | 9676 | 107G |
| IK12 | 262144 | layer | 1024 | -ictk q8_0 | 180 | 10 |
The IKM rows are straight repeats of IK1, IK3 and IK8 run later with memory capture added. They replicate to within 2 t/s (405 vs 407, 178 vs 179, 136 vs 136), which is a useful check on how repeatable these measurements are.
ik holds far more resident RAM than llama.cpp. ik sits at 106 to 108 GB used with 16 to 18 GB available, while llama.cpp under --load-mode mmap sits at 21 to 32 GB used with 93 GB available, because its model pages stay in page cache rather than counting as used. Same model, same machine. On a 128 GB box ik still fits, but there is much less headroom for anything else, and it is worth checking before running ik on a machine with less RAM.
ik_llama.cpp configurations that failed to load
| # | ctx | -ub | -ts | Extra | Failure |
|---|---|---|---|---|---|
| IK2 | 262144 | 2048 | 60,40 | OOM, 7254 MiB on device 0 | |
| IK10 | 262144 | 2048 | 60,40 | -ictk q8_0 | OOM, 7254 MiB on device 0 |
| IK11 | 262144 | 2048 | 70,30 | -ictk q8_0 | OOM, 8278 MiB on device 0 |
Engine comparison at matched settings
Context 131072, -ncmoe 48, 8k prompt.
-ub | llama.cpp | ik_llama.cpp |
|---|---|---|
| 512 | 135 | 136 |
| 1024 | 191 | 178 |
| 2048 | 303 | 407 |
Context 262144, -ncmoe 48, -ub 1024.
| # | llama.cpp | ik_llama.cpp |
|---|---|---|
| Prefill | 215 | 179 |
| Decode | 12 | 10 |
Prefill and decode versus prompt depth
llama.cpp, context 262144, -ncmoe 48 -ub 1024 -sm layer -ts 60,40, single 59k prompt. Values are the running average the server reports at each 4096-token boundary.
| Tokens | Prefill t/s | # | Tokens | Prefill t/s |
|---|---|---|---|---|
| 4096 | 220.6 | 36864 | 195.2 | |
| 8192 | 215.0 | 40960 | 193.3 | |
| 12288 | 211.5 | 45056 | 191.3 | |
| 16384 | 207.6 | 49152 | 189.3 | |
| 20480 | 204.9 | 53248 | 187.5 | |
| 24576 | 202.1 | 57344 | 185.4 | |
| 28672 | 199.8 | 58837 | 183.8 | |
| 32768 | 197.5 |
Decode over the same run fell from 12 t/s at an 8k prompt to 8 t/s at 59k.
Fitting t = a*n + b*n^2/2 to this curve gives a = 4.47e-3 and b = 3.24e-8, which extrapolates to roughly 115 t/s averaged over a full 262144-token cold prefill, about 38 minutes. That is an extrapolation from 59k and was not measured directly.
The PR thread attributes this decay to the sparse attention indexer's pooled-block cost scaling with cache length.
Parallel slots
llama.cpp, context 262144, --parallel 4, four 4k prompts fired simultaneously, 128 tokens generated each.
| Concurrent | Per-slot prefill t/s | Per-slot decode t/s | Aggregate prefill t/s |
|---|---|---|---|
| 1 | 215.5 | 10.53 | 131.1 |
| 2 | 107.6 | 4.07 | 117.4 |
| 4 | 71.2 | 2.85 | 134.5 |
Decode isolated, context 65536, --parallel 8, trivial prompts, 64 tokens each.
| Concurrent | Per-slot decode t/s | Aggregate decode t/s |
|---|---|---|
| 1 | 4.96 | 4.17 |
| 2 | 1.67 | 3.28 |
| 4 | 1.78 | 6.82 |
Note the single-slot decode of 4.96 t/s on a --parallel 8 server against about 12 t/s on a --parallel 1 server. Reserving slots costs speed even when they are idle.
Caveats
Read these before drawing conclusions from small differences.
- Single runs, no repeats. Run-to-run variance is roughly 10%. llama.cpp at
-ub 1024measured 191 t/s at 131K context but 215 t/s at 262K, which is backwards and shows the noise floor. Only the large effects (the 7x split-mode gap, the 2.2x to 3x ubatch effect, ik's 34% lead at-ub 2048) are clearly outside it. - All benchmarks ran with speculative decoding disabled. Real-world decode with
--spec-type ngram-modon repetitive coding content should be higher than the numbers here. - Decode was measured over only 64 generated tokens, which is a small sample.
- The 4-slot parallel decode figures are noisy, as shown by 2 slots scoring below 1 slot.
--n-cpu-ffnwas absent from every configuration benchmarked. The recommendation to drop it is inferred from it never being present, not from an A/B test.- The source of ik's 34% lead at
-ub 2048is not established.-fmoeaccounts for about 2% of it. The remaining 32% was not traced to a mechanism. - The full 262144-token prefill figure is extrapolated, not measured.
- Memory figures are whole-system
nvidia-smiandfreereadings, so they include the desktop and other processes.
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.