Concurrency plus nvfp4 on Blackwell
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| ~2000 tps in aggregate performing bulk captioning on images. Above is parsed from vllm log while a client runs 30 concurrent streams, each concurrent stream has 1 request with an image and prompt, then a 2nd request on the same stream (so 1st Q:A would be cached). Typical log line: Engine 000: Avg prompt throughput: 1301.0 tokens/s, Avg generation throughput: 1924.0 tokens/s, Running: 30 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.8%, Prefix cache hit rate: 0.0%, MM cache hit rate: 50.1% This is running on a RTX Pro 6000 Blackwell, but I don't think I'm actually using nearly all the VRAM yet. A 5090 should be able to get close if your individual chats are not that long as to fit into VRAM. Maybe kv cache will evict and impact perf. Here's another graph comparing to some other dense models as well using lmarena-ai/VisionArena-Chat as a test set: The quanttrio is Qwen 3.5, the rest are all Qwen 3.5. 27B isdense, 35B is moe. Unsloth is ~26GB and nvidia is ~22GB, I believe because unsloth left more unquantized layers. nvidia 35b is 23.4GB. I was actually a bit surprised that with concurrency the MOE was so far ahead, but running the Monte Carlo, about 53% (union of selected) experts are expected to be chosen per forward execution at c=24, or still only ~56% at q=0.95. Or ~61% at c=30. [link] [comments] |
More from r/LocalLLaMA
-
Unpopular opinion Qwen 3.8 is hard to understand
Aug 30
-
Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM
Aug 30
-
Oh so that's where my PCIe lanes went...
Aug 30
-
Uncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and…
Aug 30
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.