(NInfer Fork) I wanted to have a 1M context Qwen-3.8 27B, tp2, dual 5090s
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Hey! I forked NInfer (a from-scratch C++20/CUDA inference engine for Qwen models) and added two things: tensor-parallel across two GPUs, and YaRN ×4 rope scaling. Together they let Qwen3.8-27B NVFP4 run a 1,048,576-token context on two consumer 5090s — 27.4 GB per card, no NVLink. Numbers (single stream, 500 W per GPU cap):
Two GPUs are also just faster than one: 75 vs 54 tok/s at 250k, because weights and KV traffic halve per card and the ~128 cross-GPU reductions per token cost only ~0.2 ms under CUDA graphs. [link] [comments] |
More from r/LocalLLaMA
-
an unscientific qwen 3.8 flash next and glm 5.3 flash comparison
Aug 30
-
Ran Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max — speed vs context depth, 100 turns, one graph
Aug 30
-
Nemotron-3.5-Lightning at 11.77 GiB, a 16 GB option for a model that didn't have one
Aug 29
-
Humaneval benchmark for Deepseek V4 Flash 0731 vs GLM5.3 Flash on 2x DGX Spark setup
Aug 29
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.