Does PCIe matter much for inference?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Speaking of -sm tensor.
I have 2x5060ti and i get some 40-50tps form 3.8 27B q6.
I used HWinfo to see how saturated the PCIes are during inference are and as expected both the PCIe 5x16 slot and PCIe 4x4 were fully saturated.
I cant help but feel like my 2nd gpu slot is a bottleneck (4x4), im considering getting a riser for my spare NVMe 5x4 slot and plugging the card there.
I wanted to hear your experiences with PCIe bottlenecks and risers before comitting to any purchase.
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.