r/LocalLLaMA · · 1 min read

Does PCIe matter much for inference?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Speaking of -sm tensor.

I have 2x5060ti and i get some 40-50tps form 3.8 27B q6.

I used HWinfo to see how saturated the PCIes are during inference are and as expected both the PCIe 5x16 slot and PCIe 4x4 were fully saturated.

I cant help but feel like my 2nd gpu slot is a bottleneck (4x4), im considering getting a riser for my spare NVMe 5x4 slot and plugging the card there.

I wanted to hear your experiences with PCIe bottlenecks and risers before comitting to any purchase.

submitted by /u/Ok-Conflict391
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA