Fastest qwen 3.8 27b for AMD gpu?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Hey, just wondering if there are forks or exact gguf versions that give fastest prompt processing and token gen speeds for AMD gpu?
Looking to run q8 or q6
Vram 96gb
W7900 + w7800 both 48gb
With bandwidth mismatch, tensor paralleling amd equivalent not working
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.