~ 2x Speed Boost for Qwen3.8 27B on Apple Silicon
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| https://x.com/koc_z3/status/2093581036756025744?s=46 ~ 2x speed boost for Qwen3.8 27B on Apple Silicon ~ 1.5x speed boost for Qwen3.6 35B AЗB Tested on an M1 Max 64GB Mac using MTPLX with 262K (MAX) Context length. Qwen3.8-27B (Q4): - Decode ~ 21 TPS - Prefill ~ 83 TPS (Peak 111 TPS) Qwen3.6-35B-A3B (Q4): - Decode ~ 55 TPS - Prefill ~ 300 TPS (Peak 623 TPS) Three key capabilities of this framework:
[link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.