r/LocalLLaMA · · 1 min read

~ 2x Speed Boost for Qwen3.8 27B on Apple Silicon

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

~ 2x Speed Boost for Qwen3.8 27B on Apple Silicon

https://x.com/koc_z3/status/2093581036756025744?s=46

~ 2x speed boost for Qwen3.8 27B on Apple Silicon

~ 1.5x speed boost for Qwen3.6 35B AЗB

Tested on an M1 Max 64GB Mac using MTPLX with 262K (MAX) Context length.

Qwen3.8-27B (Q4):

- Decode ~ 21 TPS

- Prefill ~ 83 TPS (Peak 111 TPS)

Qwen3.6-35B-A3B (Q4):

- Decode ~ 55 TPS

- Prefill ~ 300 TPS (Peak 623 TPS)

Three key capabilities of this framework:

  1. Verified ~ 2x increase in local generation speed compared to base models.
  2. Auto-tuning: Determines the optimal MTP draft depth based on your specific chip, thermals, and memory bandwidth.
  3. Base Conversion: Transforms standard base models into MLX-ready MTP models.

Repo: github.com/youssofal/MTPLX

submitted by /u/koc_Z3
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA