Skip to content
r/LocalLLaMA

MTPLX V1: A native Swift app for running and creating MLX MTP models on Mac, doubling Qwen 3.6 27B speed

MTPLX V1: The Swift App For Running & Creating MLX MTP Models (2x TPS Qwen 3.6 27B)

Developer YoussofAl rebuilt MTPLX as a native Mac app—a 55MB DMG with the full engine bundled. The key claim is mathematically exact speculative decoding on Apple Silicon: Qwen 3.6 27B went from 28 tps to 63 tps. The new Forge feature fixes the biggest pain point from v0.1: paste a Hugging Face link, and it converts the model to MLX with MTP heads wired up, then measures real speedup on your machine. It includes a streaming chat UI, a live decode dashboard, built-in AIME 2026 benchmarking, and support for smaller models like Qwen 3.5 9B and Gemma 4. KV cache now persists to SSD so sessions survive restarts.

Why it matters: Solid local inference tool with concrete 2x speedup numbers, but audience is limited to Apple Silicon + MLX users — too niche for broader resonance. H and K both hit, R missing, just clears the featured bar. Score stays at the lower end because this is toolchain optimization, ...

Read the original ↗Export Markdown