MTPLX is a native Mac app for local AI on Apple Silicon. It runs Qwen 3.8 and other MLX models with native MTP speculative decoding: chat, coding agents, and an OpenAI and Anthropic compatible local server, all at twice the speed, with the output left exact. On an M5 Max, at each model's own sampler: 81.74 tok/s on a 27B, up to 84 tok/s on the 125B Flash-Next, 227.8 tok/s on Qwen 3.5 4B.
Twice the speed at no quality loss, at any temperature, and up to 3x on the 8-bit Quality pack. The output is identical, verified bit for bit. The record lane: 81.74 tok/s on Qwen 3.6 27B at 2.69x plain decode, M5 Max, raw logs on Benchmarks. How it works →
One click serves OpenCode, Pi, Hermes, or the web UI from your Mac. OpenAI and Anthropic compatible, so everything plugs in.
Native Swift, fully offline, in twelve languages, dark or light. Watch your models write code and prose at twice the speed.
Paste a Hugging Face link. Forge converts it to MLX and measures the speedup on your Mac.