Skip to content
Hacker News front page

DeepGrove open-sources Maple-Preview, a 20B ternary MoE model hitting 127 tok/s on iPhone

Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone

DeepGrove released Maple-Preview, a 20B-parameter, 1B-active ternary-weight reasoning model with a 5.31 GB checkpoint. It hits 218 tok/s on an M4 Mac mini and 127 tok/s on an iPhone—13× faster than 1-bit Bonsai 27B. The model scored 7/7 on IMO 2024 Problem 1 and leads its weight class on AIME and other reasoning benchmarks, trading blows with larger models. The post doesn't disclose training data, contamination checks, or specific agent-benchmark scores, and notes agentic performance may lag. I'd treat the raw reasoning numbers as solid but wait for agent evals before getting excited there.

Why it matters: Ternary-weight MoE that fits a 20B model on a phone with 127 tok/s and IMO-level reasoning earns featured. Not scoring higher because we only have the model card—no third-party benchmarks or real-world use cases yet. 82 feels right for now.

Read the original ↗Export Markdown