Skip to content
r/LocalLLaMA

Llama.cpp PR adds Maple 20B-A1B ternary MoE architecture for CPU inference

llama: add Maple 20B-A1B ternary MoE architecture (CPU) by AlexGabbia · Pull Request #27000 · ggml-org/llama.cpp

Reddit user AlexGabbia submitted PR #27000 to llama.cpp, adding the Maple 20B-A1B ternary MoE architecture. The model has 20B total parameters but activates only 1B per inference, designed for CPU execution. Ternary weights (-1, 0, 1) cut memory and compute significantly, enabling faster large-model inference on ordinary CPUs. The post is blocked by Reddit and does not disclose specific performance numbers or release timeline.

Read the original ↗Export Markdown