Llama.cpp PR adds Maple 20B-A1B ternary MoE architecture for CPU inference
llama: add Maple 20B-A1B ternary MoE architecture (CPU) by AlexGabbia · Pull Request #27000 · ggml-org/llama.cpp
Reddit user AlexGabbia submitted PR #27000 to llama.cpp, adding the Maple 20B-A1B ternary MoE architecture. The model has 20B total parameters but activates only 1B per inference, designed for CPU execution. Ternary weights (-1, 0, 1) cut memory and compute significantly, enabling faster large-model inference on ordinary CPUs. The post is blocked by Reddit and does not disclose specific performance numbers or release timeline.