Skip to content
Hacker News front page

Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

Running Gemma 4 26B at 5 tokens/SEC on a 13-year-old Xeon with no GPU

The author got Google's Gemma 4 26B MoE model running on a dual Xeon E5-2690 v2 server from 2013 with no GPU, costing under $300. The CPUs only support AVX1, but ik_llama.cpp's optimized kernels require AVX2, causing silent gibberish output. Claude diagnosed that the graph builder unconditionally emitted MOE_FUSED_UP_GATE ops while the dispatcher had no matching case, leaving ~240 tensors per forward pass reading uninitialized memory. After the fix, decode reaches ~5.2 tokens/sec and prompt eval ~16 tokens/sec. A PR is open but not yet merged. The post doesn't disclose quantized model memory usage or power draw.

Why it matters: A first-person experiment with real numbers, not a generic 'run LLMs locally' tutorial. Gemma 4 26B MoE on a 13-year-old Xeon, no GPU, sub-$300 total cost — every detail is concrete. HKR all hit, but it's a personal blog experiment, not a product launch or research breakthroug...

Read the original ↗Export Markdown