TurboFieldfare: Run Gemma 4 26B on any M-series Mac with 2 GB RAM
Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
A Swift + Metal inference engine runs 4-bit Gemma 4 26B-A4B-IT using about 2 GB RAM. The 14 GB weights won't fit conventional tools on 8 GB Macs. It keeps shared layers and KV cache in RAM, streams routed experts per token from SSD, and hides SSD latency with a small expert cache plus parallel preads. Hits 5–6 tok/s on an 8 GB M2 MacBook Air, 31–35 tok/s on an M5 MacBook Pro. Includes an experimental OpenAI-compatible local server with streaming and tool calls. The post doesn't spell out quantization details or expert cache hit rates.
Why it matters: An open-source inference engine gets Gemma 4 26B running on a Mac with 2GB RAM, with concrete technical details (4-bit quantization + per-token expert streaming from SSD). High practical value for M-series Mac users. Score capped at 78 because it's a solo project with no bench...