Skip to content
Hacker News front page

Local LLM setup on M4 Pro Mac Mini: Qwen3.6 and Gemma-4 on 48GB RAM

My local model setup on an M4 Pro Mac Mini

Kevin Lewis shares his local LLM setup on an M4 Pro Mac Mini with 48GB RAM. His main model is Qwen3.6-35B-A3B-OptiQ-4bit (MoE, 3B active params per token, ~20GB RAM), and a lightweight Gemma-4-E4B-it-OptiQ-4bit (~2.4GB). He uses oMLX as the inference server, Tailscale to connect iPhone and MacBook, and runs Hermes agent backend, Apollo chat, Pi coding, and Raycast queries. His point: local isn't meant to replace cloud APIs entirely, but to handle 80% of daily requests, avoiding price hikes, rate limits, model degradation, and data privacy risks. The post doesn't disclose specific inference speed or power consumption figures.

Read the original ↗Export Markdown