Skip to content
Hacker News front page

Running local models is good now: Vicki Boykis's hands-on take

Running local models is good now

Vicki Boykis has been running local models on an M2 Mac with 64GB RAM for over a year, and now finds them genuinely useful. Gemma 4 26B and the newer 12B qat variant let her do agentic coding locally at roughly 75% of frontier-model accuracy. She uses Pi as the agent harness and LM Studio for inference, all inside a Docker container with restricted permissions. The post doesn't give token speeds or latency numbers, but notes the KV cache can eat all 64GB of RAM.

Why it matters: Vicki Boykis is a respected technical blogger; this is a first-person long-term usage report with specific hardware, models, and a quality comparison — not marketing. The score stays at the featured threshold of 72 because the post lacks key deployment data (latency, generatio...

Read the original ↗Export Markdown