Skip to content
r/LocalLLaMA

I built a coding agent that gets 87% on benchmarks with a 4B parameter model

I built a coding agent that gets 87% on benchmarks with a 4B parameter model, here's how

SmallCode passes 87 of 100 benchmark tasks with Gemma 4 activating 4B parameters per token. The author attributes the result to compound tools, compile and lint feedback, task decomposition after two repeated failures, and optional escalation to Claude or OpenAI for one task.

Why it matters: HKR-H/K/R all pass, but this is a single Reddit post and the benchmark identity plus replication details are incomplete. It fits a concrete first-person experiment above the featured bar, not the 78+ band.

Read the original ↗Export Markdown