Skip to content
Hacker News front page

Agent Harness Engineering: The Scaffolding Matters More Than the Model

Agent Harness Engineering

Addy Osmani argues that a coding agent's effectiveness depends as much on its harness—prompts, tools, hooks, sandboxes, and feedback loops—as on the underlying model. Citing Viv Trivedy and others, he notes a decent model with a great harness beats a great model with a bad one. A key data point: Claude Opus 4.6 jumped from Top 30 to Top 5 on Terminal Bench 2.0 solely by switching to a custom harness. The core discipline is a ratchet: every agent mistake becomes a permanent rule or check in AGENTS.md or a pre-commit hook, so it never repeats.

Why it matters: Addy Osmani's post crystallizes 'Agent Harness Engineering' with a killer data point: Claude Opus 4.6 went from 28% to 72% on Terminal Bench 2.0 just by changing the harness. The idea isn't brand new, but naming and systematizing scattered engineering practices into a coherent...

Read the original ↗Export Markdown