Skip to content
Hacker News front page

How should you design the harness for a coding agent? This paper tests 176 configurations

An Empirical Study of Harness Design for Coding Agents

The paper breaks a coding agent harness into three swappable parts—planning, action space, and context management—and runs 176 matched comparisons on SWE-Bench Verified and Terminal-Bench 2.1. Context management matters most when the context budget is tight, mainly by preventing overflow failures. Staging rule-based elision before LLM summarization gives the best efficiency; making elided content recoverable adds complexity models rarely use. Planning helps weaker models with accuracy but mainly saves cost for stronger ones. Bash-capable models work well with a bash-only interface at lower cost; predefined tools only help models with weak bash skills. The post does not name the four models tested or give exact cost figures.

Read the original ↗Export Markdown