Models keep absorbing the agent harness — what's left will manage human attention, not the model
The Evolution of the Agent Harness
Dan McAteer traces the tug-of-war between agent harnesses (tools, memory, guardrails outside model weights) and model capability. ReAct in late 2022 was a paper loop; AutoGPT in spring 2023 handed models autonomy they couldn't handle — 95% per-step reliability over 20 steps yields ~36% success. Cursor and Copilot pulled the harness back below the model curve by keeping humans in the loop. The curves inverted when o1 reasoning models arrived in late 2024, and Claude Code in February 2025 made them truly cross. The thesis: models will keep absorbing harness functions into their weights, engineers will delete what gets absorbed, and the remaining harness will manage human attention rather than the model. The post does not provide a timeline or product roadmap.
Why it matters: Dan McAteer uses concrete reliability math to trace the agent harness evolution with a sharp, original angle. Score stays at 78 because this is a commentary piece, not a product launch or first-party release—the signal is in the framing, not in breaking news.