Opus 4.8’s useful move is Claude Code writing an orchestration script before launching parallel subagents. That order matters. Anthropic is not proving free-form multi-agent swarms work; it is turning task decomposition, dependencies, and checks into a deterministic wrapper around smaller agent loops.
The evidence is messy in a familiar way. Simon Willison calls 4.8 modest but useful, mainly because it admits uncertainty and catches more flaws in its own code. Every says it jumps from 4.7 and competes with GPT-5.5 on an internal senior-engineer benchmark. Datacurve puts it below GPT-5.5, barely above 5.4, while using far more tokens. The ARC-AGI-3 claim says it triples 5.5’s score, but the harness is doing too much work here. I’d trust the Claude Code workflow change before I trust the leaderboard flex.