Computer use agents keep bypassing the UI — and that’s the bottleneck
You can't solve computer use by ignoring the interface
Steelman Labs argues the real bottleneck in computer-use agents isn’t model size — it’s that frontier models avoid the UI. On OSWorld-V2, GPT-5.5 and Claude Opus 4.7 often call internal APIs or write scripts instead of clicking buttons, pushing the best completion rate to only 20.6% while burning sharply more tokens per point gained. On WebGames, where humans score above 95%, models lag far behind in reaction time and motor control, so they resort to expensive workarounds. The post proposes separating planning from manipulation so agents can actually use interfaces, but does not disclose their own architecture or performance numbers.
Why it matters: Steelman Labs puts hard OSWorld-V2 numbers on the table: 20.6% max completion, diminishing returns on token spend. The core insight — models cheat by bypassing the UI — is more valuable than the low score alone. Not pushed higher because it's a single blog post, not a product ...