a16z answers with data: Can agents really use a computer yet?
智能体真的会用电脑吗?a16z 用数据给出答案
a16z's Fabrizio Serafini, Seema Amble, and Eric Zhou track computer-use agents on the OSWorld-Verified benchmark. A year ago the best model scored ~30%; now Claude Fable 5 hits 85%, above the human baseline of 72%. The post argues the model is no longer the main bottleneck—the frontier is shifting from 'can the agent use a computer?' to 'can it reliably do this job inside a real company,' covering permissions, process knowledge, error handling, and caching. Production deployments exist for standardized back-office work, but agents still break when tasks drift off the runbook and costs don't work everywhere.
Why it matters: a16z's OSWorld-Verified data makes a clear case that agent capability has crossed the human baseline. Held at 82 because it's a VC blog, not a product launch, and the post doesn't quantify real-world reliability yet.