Skip to content
AI HOT (Curated Pool)

GPT-6 Astra hits 99% on ARC-AGI-3; Greg Brockman says the benchmark is saturated

Greg Brockman 转发:GPT-6 Astra 在 ARC-AGI-3 达到 SOTA,基准趋于饱和

OpenAI's GPT-6 Astra scored 99% on ARC-AGI-3, beating human performance on 96% of tasks. The standard harness gave only 63%; a new Provider Adapter harness pushed it to 99%. Higher reasoning tiers cost less because Astra solves tasks in fewer actions, cutting model calls and tokens. Greg Brockman reposted the result and said the benchmark is saturated.

Why it matters: GPT-6 Astra's 99% on ARC-AGI-3 is a real industry event, amplified by Greg Brockman's repost. Not a 95 because the score depends on the Provider Adapter framework rather than the default run, and the benchmark itself is nearing saturation—future differentiation is in question.

Read the original ↗Export Markdown