Skip to content
AI HOT (Curated Pool)

GPT-6 Astra benchmarks clash, but its human-beating efficiency on ARC-AGI-3 pulls Chollet's AGI forecast forward

GPT-6 Astra 基准表现分歧,ARC-AGI-3 效率超人类令 Chollet 提前 AGI 预测

GPT-6 Astra gets contradictory scores: Epoch AI ranks it first, while Artificial Analysis says it ties the previous model. The real signal is ARC-AGI-3, where Astra hits 62.7% in unfamiliar game worlds—up from Sol's 7.8%—and for the first time beats average human efficiency. ARC Prize's François Chollet says progress is about 2x faster than he expected and is moving his AGI timeline forward. Astra also solved 2 open Erdős math problems at $300 per attempt, and its hallucination rate dropped from 92% to 51%, though it lost ground on long-context reasoning and some coding tests.

Why it matters: GPT-6 Astra beat human efficiency on ARC-AGI-3 for the first time, and Chollet moved his AGI forecast forward — that's a hard signal. The split between Epoch AI and Artificial Analysis rankings adds narrative tension. Not scoring higher because the post only gives the 62.7% fi...

Read the original ↗Export Markdown