Gary Marcus on GPT-6 Astra: Real progress, but robustness and monitorability are open questions
GPT-6 Astra scores 63% on ARC-AGI-3 and 99% with a provider adapter, while building symbolic world models to solve tasks. Gary Marcus calls the direction vindicating but warns the post doesn't disclose how robust this capability is in open-ended settings. The system also appears less monitorable than prior versions, which raises safety concerns. He cautions against AGI claims until more technical details and independent testing emerge.
Why it matters: Gary Marcus's take on GPT-6 Astra carries built-in narrative weight — the ARC-AGI-3 63%/99% numbers are hard data, and he directly challenges Brockman's AGI framing, hitting all three HKR axes. Score capped below 85 because it's a third-party commentary rather than a first-par...