Schema harness pushes frontier models to ~99% on ARC-AGI-3 Public
Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
Impossible Research released Schema, a harness that gets Claude Opus 4.8 and GPT-5.6 Sol to 99% and 95.35% RHAE on the ARC-AGI-3 Public set. It doesn't touch model weights. Instead, it makes models act like physicists: turn raw grid observations into an executable state program, discover the game's mechanism, and keep both in one editable program so a failed prediction can revise the state definition itself. Both scores are self-reported and not yet verified by ARC Prize. The post does not disclose inference latency or per-run cost.
Why it matters: 99% on ARC-AGI-3 is a hard number, and the method—a scheduling harness rather than a new model—has direct implications for agent design. Downside: only a project page is available, no full paper yet, so mechanism details need verification.