Skip to content
Hacker News front page

Grok 4.6 matches GPT-5.6 Sol on intelligence, leads on agentic cost efficiency

SpaceXAI's Grok 4.6 Scores 61 on the Artificial Analysis Intelligence Index

Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, a 5-point gain over Grok 4.5, matching GPT-5.6 Sol (max) and trailing only Anthropic's Claude Opus 5 (63) and Claude Fable 5 (62). It shines on agentic tasks: GDPval-AA v2 Elo of 1753, behind only Claude Opus 5; 50.7% on 𝜏³-Banking and 88.4% on Terminal-Bench v2.1, both top-tier. Pricing stays at $2/$6 per 1M input/output tokens, over 60% cheaper than Claude Opus 5 and GPT-5.6 Sol. On AA-Briefcase, a long-horizon knowledge-work benchmark, it scores Elo 1577 (Fable 5-tier) but finishes tasks in ~53 turns and ~0.5B input tokens vs. ~103 turns and ~2.0B for Claude Opus 5, giving it a large cost edge. Context window remains 500k tokens; cache-hit pricing rose from $0.3 to $0.5 per 1M tokens.

Why it matters: Grok 4.6 ties GPT-5.6 Sol on the Intelligence Index with standout agentic scores and a clear cost edge. Not scoring higher because this is a third-party benchmark, not an official launch, and the one-month gap from Grok 4.5 warrants more independent confirmation.

Read the original ↗Export Markdown