Grok 4.7 hits the frontier on agentic knowledge work, Coding Agent Index reaches 56
Artificial Analysis 评测 Grok 4.7:智能体知识工作跻身前沿,编码代理得分升至 56
Grok 4.7 scores 1657 Elo on AA-Briefcase, up 111 points from Grok 4.6, landing just behind Claude Opus 5 and Claude Fable 5.1 on long-horizon agentic knowledge work. Its Coding Agent Index jumps from 47 to 56, with DeepSWE rising from 65% to 73% and Terminal-Bench doubling to 33%. The gains come at a cost: 81k output tokens per task on average, nearly 3× what GPT-6 Astra uses. Pricing stays at $2/$6 per 1M input/output tokens, context window unchanged at 500k. Hallucination rate drops from 34% to 29%, accuracy is flat.
Why it matters: Grok 4.7 reaches the frontier tier on agentic knowledge work and coding agents, with 1657 Elo on AA-Briefcase and 56 on the Coding Agent Index — both directly comparable numbers. Not scored higher because gains outside agent tasks are incremental and the body is truncated, lea...