Grok 4.7 hits the frontier on agentic knowledge work, Coding Agent Index reaches 56
What happened
Grok 4.7 在模拟真实工作任务的 AA-Briefcase 测试中拿到 1657 分,比上一代高出 111 分,分析质量提升明显,但呈现质量反而微降,总分排在 Claude Opus 5 和 Claude Fable 5.1 之后。编码代理指数从 47 跳到 56,其中 Terminal-Bench 成绩翻倍到 33%,DeepSWE 从 65%...
Coverage
Follow the reports to see the story from different sides.
- AI HOT (Curated Pool)PickGrok 4.7 hits the frontier on agentic knowledge work, Coding Agent Index reaches 56
Grok 4.7 scores 1657 Elo on AA-Briefcase, up 111 points from Grok 4.6, landing just behind Claude Opus 5 and Claude Fable 5.1 on long-horizon agentic knowledge work. Its Coding Agent Index jumps from 47 to 56, with DeepSWE rising from 65% to 73% and Terminal-Bench doubling to 33%. The gains come at a cost: 81k output tokens per task on average, nearly 3× what GPT-6 Astra uses. Pricing stays at $2/$6 per 1M input/output tokens, context window unchanged at 500k. Hallucination rate drops from 34% to 29%, accuracy is flat.