Skip to content
Trending storyPast story

Grok 4.7 hits the frontier on agentic knowledge work, Coding Agent Index reaches 56

1 report1 sourceupdated 8 days ago

What happened

Summary

Grok 4.7 在模拟真实工作任务的 AA-Briefcase 测试中拿到 1657 分,比上一代高出 111 分,分析质量提升明显,但呈现质量反而微降,总分排在 Claude Opus 5 和 Claude Fable 5.1 之后。编码代理指数从 47 跳到 56,其中 Terminal-Bench 成绩翻倍到 33%,DeepSWE 从 65%...

Coverage

Follow the reports to see the story from different sides.

Sep 22
  1. AI HOT (Curated Pool)Pick
    Grok 4.7 hits the frontier on agentic knowledge work, Coding Agent Index reaches 56

    Grok 4.7 scores 1657 Elo on AA-Briefcase, up 111 points from Grok 4.6, landing just behind Claude Opus 5 and Claude Fable 5.1 on long-horizon agentic knowledge work. Its Coding Agent Index jumps from 47 to 56, with DeepSWE rising from 65% to 73% and Terminal-Bench doubling to 33%. The gains come at a cost: 81k output tokens per task on average, nearly 3× what GPT-6 Astra uses. Pricing stays at $2/$6 per 1M input/output tokens, context window unchanged at 500k. Hallucination rate drops from 34% to 29%, accuracy is flat.