Skip to content
Trending storyPast story

xAI releases Grok 4.6, leads on agentic task efficiency

2 reports2 sourcesupdated Aug 13, 2026

What happened

From the coverage

Grok 4.6 在 Grok 4.5 的基础上,专门训练了模型处理长链条任务的能力,比如做研究、分析信息、跨代码库工作,或者把一个想法直接做成能用的应用。在 AA 智能指数这个综合跑分上,它和 GPT-5.6 Sol 打平,都是 61 分;在 DeepSWE 1.1 这个软件工程测试里,分数从 4.5 的 54% 跳到了 65.9%。xAI 说,任务...

From AI HOT 精选

Coverage

Follow the reports to see the story from different sides.

Aug 13
  1. Hacker News front pagePick
    Grok 4.6 matches GPT-5.6 Sol on intelligence, leads on agentic cost efficiency

    Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, a 5-point gain over Grok 4.5, matching GPT-5.6 Sol (max) and trailing only Anthropic's Claude Opus 5 (63) and Claude Fable 5 (62). It shines on agentic tasks: GDPval-AA v2 Elo of 1753, behind only Claude Opus 5; 50.7% on 𝜏³-Banking and 88.4% on Terminal-Bench v2.1, both top-tier. Pricing stays at $2/$6 per 1M input/output tokens, over 60% cheaper than Claude Opus 5 and GPT-5.6 Sol. On AA-Briefcase, a long-horizon knowledge-work benchmark, it scores Elo 1577 (Fable 5-tier) but finishes tasks in ~53 turns and ~0.5B input tokens vs. ~103 turns and ~2.0B for Claude Opus 5, giving it a large cost edge. Context window remains 500k tokens; cache-hit pricing rose from $0.3 to $0.5 per 1M tokens.

Aug 12
  1. AI HOT (Curated Pool)Pick
    xAI releases Grok 4.6, focused on long-running agent capabilities

    Grok 4.6 builds on Grok 4.5 with a focus on long-running agents that can research, analyze, code, or turn an idea into a working app across many steps. It matches GPT-5.6 Sol on the AA Intelligence Index at 61, and jumps from 54% to 65.9% on DeepSWE 1.1. xAI reports the model shows more self-testing and verification on longer trajectories. Pricing is $2/M input tokens and $6/M output tokens, with a fast variant at double the price. Available today in Cursor and Grok Build, with 2x included usage for the first week.