Skip to content

xAI / Grok

Everything xAI and Grok: model iterations, compute build-out and integration with X.

Latest picks

1–20 of 100

Sep 28Monday

AI HOT (Curated Pool)

xAI launches Team Bots: Grok agents that learn and work alongside your team

xAI turned Grok Bot into shared AI coworkers. You give a Team Bot files, app access, and credentials, then the whole team works from the same context. SpaceXAI already uses them: a sales Bot posts daily account briefings in Slack, an engineering Bot coordinates PR reviews and bug fixes, and a marketing Bot checks drafts against brand guidelines and ships website updates directly. Each Bot remembers team decisions so knowledge stays when people rotate. Harper Insurance built one in 24 hours to recover lapsed policies, saving customers over $120,000. The post doesn't disclose pricing or a public launch date.

Why it matters: xAI turns Grok Bot into a shared team agent with persistent memory and concrete deployment examples, not vaporware. But only one customer (SpaceXAI) is named, and pricing/availability aren't spelled out, so it stays at 78.

Sep 26Saturday

AI Chat-Group Daily (群聊日报)

OpenAI Codex code confirms Pro Max pricing; Astra 3D printing pipeline works end-to-end

An OpenAI Codex repo commit reveals Pro Max at $600/month ($500 pre-tax), with three clear tiers: $100 Lite, $200 Pro, $500 Max. DevDay next Tuesday is the likely launch. The group also spotted an unlisted model name: gpt-6.1-astra-max. Separately, multiple users verified Astra's end-to-end 3D printing pipeline—from verbal modeling and watertightness checks to driving Bambu Studio directly. One printed a play supermarket; another printed a phone stand that couldn't hold a phone. On Terminal-Bench-Science 0.1, GPT-6 Astra leads at 63.3%, but Opus 5.5 xhigh trails by under two points at significantly lower cost. xAI disclosed full Colossus cluster specs for the first time. Microsoft launched Copilot Code to compete with Codex and Claude Code. Meta released Horizon Create and Studio for AI game creation.

Why it matters: Code-level confirmation of Pro Max tier in OpenAI's Codex repo, with clear three-tier pricing and an unlisted model name. Source is a chatgroup daily, not an official announcement, so capped below 85. But the DevDay countdown + pricing leak combo is enough to make paying users...

Sep 24Thursday

AI HOT (Curated Pool)

OpenAI says its ChatGPT deal with Apple fell far short of expectations

OpenAI stated in court filings that its 2024 deal to integrate ChatGPT into Apple Intelligence underperformed significantly. iPhone user uptake was weak from the first month, and by summer 2025 OpenAI confirmed the integration fell far short of forecasts, cutting weekly active user estimates. The relationship soured afterward; Apple switched to Google Gemini for a rebuilt Siri AI in January 2026. The filings emerged from an antitrust suit by xAI. OpenAI argued the Apple deal did not boost its market position and coincided with a share decline against Google, Anthropic, Meta, and Grok.

Why it matters: OpenAI's court filing self-reports the Apple deal as a flop — first official confirmation with concrete details: slow first-month growth, downward-revised WAU forecasts, and a summer 2025 acknowledgment of no real benefit. HKR all hit, but the info comes from a legal filing ra...

Sep 21Monday

AI HOT (Curated Pool)

xAI launches Grok 4.7, twice as fast as Grok 4.6 at the same price

Grok 4.7 uses a larger base model and a longer RL run on harder, multi-hour tasks. It scores 46.3% on CursorBench 4.0, ahead of GPT-5.6 Sol Max (41.7%) but behind Fable 5.1 Max (51.8%). Pricing stays at $2/$6 per million input/output tokens, same as Grok 4.6, with double the speed. Safety stack is new: only 3.3% of risky cyber prompts get through, and it hits 62.4% on LatchBio's biosafety benchmark. Available today in Cursor, Grok Build, and the API.

Why it matters: xAI drops Grok 4.7 targeting coding and knowledge work, hitting 46.3% on CursorBench 4.0 — above GPT-5.6 Sol Max but behind Fable. Concrete benchmark and training details clear all three HKR axes. Held below 85 because the post doesn't disclose model size, architecture changes...

Sep 16Wednesday

AI HOT (Curated Pool)

Grok Build adds memory that carries project conventions and decisions across sessions

Grok Build now writes project conventions, decisions, and facts in the background and reads them back in later sessions. It captures durable details like team code style and test commands, skipping transient state and secrets. /memory browses all notes, and /dream organizes them into topic files. The feature is live for new sessions.

Why it matters: Grok Build's memory isn't just session history — it auto-extracts project conventions and proactively applies them in later sessions, with /memory for browsing and /dream for organizing. This is a step beyond Cursor's Rules in automation, but it's fresh out the gate and only w...

Sep 10Thursday

Computing Life · Share · Yage

Cloud agents aren't new—custody is

Meta Muse, xAI Grok Bot, and Manus Cloud Computer all gave agents a persistent cloud desktop within months. The post traces a four-generation shift from chat window to always-on home, arguing that personal agents need a place to keep logins, files, and habits. The real variable isn't cloud vs. local—it's who holds custody of that operational state, which shapes lock-in, maintenance burden, and the subscription model behind it.

Why it matters: Three independent vendors converging on the same architecture — persistent cloud VMs for personal agents — within months is a genuine signal. The piece connects Manus My Computer → Cloud Computer → Grok Bot → Muse into a clean evolution line, not isolated reporting. Deduction:...

Sep 4Friday

AI Chat-Group Daily (群聊日报)

GPT-6 Astra launch day saw OpenAI, Anthropic, and xAI all go down; Cerebras launched Qwen 3.8 27B inference

OpenAI released GPT-6 Astra with 99.9% on ARC-AGI-3, but most paid users couldn't access it on launch day. Tibo announced daily banked reset compensation, which users actually welcomed. OpenAI, Anthropic, and xAI all experienced outages around the launch, leaving Gemini briefly as the only available model in North America. Cerebras launched Qwen 3.8 27B inference the same day, hitting 1,806 tok/s in real tests. Zhipu ZCode started a 15-day free promotion. The group also discussed Mac M5 Max local inference bottlenecks, DSH's unstable dev experience, and the real makeup of 10x automation gains—mostly from tooling improvements, not full automation.

Why it matters: GPT-6 Astra launch is the day's biggest story, with ARC-AGI-3 hitting 99.9% as a striking number. But the source is a curated chat digest, not a primary report — high signal density but lower authority, so 78 featured rather than p1.

AI HOT (Curated Pool)

xAI set Grok Bot loose on procurement — Haggle Bot found over $100K in direct savings

xAI built an internal procurement agent called Haggle Bot on Grok Bot, giving it access to vendor spend, contracts, and usage data. It has already identified over $100,000 in direct savings by flagging unused SaaS seats, negotiating renewals, and shopping around for office supplies. xAI published the full system prompt, which hardcodes permission lines, negotiation anchors, and a strict 'strong finding' standard — every recommendation must cite live spend data, a specific savings mechanism, and the next step already taken. Grain of salt: this is xAI's own case study with no third-party verification, but the prompt's constraints on evidence and decision authority are concrete and reusable.

Why it matters: xAI published the full prompt and a $100K savings case for an internal procurement agent — concrete numbers and design details make it a strong reference for enterprise agent builders. Not scored higher because it's a single-company experiment, not a reproducible product or op...

Sep 3Thursday

The Verge · AI

ChatGPT, Grok, and Claude all went down at the same time on Thursday

Around 11AM ET Thursday, ChatGPT, Grok, and Claude all started having issues at roughly the same time. ChatGPT returned errors across chat, login, file uploads, voice, search, deep research, and image generation; its status page cited elevated errors for ChatGPT and Codex. Anthropic's Claude chatbot and Claude Code were also affected. The post doesn't detail Grok's specific symptoms, the recovery timeline for each service, or whether the outages share a root cause.

Why it matters: A simultaneous outage across ChatGPT, Grok, and Claude is a rare event that directly disrupts workflows for a huge user base. Missing root cause and recovery timeline keeps it from 95+, but the topic is strong enough for featured.

Hacker News front page

OpenAI, Claude, and Grok all went down at once—users suspect a Cloudflare cascade

A Hacker News thread noted that OpenAI, Claude, and Grok all went down around the same time. Users pointed to Downdetector spikes for Cloudflare, Azure, AWS, and Google Cloud near 7:30, suspecting a cascade from Cloudflare or another shared dependency. Other guesses include user migration overload and deliberate attack, but the post is community speculation—no official root cause is confirmed.

Why it matters: Simultaneous outage across OpenAI, Claude, and Grok with high HN engagement. Downdetector data points to Cloudflare or shared infra as a possible common cause. The event is conversation-worthy but lacks a confirmed root cause, so it lands at the 78 featured threshold rather th...

AI HOT (Curated Pool)

xAI launches Grok Bot for Enterprise, free for Grok and Cursor Enterprise customers for two weeks

xAI brings Grok Bot to enterprises. Each Bot runs as an isolated cloud worker that can use apps and websites like a person. You teach it a workflow once, and it runs autonomously after that. Bots can message each other and share context. The enterprise release adds access, network, and audit controls. The post lists five use cases—sales, recruiting, marketing, finance, and engineering—with a finance Bot surfacing tens of thousands of dollars in savings across SaaS and recurring purchases. Grok and Cursor Enterprise customers get free access for two weeks and can invite their whole org, including people without a seat. The post does not disclose pricing after the two-week window.

Why it matters: xAI launched Grok Bot for enterprises with access, network, and audit controls, plus a two-week free trial for Grok and Cursor Enterprise users. The product goes beyond standard chatbots, but the post lacks pricing and named customer examples, capping the score below 85.

AI HOT (Curated Pool)

xAI unveils Grok Bot design: moving AI from a chat window to persistent agents that work on their own

On Sep 3, xAI shared the design philosophy behind Grok Bot. The core shift is treating Bots—not chat sessions—as the primary object. Each Bot has its own name, avatar, memory, and tools, remembers past conversations, and can keep working without the user watching. The sidebar becomes a roster of Bots with presence indicators, not a list of disposable chats. The post does not disclose a launch date or pricing.

Why it matters: xAI published an official design piece on Grok Bot, positioning bots as persistent contacts with their own computer and offline work capability. Directly useful for agent product builders, but it's a design philosophy post rather than a feature launch, so it lands at the 72 fe...

Aug 19Wednesday

AI HOT (Curated Pool)

Grok Build is now available on web and mobile for all plans

xAI moved Grok Build from Early Beta to general availability. Any plan user on web, iOS, or Android can now describe an app, game, or dashboard in natural language and get a working version live in chat. Published apps get a grok.me link, support custom domains, can export to GitHub, and can call Grok's APIs for chat, images, and voice. Shared apps render as inline cards on X. The post shows four real published examples: a forest driving game, an isometric city builder, a 3D physics playground, and a browser beat machine.

Why it matters: xAI pushed Grok Build from paid beta to full launch with publishing, custom domains, and API access — a real product step. Not scoring higher because only the official announcement is available, with no third-party testing or specific limitations disclosed.

Aug 16Sunday

TechCrunch · AI

Woman claims stepfather used Grok to turn her childhood photo into 7,000+ explicit images

A woman, Jane Doe 4, joined a class-action lawsuit against xAI, alleging her stepfather used Grok to create over 7,000 explicit images from a photo taken when she was 11. Her stepfather died by suicide two days after a law enforcement raid uncovered the images. Three Tennessee teenagers had previously sued xAI, claiming Grok lacked basic safeguards to prevent generating explicit imagery of real people, including minors. Earlier this year, X was flooded with millions of Grok-generated sexualized images. TechCrunch has reached out to xAI; the post does not include a response.

Why it matters: This isn't a product update or a paper — it's a new filing in a class-action lawsuit that pins Grok's safety gaps to a horrifyingly specific case. 7,000+ images, an age-11 source photo, a suicide — every detail forces the question of where xAI's content moderation line actuall...

Aug 14Friday

AI HOT (Curated Pool)

Cursor has been acquired by SpaceX, gaining access to the world's largest GPU fleet

Cursor announced it has been acquired by SpaceX, closing a deal that began in April. The acquisition gives Cursor access to SpaceX's massive GPU fleet to build stronger, cheaper-to-run models. Grok 4.6, released Wednesday, is the first preview of what the combined effort can produce. The team says the product direction stays the same: help people write less code and solve harder problems.

Why it matters: Cursor's acquisition by SpaceX is one of the biggest structural moves in AI tooling this year. The deal was in talks since April and just closed; Cursor now gets direct access to SpaceX's GPU cluster, and Grok 4.6 already shipped as the first post-merger preview. The team says...

Aug 13Thursday

Latent Space

xAI drops Grok 4.6 and Grok Bot, a strong new entrant in the AI teammate race

xAI launched Grok 4.6 and the Grok Bot early beta. Grok Bot logs into your tools, operates them like a human, and returns finished work—positioned as an AI teammate. The 1.5T-parameter Grok 4.6 scores near GPT-5.6 Sol Max on the AA-Briefcase knowledge-work benchmark but costs far less: $2/M input tokens, $6/M output. Training reused Grok 4.5 to regenerate SFT traces and added agentic RL across coding, web, CAD, and kernel optimization. Elon says Grok 4.7 is already training. The same day, Qwen3.8-Max dropped as open weights: a 2.4T total / 95B active MoE.

Why it matters: Grok 4.6 matches GPT-5.6 Sol Max on a knowledge-work benchmark at an order-of-magnitude lower price, while the simultaneously launched Grok Bot enters the AI teammate race built by the ex-Cursor team with positive early feedback. Score isn't higher because the Bot is still in ...

Aug 12Wednesday

AI HOT (Curated Pool)

xAI releases Grok 4.6, focused on long-running agent capabilities

Grok 4.6 builds on Grok 4.5 with a focus on long-running agents that can research, analyze, code, or turn an idea into a working app across many steps. It matches GPT-5.6 Sol on the AA Intelligence Index at 61, and jumps from 54% to 65.9% on DeepSWE 1.1. xAI reports the model shows more self-testing and verification on longer trajectories. Pricing is $2/M input tokens and $6/M output tokens, with a fast variant at double the price. Available today in Cursor and Grok Build, with 2x included usage for the first week.

Why it matters: xAI releases Grok 4.6 with a focus on long-running agents, matching GPT-5.6 Sol on the AA Intelligence Index and showing a clear jump on DeepSWE. This is a substantive update from a major lab with concrete benchmarks and a direct competitor comparison, earning featured. Not sc...

Hacker News front page

xAI launches Grok Bot: an AI teammate that signs into your tools and finishes work

xAI released an early beta of Grok Bot, positioned as an AI teammate that operates browsers and apps, not just a chat assistant. You assign it tasks; it signs into tools like Zendesk, clicks through workflows, and returns with finished work. Multiple bots run in parallel, hand off tasks to each other, and retain context and preferences. Pricing: Cursor Ultra at $200/month for individuals, Cursor Premium Teams at $120/seat/month. The post does not disclose the underlying model, available regions, or any quality benchmarks.

Why it matters: xAI launches Grok Bot — an AI teammate that operates browsers and apps directly, with multi-bot parallelism and task handoff. Personal plan at $200/month. Product shape is more concrete than most agent offerings, but macOS-only early beta with no reliability data yet — scores 82.

Aug 4Tuesday

Bloomberg Technology

Big AI bets are splitting venture capital, leaving smaller funds behind

Bloomberg maps how AI's capital intensity is concentrating power among mega-funds. Rounds for OpenAI, Anthropic, and xAI now run into tens of billions, playable only by Tiger Global, SoftBank, and a16z. Smaller funds are locked out of the best deals and pushed into seed or niche apps. LPs and GPs quoted say the traditional spray-and-pray VC model breaks when AI demands so much cash and returns cluster in so few names. The piece is a trend sketch—it doesn't give hard failure rates or return comparisons for small funds.

Why it matters: Bloomberg's trend piece lays out the structural split in AI fundraising clearly: $10B+ rounds are only for Tiger Global, SoftBank, a16z, and smaller funds are getting squeezed out. HKR all hit, but it's a feature sketch rather than hard news—no new data point or exclusive scoo...

Jul 31Friday

Hacker News front page

SWE-rebench leaderboard: 13 models and 4 agents benchmarked on real-world bug fixes across Go, Java, Python, Rust, and TypeScript

Nebius built this benchmark using 111 real GitHub issues from 65 repos across Go, Java, Python, Rust, and TypeScript. Fable 5 leads with a 64.5% resolved rate at $4.40 per problem. Grok 4.5 and Opus 5 both hit above 63%, but Grok 4.5 costs only $1.47 per problem—much cheaper. Among agents, Junie scores 61.8% at $0.81, while Claude Code gets 60.4% at $3.39. DeepSeek-V4 Pro resolves 40.2% at just $0.15 per problem, the cheapest in the top 14. The post does not break down per-language performance or explain why many models—from Claude Opus 4.1 through Sonnet 4.6—are listed as N/A.

Why it matters: Nebius built this benchmark from 111 real GitHub issues across 65 repos, mixing models and agents with transparent cost data. All three HKR axes hit, but it's a third-party eval, not a model release—caps below 85.