Skip to content

xAI / Grok

Everything xAI and Grok: model iterations, compute build-out and integration with X.

Latest picks

21–40 of 100

Jul 30Thursday

The Verge · AI

xAI sues to block Minnesota's anti-nudification app law at the last minute

Minnesota's law banning nudification apps is about to take effect, and xAI filed a last-minute lawsuit to block it. xAI argues the law is overbroad and would restrict Grok's image generation, violating First Amendment free speech. In the filing, xAI describes Grok as an opinionated, sarcastic AI assistant whose explicit images are a form of expression. The state attorney general counters that the law only targets non-consensual fake nudes and has nothing to do with free speech. The case has just been filed and hasn't been heard yet.

Why it matters: xAI sues Minnesota over its anti-deepfake-nudity law, tying Grok's image generation to a First Amendment defense — the legal conflict is sharp. Score held back because it's just a filing so far; no ruling yet, so real-world impact is pending.

Jul 21Tuesday

AI HOT (Curated Pool)

Grok for Excel: ask questions, write formulas, and run scenarios in plain English inside the workbook

xAI released a free Microsoft 365 add-in that brings Grok into Excel. Select a range and ask what moved or why—answers cite source cells and charts drop into the sheet. Describe the outcome you want and Grok writes the formula; edits land in the formula bar so you can still tweak them. The add-in can pull context from SharePoint or Google Drive via Grok connectors. It's available now on the Microsoft Marketplace, with Word and PowerPoint versions also listed.

Why it matters: xAI released a free Grok add-in for Excel with in-workbook natural language querying, formula generation, and scenario running. Feature descriptions are concrete (cell citation, editable formulas, external data connectors), but this is day-one announcement with no third-party ...

Jul 19Sunday

Computing Life · Share · Yage

Grok Build open-sourced its client harness, not the model or cloud

xAI released the Rust client harness that handles local files, commands, and permissions for Grok Build under Apache-2.0. The Grok model, cloud services, and the official binary build chain remain closed. The repo doesn't accept external PRs. The commit from the earlier upload controversy isn't in the public history, so the current code can't close that case. The real win: you can now pin a public commit, build it yourself, and compare its behavior against the official binary.

Why it matters: xAI open-sourcing Grok Build's client harness is substantive—Apache-2.0, headless mode, and ACP support go beyond signaling. But the model and build chain remain closed, and the repo rejects PRs, capping it below 85. All three HKR axes hit, so featured.

Jul 17Friday

AI HOT (Curated Pool)

xAI can't deny Grok makes CSAM anymore, so it's suing users

xAI filed its first lawsuit against a Grok user accused of generating child sex abuse images. The company had long claimed such outputs were user-created, but this suit effectively admits the model can be misused to produce illegal content. The post doesn't disclose specific safeguards, model versions, or how many users are affected. This reads more like legal damage control than a technical fix.

Why it matters: xAI's first lawsuit against a user for generating CSAM with Grok amounts to a legal admission that the model can be abused — a pivot from denial to damage control. The article lacks specifics on safety measures, model version, or user count, capping the score below 85. But the...

Jul 16Thursday

The Verge · AI

xAI sues a man for using Grok to generate CSAM deepfakes

xAI filed a federal lawsuit against Terry Harwood, accusing him of bypassing Grok's safeguards to generate CSAM deepfakes. The company claims Harwood used prompt injection and other methods, and is seeking reputational and legal damages. It's a rare case of an AI company proactively suing a user for generating illegal content, though the post doesn't disclose the specific techniques or volume of images produced.

Why it matters: xAI proactively suing a user for bypassing Grok's guardrails to generate CSAM deepfakes is a rare case of an AI company pursuing end-user abuse. Score held back because the post lacks technical specifics and generation volume — strong topic, thin on hard facts.

AI HOT (Curated Pool)

xAI open-sources Grok Build coding agent and terminal UI

xAI released the full Grok Build codebase on GitHub, covering the agent loop, tool dispatch, terminal UI, and extension system. You can read the source to see how context assembly and tool calls work, or compile it yourself and point it at a local inference setup.

Why it matters: xAI open-sourced Grok Build's full codebase — agent loop, TUI, extension system, local-first support. Hits all three HKR axes for the dev audience. Score stays at the featured threshold because we only have the official announcement so far; no third-party benchmarks or hands-o...

Jul 15Wednesday

The Verge · AI

SpaceXAI's Grok coding tool uploaded users' entire codebases to cloud storage

SpaceXAI's Grok coding tool was silently uploading users' entire local repositories to cloud storage by default. Elon Musk responded that all previously uploaded data will be deleted, but the post doesn't say how long this behavior was live or how many users were affected. If you're using it, I'd check whether your repos got synced.

Why it matters: Default full-repo upload is a serious product incident with broad exposure and zero user awareness. The Verge broke it and Musk responded, making it verifiable. Not scoring higher because the post doesn't disclose how long this ran or how many users were affected — key facts a...

Jul 13Monday

AI HOT (Curated Pool)

xAI's Official Grok CLI Caught Silently Uploading Entire Codebase and User Keys

A security researcher found that xAI's Grok CLI silently packages and uploads your entire working directory. Version 0.2.93 of the npm package compresses the codebase into tar.gz files before and after every task, sending them through a separate side channel to xAI's Google Cloud bucket—even when the model replies with a single word. Worse, the uploads also included ~/.claude.json, Claude Code settings, global agent rules, 30+ skill files, and an API key. On July 13, xAI pushed a remote server-side toggle adding a disable_codebase_upload field to turn off the default behavior, but it had been on by default until then. The post doesn't disclose how long this was active or how many users were affected.

Why it matters: Security researcher confirms xAI's official CLI silently uploads entire working directories and key files, with specific version, upload path, and affected file list. All three HKR axes hit. Industry-level incident, importance 92.

Jul 12Sunday

Hacker News front page

Wire analysis: xAI's Grok Build CLI uploads your .env and entire repo to xAI

A packet capture of Grok Build CLI (v0.2.93) shows it uploads the entire project repo to xAI's GCS bucket by default, including plaintext secrets in .env and full git history. Even with a prompt telling the model to reply 'OK' and read no files, the whole repo is still uploaded. On a 12 GB test repo, the storage upload hit 5.10 GiB—roughly 27,800× the model-turn channel data. Disabling 'Improve the model' does not stop the upload.

Why it matters: A wire-level analysis shows Grok Build CLI uploads the entire repo — including plaintext .env secrets and full git history — to xAI's GCS bucket by default, even when the prompt says 'don't read any files.' This is hard evidence on AI coding tool privacy, not speculation. Scor...

Jul 11Saturday

r/LocalLLaMA

Grok Build CLI uploads your entire repo—full git history and .env secrets—to xAI's cloud, and the opt-out doesn't stop it

A Reddit user captured network traffic showing that xAI's Grok Build CLI uploads the entire project repo to the cloud, including full git history and .env secrets. The opt-out setting does not stop the upload. The post body is inaccessible due to a Reddit block, so trigger conditions, affected versions, and xAI's response remain undisclosed.

Why it matters: A security/privacy incident with wire-capture evidence — high credibility and direct relevance to any dev considering Grok Build CLI. Score capped below 85 because the Reddit body is blocked, leaving trigger conditions, affected versions, and xAI's response unknown.

Jul 9Thursday

Latent Space

SpaceXAI launches Grok 4.5, first Opus-class model co-trained with Cursor

SpaceXAI dropped Grok 4.5 one day before GPT-5.6, positioning it as an Opus-class coding and agent model co-trained with Cursor. Musk called it roughly comparable to Opus 4.7 but faster and cheaper—$2/$6 per million tokens, undercutting both GPT-5.6 and Opus 4.8. It's 1.5T parameters, 3x larger than Grok 4.3, with a 500k context window that may return to 1M next week. Cursor says this is their first model built beyond software engineering and offers double usage for the first week. The post doesn't disclose specific benchmark scores; it notes SWE-Bench Pro is now considered saturated by OpenAI's evals team.

Why it matters: SpaceXAI dropped Grok 4.5 a day before GPT-5.6 — the timing alone is a story. 1.5T params, 3x the previous generation, and $2/M input tokens give a clear performance and cost picture. It's Cursor's first post-acquisition move beyond pure coding, which matters directly to agent...

Hacker News front page

Grok 4.5, GPT-5.5, and Claude build the same apps: speed, cost, and quality compared

TryAI gave Grok 4.5, GPT-5.5, Claude Opus 4.8, and Fable 5 the same three app prompts and measured latency and cost. Claude models nailed the 3D Rubik's cube first try; Grok 4.5 needed its one allowed retry after a blank render, and GPT-5.5 only drew a single dark face. All four shipped a working particle sandbox and a playable Breakout game. Grok 4.5 led on speed: 0.44s first token, ~110 tok/s throughput, and the cheapest per reply. Fable 5 was slowest and priciest. The post doesn't disclose parameter counts or training details.

Why it matters: First-hand coding shootout with concrete failure cases and cost data, not just benchmark scores. Score isn't higher because TryAI isn't a tier-1 evaluator and the excerpt only gives a summary — full data requires clicking through.

AI HOT (Curated Pool)

Lawsuit: Man used Grok to make 7K sex images of stepdaughter, then shot himself

A new lawsuit alleges xAI's Grok was used to create over 7,000 child sexual abuse images of the user's stepdaughter. The man later shot himself. xAI reported only one gang-rape prompt to NCMEC and did not report the thousands of other CSAM generations. The suit accuses X and xAI of shielding child predators. The post does not spell out why xAI's safety filters missed the bulk of the images, nor whether Grok's image generation disables real-face simulation by default.

Why it matters: Ars Technica exclusive on a lawsuit revealing a severe gap in xAI's safety reporting: only one prompt flagged out of 7K CSAM images generated. Involves a minor, suicide, and platform liability — all three HKR axes hit. Score capped below 90 because the article doesn't explain ...

Jul 1Wednesday

AI HOT (Curated Pool)

xAI launches Voice Agent Builder beta: build a voice agent in under 2 minutes

xAI packaged Grok Voice into a no-code platform, now in beta as of July 1. You describe the call flow in plain language, upload docs as a knowledge base, and connect tools like calendars or ticketing systems—then you get a working voice agent. It uses a speech-to-speech path instead of chaining ASR→LLM→TTS, which xAI claims cuts latency and failure points. Pricing is $0.05/min of audio plus $0.01/min for a platform-provided number. xAI also published τ-voice Bench scores: Grok Voice Think Fast 1.0 hit 67.3% overall, versus 43.8% for Gemini 3.1 Flash Live and 35.3% for GPT Realtime 1.5. Take the benchmark with a grain of salt—it's xAI's own test, and third-party results aren't out yet.

Why it matters: xAI turned Grok Voice into a no-code platform with voice-to-voice direct pipeline, $0.05/min, and 2-minute setup — all concrete numbers. Not scoring higher because we only have the official announcement so far, no third-party testing or head-to-head comparisons; scoring 78 as ...

Jun 28Sunday

AI HOT (Curated Pool)

Grok 4.5 enters private testing at SpaceX and Tesla, performance near Opus

Elon Musk says Grok 4.5 is built on a 1.5T-parameter V9 base model with Cursor data added during supplementary training, now in private testing at SpaceX and Tesla. Early evals show performance close to or possibly exceeding Opus. RL is still improving the model, and the Grok Build toolchain is maturing. SpaceX will also release a fully from-scratch trained model every month this year. The post doesn't specify which Opus model, benchmarks, or testing scale.

Why it matters: Musk's own tease of Grok 4.5 vs Opus with Cursor data injection is strong signal. But no benchmark names, Opus version, or sample size disclosed — caps at 78.

AI HOT (Curated Pool)

SpaceX files SpaceXAI trademark, Musk says xAI will merge into SpaceX

SpaceX has filed a trademark for SpaceXAI. Musk says xAI will dissolve as a standalone company and become SpaceX's AI product line. The post doesn't disclose a merger timeline, team structure, or what happens to existing xAI products like Grok.

Why it matters: xAI folding into SpaceX is a structural shift, not a routine product update. The trademark filing provides hard evidence, but the post lacks timeline and team integration details, capping the score below 85.

Jun 24Wednesday

Hacker News front page

Reid Hoffman: SpaceX is 'not an AI company,' xAI is a 'complete train wreck'

LinkedIn co-founder Reid Hoffman called SpaceX 'not an AI company' and xAI 'a complete train wreck' on a podcast. He said SpaceX's post-IPO Cursor acquisition is buying relevance, and its compute leasing is just 'a premium-priced CoreWeave.' On xAI, all 11 co-founders have left and the company is on its third restart. Hoffman also criticized the U.S. government's forced takedown of Anthropic's Fable and Mythos models as 'autocratic willy-nilly,' troubled by the asymmetry with OpenAI. He invests in both Anthropic and OpenAI and sees room for both to win.

Why it matters: A well-known investor publicly sizes up the AI landscape with specific claims and sharp language — all three HKR axes hit. Deduction because it's a podcast opinion, not a product/research release with verifiable new capabilities, so it lands at the 78 featured threshold.

Jun 22Monday

AI HOT (Curated Pool)

Grok Build adds /goal mode for long-running autonomous task execution

xAI added /goal to Grok Build: give the agent an objective and it plans, breaks work into a checklist, and executes until done. You can check status, pause, resume, or clear the goal mid-run. The post doesn't disclose max run time, resource costs, or specific pricing.

Why it matters: xAI added /goal mode to Grok Build, letting the agent autonomously complete a task — similar in shape to Cursor Agent and Claude Code's long-running execution. Concrete interaction details are present, but the post doesn't disclose max runtime, resource consumption, or extra p...

Jun 19Friday

Latent Space

Anjney Midha on AI compute waste: frontier labs run sub-10% MFU, AMP plans an independent compute grid

Anjney Midha discusses hidden AI infrastructure waste on Latent Space. xAI's training MFU is under 10%, while Google treated 95% utilization as an outage; best-in-class today is 60–70%. He invested in Anthropic, Mistral, and Black Forest Labs, and now runs AMP, aiming for a 1.2 GW base-load compute grid with 6 GW spike capacity. He also flags DeepMind's unpublished research as a market failure and notes Anthropic prioritized coding as P0 from day one. The post does not disclose a timeline for AMP's grid.

Why it matters: Anjney Midha puts a specific number on xAI's training MFU (under 10%), turning vague complaints about compute waste into a quantifiable discussion. Score capped at 78 because it's a podcast opinion, not a product launch or paper — no reproducible verification path.

Jun 18Thursday

Hacker News front page

OpenRouter ran 11 LLMs in a 30-game battle royale — Grok 4.1 Fast won 43%

OpenRouter's Jacky Liang dropped 11 LLMs into a 2D battle royale for 30 matches. Grok 4.1 Fast won 13 games at $0.97 per win; Claude Sonnet 4.6 won 5 at $26.78 per win — a 27x gap. GPT 5.4 had the most kills (38) but only 2 wins, so killing more didn't mean winning more. GPT 5.4-mini, DeepSeek 4 Flash, and Kimi K2.6 spent $57 combined and won zero games. The models reasoned, called tools, and updated memory each turn — they weren't just generating control code. The post doesn't provide the full leaderboard or detailed behavioral differences across all models.

Why it matters: OpenRouter's official blog, author Jacky Liang ran 30 games himself with full data and replays. Grok 4.1 Fast's cost advantage is stark, Claude Sonnet 4.6 is expensive but consistent, GPT 5.4 is the kill leader but can't close — all three takeaways are concrete and verifiable....