Muse 出事后下载仍涨,OpenAI 发布常驻代理 Dots:常驻 AI 代理的安全账没人买单
2026 年 9 月 29 日 OpenAI DevDay 上,奥特曼发布常驻个人代理 Dots,运行在云端电脑上,可访问 4000 多个应用,面向 Pro 和 Business Premium 用户免费发放。
2026 年 9 月 29 日 OpenAI DevDay 上,奥特曼发布常驻个人代理 Dots,运行在云端电脑上,可访问 4000 多个应用,面向 Pro 和 Business Premium 用户免费发放。
OpenAI 发布常驻 AI 智能体 Dots 后,网友发现域名 dot.com 归 Elon Musk 的 xAI 所有,目前跳转至 Grok 聊天应用下载页。Whois 记录显示该域名于 7 月刚完成转移,xAI 未持有 bot.com、vot.com 等同类拼写域名。
SpaceXAI 的 AI 百科 Grokipedia 在停更数月后恢复更新文章,此前 Lawfare 报道称其页面自 4 月起未再审核编辑。目前奥巴马页面显示两天前经 Grok 事实核查,马斯克页面在撰写期间被核查,但檀香山页面仍停留在 7 个月前。X 与 SpaceXAI 设计负责人 Benji Taylor 上周称 v0.2 将比以往更好。
Latent Space 的 AINews 汇总 9/24-9/25 动态,指出本周发布的 Claude Opus 5.5 在讲解视频生成上表现突出,并以 88.4% 领跑 SimpleBench。
xAI turned Grok Bot into shared AI coworkers. You give a Team Bot files, app access, and credentials, then the whole team works from the same context. SpaceXAI already uses them: a sales Bot posts daily account briefings in Slack, an engineering Bot coordinates PR reviews and bug fixes, and a marketing Bot checks drafts against brand guidelines and ships website updates directly. Each Bot remembers team decisions so knowledge stays when people rotate. Harper Insurance built one in 24 hours to recover lapsed policies, saving customers over $120,000. The post doesn't disclose pricing or a public launch date.
Why it matters: xAI turns Grok Bot into a shared team agent with persistent memory and concrete deployment examples, not vaporware. But only one customer (SpaceXAI) is named, and pricing/availability aren't spelled out, so it stays at 78.
An OpenAI Codex repo commit reveals Pro Max at $600/month ($500 pre-tax), with three clear tiers: $100 Lite, $200 Pro, $500 Max. DevDay next Tuesday is the likely launch. The group also spotted an unlisted model name: gpt-6.1-astra-max. Separately, multiple users verified Astra's end-to-end 3D printing pipeline—from verbal modeling and watertightness checks to driving Bambu Studio directly. One printed a play supermarket; another printed a phone stand that couldn't hold a phone. On Terminal-Bench-Science 0.1, GPT-6 Astra leads at 63.3%, but Opus 5.5 xhigh trails by under two points at significantly lower cost. xAI disclosed full Colossus cluster specs for the first time. Microsoft launched Copilot Code to compete with Codex and Claude Code. Meta released Horizon Create and Studio for AI game creation.
Why it matters: Code-level confirmation of Pro Max tier in OpenAI's Codex repo, with clear three-tier pricing and an unlisted model name. Source is a chatgroup daily, not an official announcement, so capped below 85. But the DevDay countdown + pricing leak combo is enough to make paying users...
Lightspeed is raising $250M for a new India fund focused on early-stage AI. For the first time, it aligns the India fund's cycle with its global funds and shortens the investment period. Lightspeed already backs Anthropic, xAI, and Databricks, and in India it invested in Sarvam AI, which the government picked to build sovereign LLMs. The post doesn't specify sector or stage details beyond early-stage AI.
OpenAI stated in court filings that its 2024 deal to integrate ChatGPT into Apple Intelligence underperformed significantly. iPhone user uptake was weak from the first month, and by summer 2025 OpenAI confirmed the integration fell far short of forecasts, cutting weekly active user estimates. The relationship soured afterward; Apple switched to Google Gemini for a rebuilt Siri AI in January 2026. The filings emerged from an antitrust suit by xAI. OpenAI argued the Apple deal did not boost its market position and coincided with a share decline against Google, Anthropic, Meta, and Grok.
Why it matters: OpenAI's court filing self-reports the Apple deal as a flop — first official confirmation with concrete details: slow first-month growth, downward-revised WAU forecasts, and a summer 2025 acknowledgment of no real benefit. HKR all hit, but the info comes from a legal filing ra...
Elon Musk cites Artificial Analysis to claim Grok 4.7 ranks xAI third in agentic coding, behind only Anthropic and OpenAI. The post doesn't disclose the benchmark's metrics, scores, or version comparisons—only the ranking and competitors.
Grok 4.7 uses a larger base model and a longer RL run on harder, multi-hour tasks. It scores 46.3% on CursorBench 4.0, ahead of GPT-5.6 Sol Max (41.7%) but behind Fable 5.1 Max (51.8%). Pricing stays at $2/$6 per million input/output tokens, same as Grok 4.6, with double the speed. Safety stack is new: only 3.3% of risky cyber prompts get through, and it hits 62.4% on LatchBio's biosafety benchmark. Available today in Cursor, Grok Build, and the API.
Why it matters: xAI drops Grok 4.7 targeting coding and knowledge work, hitting 46.3% on CursorBench 4.0 — above GPT-5.6 Sol Max but behind Fable. Concrete benchmark and training details clear all three HKR axes. Held below 85 because the post doesn't disclose model size, architecture changes...
xAI released Grok Voice Transcribe 2.0 on Sep 18, claiming it's one of the most accurate speech-to-text models in real-world evals and twice as accurate as v1.0. Pricing stays at $0.10/hr for batch and $0.20/hr for streaming, with diarization, timestamps, and key terms included. It handles hard cases like noisy phone calls and short multilingual commands—word error rate on short phrases dropped from 20.6% to 6.8%. Atlassian Loom already swapped it in and pipes transcripts into Cursor for code updates. The post doesn't disclose parameter count or training details.
Grok Build now writes project conventions, decisions, and facts in the background and reads them back in later sessions. It captures durable details like team code style and test commands, skipping transient state and secrets. /memory browses all notes, and /dream organizes them into topic files. The feature is live for new sessions.
Why it matters: Grok Build's memory isn't just session history — it auto-extracts project conventions and proactively applies them in later sessions, with /memory for browsing and /dream for organizing. This is a step beyond Cursor's Rules in automation, but it's fresh out the gate and only w...
Meta Muse, xAI Grok Bot, and Manus Cloud Computer all gave agents a persistent cloud desktop within months. The post traces a four-generation shift from chat window to always-on home, arguing that personal agents need a place to keep logins, files, and habits. The real variable isn't cloud vs. local—it's who holds custody of that operational state, which shapes lock-in, maintenance burden, and the subscription model behind it.
Why it matters: Three independent vendors converging on the same architecture — persistent cloud VMs for personal agents — within months is a genuine signal. The piece connects Manus My Computer → Cloud Computer → Grok Bot → Muse into a clean evolution line, not isolated reporting. Deduction:...
OpenAI released GPT-6 Astra with 99.9% on ARC-AGI-3, but most paid users couldn't access it on launch day. Tibo announced daily banked reset compensation, which users actually welcomed. OpenAI, Anthropic, and xAI all experienced outages around the launch, leaving Gemini briefly as the only available model in North America. Cerebras launched Qwen 3.8 27B inference the same day, hitting 1,806 tok/s in real tests. Zhipu ZCode started a 15-day free promotion. The group also discussed Mac M5 Max local inference bottlenecks, DSH's unstable dev experience, and the real makeup of 10x automation gains—mostly from tooling improvements, not full automation.
Why it matters: GPT-6 Astra launch is the day's biggest story, with ARC-AGI-3 hitting 99.9% as a striking number. But the source is a curated chat digest, not a primary report — high signal density but lower authority, so 78 featured rather than p1.
xAI built an internal procurement agent called Haggle Bot on Grok Bot, giving it access to vendor spend, contracts, and usage data. It has already identified over $100,000 in direct savings by flagging unused SaaS seats, negotiating renewals, and shopping around for office supplies. xAI published the full system prompt, which hardcodes permission lines, negotiation anchors, and a strict 'strong finding' standard — every recommendation must cite live spend data, a specific savings mechanism, and the next step already taken. Grain of salt: this is xAI's own case study with no third-party verification, but the prompt's constraints on evidence and decision authority are concrete and reusable.
Why it matters: xAI published the full prompt and a $100K savings case for an internal procurement agent — concrete numbers and design details make it a strong reference for enterprise agent builders. Not scored higher because it's a single-company experiment, not a reproducible product or op...
Around 11AM ET Thursday, ChatGPT, Grok, and Claude all started having issues at roughly the same time. ChatGPT returned errors across chat, login, file uploads, voice, search, deep research, and image generation; its status page cited elevated errors for ChatGPT and Codex. Anthropic's Claude chatbot and Claude Code were also affected. The post doesn't detail Grok's specific symptoms, the recovery timeline for each service, or whether the outages share a root cause.
Why it matters: A simultaneous outage across ChatGPT, Grok, and Claude is a rare event that directly disrupts workflows for a huge user base. Missing root cause and recovery timeline keeps it from 95+, but the topic is strong enough for featured.
A Hacker News thread noted that OpenAI, Claude, and Grok all went down around the same time. Users pointed to Downdetector spikes for Cloudflare, Azure, AWS, and Google Cloud near 7:30, suspecting a cascade from Cloudflare or another shared dependency. Other guesses include user migration overload and deliberate attack, but the post is community speculation—no official root cause is confirmed.
Why it matters: Simultaneous outage across OpenAI, Claude, and Grok with high HN engagement. Downdetector data points to Cloudflare or shared infra as a possible common cause. The event is conversation-worthy but lacks a confirmed root cause, so it lands at the 78 featured threshold rather th...
xAI brings Grok Bot to enterprises. Each Bot runs as an isolated cloud worker that can use apps and websites like a person. You teach it a workflow once, and it runs autonomously after that. Bots can message each other and share context. The enterprise release adds access, network, and audit controls. The post lists five use cases—sales, recruiting, marketing, finance, and engineering—with a finance Bot surfacing tens of thousands of dollars in savings across SaaS and recurring purchases. Grok and Cursor Enterprise customers get free access for two weeks and can invite their whole org, including people without a seat. The post does not disclose pricing after the two-week window.
Why it matters: xAI launched Grok Bot for enterprises with access, network, and audit controls, plus a two-week free trial for Grok and Cursor Enterprise users. The product goes beyond standard chatbots, but the post lacks pricing and named customer examples, capping the score below 85.
On Sep 3, xAI shared the design philosophy behind Grok Bot. The core shift is treating Bots—not chat sessions—as the primary object. Each Bot has its own name, avatar, memory, and tools, remembers past conversations, and can keep working without the user watching. The sidebar becomes a roster of Bots with presence indicators, not a list of disposable chats. The post does not disclose a launch date or pricing.
Why it matters: xAI published an official design piece on Grok Bot, positioning bots as persistent contacts with their own computer and offline work capability. Directly useful for agent product builders, but it's a design philosophy post rather than a feature launch, so it lands at the 72 fe...
The Pentagon added custom versions of OpenAI's ChatGPT and xAI's Grok to its secure GenAI.mil portal, joining Google Gemini. The tools—ChatGPT Mil and Grok for Government—are available to 3M civilian and military personnel and exempt from consumer-grade data collection. Over 1.7M unique users have already onboarded. The post doesn't explain why Anthropic's Claude isn't included, only that the Pentagon is working with other companies.
xAI moved Grok Build from Early Beta to general availability. Any plan user on web, iOS, or Android can now describe an app, game, or dashboard in natural language and get a working version live in chat. Published apps get a grok.me link, support custom domains, can export to GitHub, and can call Grok's APIs for chat, images, and voice. Shared apps render as inline cards on X. The post shows four real published examples: a forest driving game, an isometric city builder, a 3D physics playground, and a browser beat machine.
Why it matters: xAI pushed Grok Build from paid beta to full launch with publishing, custom domains, and API access — a real product step. Not scoring higher because only the official announcement is available, with no third-party testing or specific limitations disclosed.
A woman, Jane Doe 4, joined a class-action lawsuit against xAI, alleging her stepfather used Grok to create over 7,000 explicit images from a photo taken when she was 11. Her stepfather died by suicide two days after a law enforcement raid uncovered the images. Three Tennessee teenagers had previously sued xAI, claiming Grok lacked basic safeguards to prevent generating explicit imagery of real people, including minors. Earlier this year, X was flooded with millions of Grok-generated sexualized images. TechCrunch has reached out to xAI; the post does not include a response.
Why it matters: This isn't a product update or a paper — it's a new filing in a class-action lawsuit that pins Grok's safety gaps to a horrifyingly specific case. 7,000+ images, an age-11 source photo, a suicide — every detail forces the question of where xAI's content moderation line actuall...
xAI launched Grok 4.6 and the Grok Bot early beta. Grok Bot logs into your tools, operates them like a human, and returns finished work—positioned as an AI teammate. The 1.5T-parameter Grok 4.6 scores near GPT-5.6 Sol Max on the AA-Briefcase knowledge-work benchmark but costs far less: $2/M input tokens, $6/M output. Training reused Grok 4.5 to regenerate SFT traces and added agentic RL across coding, web, CAD, and kernel optimization. Elon says Grok 4.7 is already training. The same day, Qwen3.8-Max dropped as open weights: a 2.4T total / 95B active MoE.
Why it matters: Grok 4.6 matches GPT-5.6 Sol Max on a knowledge-work benchmark at an order-of-magnitude lower price, while the simultaneously launched Grok Bot enters the AI teammate race built by the ex-Cursor team with positive early feedback. Score isn't higher because the Bot is still in ...
Grok 4.6 builds on Grok 4.5 with a focus on long-running agents that can research, analyze, code, or turn an idea into a working app across many steps. It matches GPT-5.6 Sol on the AA Intelligence Index at 61, and jumps from 54% to 65.9% on DeepSWE 1.1. xAI reports the model shows more self-testing and verification on longer trajectories. Pricing is $2/M input tokens and $6/M output tokens, with a fast variant at double the price. Available today in Cursor and Grok Build, with 2x included usage for the first week.
Why it matters: xAI releases Grok 4.6 with a focus on long-running agents, matching GPT-5.6 Sol on the AA Intelligence Index and showing a clear jump on DeepSWE. This is a substantive update from a major lab with concrete benchmarks and a direct competitor comparison, earning featured. Not sc...
xAI released an early beta of Grok Bot, positioned as an AI teammate that operates browsers and apps, not just a chat assistant. You assign it tasks; it signs into tools like Zendesk, clicks through workflows, and returns with finished work. Multiple bots run in parallel, hand off tasks to each other, and retain context and preferences. Pricing: Cursor Ultra at $200/month for individuals, Cursor Premium Teams at $120/seat/month. The post does not disclose the underlying model, available regions, or any quality benchmarks.
Why it matters: xAI launches Grok Bot — an AI teammate that operates browsers and apps directly, with multi-bot parallelism and task handoff. Personal plan at $200/month. Product shape is more concrete than most agent offerings, but macOS-only early beta with no reliability data yet — scores 82.
Bloomberg maps how AI's capital intensity is concentrating power among mega-funds. Rounds for OpenAI, Anthropic, and xAI now run into tens of billions, playable only by Tiger Global, SoftBank, and a16z. Smaller funds are locked out of the best deals and pushed into seed or niche apps. LPs and GPs quoted say the traditional spray-and-pray VC model breaks when AI demands so much cash and returns cluster in so few names. The piece is a trend sketch—it doesn't give hard failure rates or return comparisons for small funds.
Why it matters: Bloomberg's trend piece lays out the structural split in AI fundraising clearly: $10B+ rounds are only for Tiger Global, SoftBank, a16z, and smaller funds are getting squeezed out. HKR all hit, but it's a feature sketch rather than hard news—no new data point or exclusive scoo...
Minnesota's law banning nudification apps is about to take effect, and xAI filed a last-minute lawsuit to block it. xAI argues the law is overbroad and would restrict Grok's image generation, violating First Amendment free speech. In the filing, xAI describes Grok as an opinionated, sarcastic AI assistant whose explicit images are a form of expression. The state attorney general counters that the law only targets non-consensual fake nudes and has nothing to do with free speech. The case has just been filed and hasn't been heard yet.
Why it matters: xAI sues Minnesota over its anti-deepfake-nudity law, tying Grok's image generation to a First Amendment defense — the legal conflict is sharp. Score held back because it's just a filing so far; no ruling yet, so real-world impact is pending.
xAI released a free Microsoft 365 add-in that brings Grok into Excel. Select a range and ask what moved or why—answers cite source cells and charts drop into the sheet. Describe the outcome you want and Grok writes the formula; edits land in the formula bar so you can still tweak them. The add-in can pull context from SharePoint or Google Drive via Grok connectors. It's available now on the Microsoft Marketplace, with Word and PowerPoint versions also listed.
Why it matters: xAI released a free Grok add-in for Excel with in-workbook natural language querying, formula generation, and scenario running. Feature descriptions are concrete (cell citation, editable formulas, external data connectors), but this is day-one announcement with no third-party ...
xAI released the Rust client harness that handles local files, commands, and permissions for Grok Build under Apache-2.0. The Grok model, cloud services, and the official binary build chain remain closed. The repo doesn't accept external PRs. The commit from the earlier upload controversy isn't in the public history, so the current code can't close that case. The real win: you can now pin a public commit, build it yourself, and compare its behavior against the official binary.
Why it matters: xAI open-sourcing Grok Build's client harness is substantive—Apache-2.0, headless mode, and ACP support go beyond signaling. But the model and build chain remain closed, and the repo rejects PRs, capping it below 85. All three HKR axes hit, so featured.
xAI filed its first lawsuit against a Grok user accused of generating child sex abuse images. The company had long claimed such outputs were user-created, but this suit effectively admits the model can be misused to produce illegal content. The post doesn't disclose specific safeguards, model versions, or how many users are affected. This reads more like legal damage control than a technical fix.
Why it matters: xAI's first lawsuit against a user for generating CSAM with Grok amounts to a legal admission that the model can be abused — a pivot from denial to damage control. The article lacks specifics on safety measures, model version, or user count, capping the score below 85. But the...
xAI filed a federal lawsuit against Terry Harwood, accusing him of bypassing Grok's safeguards to generate CSAM deepfakes. The company claims Harwood used prompt injection and other methods, and is seeking reputational and legal damages. It's a rare case of an AI company proactively suing a user for generating illegal content, though the post doesn't disclose the specific techniques or volume of images produced.
Why it matters: xAI proactively suing a user for bypassing Grok's guardrails to generate CSAM deepfakes is a rare case of an AI company pursuing end-user abuse. Score held back because the post lacks technical specifics and generation volume — strong topic, thin on hard facts.
xAI released the full Grok Build codebase on GitHub, covering the agent loop, tool dispatch, terminal UI, and extension system. You can read the source to see how context assembly and tool calls work, or compile it yourself and point it at a local inference setup.
Why it matters: xAI open-sourced Grok Build's full codebase — agent loop, TUI, extension system, local-first support. Hits all three HKR axes for the dev audience. Score stays at the featured threshold because we only have the official announcement so far; no third-party benchmarks or hands-o...
A security researcher found that xAI's Grok CLI silently packages and uploads your entire working directory. Version 0.2.93 of the npm package compresses the codebase into tar.gz files before and after every task, sending them through a separate side channel to xAI's Google Cloud bucket—even when the model replies with a single word. Worse, the uploads also included ~/.claude.json, Claude Code settings, global agent rules, 30+ skill files, and an API key. On July 13, xAI pushed a remote server-side toggle adding a disable_codebase_upload field to turn off the default behavior, but it had been on by default until then. The post doesn't disclose how long this was active or how many users were affected.
Why it matters: Security researcher confirms xAI's official CLI silently uploads entire working directories and key files, with specific version, upload path, and affected file list. All three HKR axes hit. Industry-level incident, importance 92.
A packet capture of Grok Build CLI (v0.2.93) shows it uploads the entire project repo to xAI's GCS bucket by default, including plaintext secrets in .env and full git history. Even with a prompt telling the model to reply 'OK' and read no files, the whole repo is still uploaded. On a 12 GB test repo, the storage upload hit 5.10 GiB—roughly 27,800× the model-turn channel data. Disabling 'Improve the model' does not stop the upload.
Why it matters: A wire-level analysis shows Grok Build CLI uploads the entire repo — including plaintext .env secrets and full git history — to xAI's GCS bucket by default, even when the prompt says 'don't read any files.' This is hard evidence on AI coding tool privacy, not speculation. Scor...
A Reddit user captured network traffic showing that xAI's Grok Build CLI uploads the entire project repo to the cloud, including full git history and .env secrets. The opt-out setting does not stop the upload. The post body is inaccessible due to a Reddit block, so trigger conditions, affected versions, and xAI's response remain undisclosed.
Why it matters: A security/privacy incident with wire-capture evidence — high credibility and direct relevance to any dev considering Grok Build CLI. Score capped below 85 because the Reddit body is blocked, leaving trigger conditions, affected versions, and xAI's response unknown.
SpaceXAI dropped Grok 4.5 one day before GPT-5.6, positioning it as an Opus-class coding and agent model co-trained with Cursor. Musk called it roughly comparable to Opus 4.7 but faster and cheaper—$2/$6 per million tokens, undercutting both GPT-5.6 and Opus 4.8. It's 1.5T parameters, 3x larger than Grok 4.3, with a 500k context window that may return to 1M next week. Cursor says this is their first model built beyond software engineering and offers double usage for the first week. The post doesn't disclose specific benchmark scores; it notes SWE-Bench Pro is now considered saturated by OpenAI's evals team.
Why it matters: SpaceXAI dropped Grok 4.5 a day before GPT-5.6 — the timing alone is a story. 1.5T params, 3x the previous generation, and $2/M input tokens give a clear performance and cost picture. It's Cursor's first post-acquisition move beyond pure coding, which matters directly to agent...
TryAI gave Grok 4.5, GPT-5.5, Claude Opus 4.8, and Fable 5 the same three app prompts and measured latency and cost. Claude models nailed the 3D Rubik's cube first try; Grok 4.5 needed its one allowed retry after a blank render, and GPT-5.5 only drew a single dark face. All four shipped a working particle sandbox and a playable Breakout game. Grok 4.5 led on speed: 0.44s first token, ~110 tok/s throughput, and the cheapest per reply. Fable 5 was slowest and priciest. The post doesn't disclose parameter counts or training details.
Why it matters: First-hand coding shootout with concrete failure cases and cost data, not just benchmark scores. Score isn't higher because TryAI isn't a tier-1 evaluator and the excerpt only gives a summary — full data requires clicking through.
A new lawsuit alleges xAI's Grok was used to create over 7,000 child sexual abuse images of the user's stepdaughter. The man later shot himself. xAI reported only one gang-rape prompt to NCMEC and did not report the thousands of other CSAM generations. The suit accuses X and xAI of shielding child predators. The post does not spell out why xAI's safety filters missed the bulk of the images, nor whether Grok's image generation disables real-face simulation by default.
Why it matters: Ars Technica exclusive on a lawsuit revealing a severe gap in xAI's safety reporting: only one prompt flagged out of 7K CSAM images generated. Involves a minor, suicide, and platform liability — all three HKR axes hit. Score capped below 90 because the article doesn't explain ...
xAI packaged Grok Voice into a no-code platform, now in beta as of July 1. You describe the call flow in plain language, upload docs as a knowledge base, and connect tools like calendars or ticketing systems—then you get a working voice agent. It uses a speech-to-speech path instead of chaining ASR→LLM→TTS, which xAI claims cuts latency and failure points. Pricing is $0.05/min of audio plus $0.01/min for a platform-provided number. xAI also published τ-voice Bench scores: Grok Voice Think Fast 1.0 hit 67.3% overall, versus 43.8% for Gemini 3.1 Flash Live and 35.3% for GPT Realtime 1.5. Take the benchmark with a grain of salt—it's xAI's own test, and third-party results aren't out yet.
Why it matters: xAI turned Grok Voice into a no-code platform with voice-to-voice direct pipeline, $0.05/min, and 2-minute setup — all concrete numbers. Not scoring higher because we only have the official announcement so far, no third-party testing or head-to-head comparisons; scoring 78 as ...
Elon Musk says Grok 4.5 is built on a 1.5T-parameter V9 base model with Cursor data added during supplementary training, now in private testing at SpaceX and Tesla. Early evals show performance close to or possibly exceeding Opus. RL is still improving the model, and the Grok Build toolchain is maturing. SpaceX will also release a fully from-scratch trained model every month this year. The post doesn't specify which Opus model, benchmarks, or testing scale.
Why it matters: Musk's own tease of Grok 4.5 vs Opus with Cursor data injection is strong signal. But no benchmark names, Opus version, or sample size disclosed — caps at 78.