Skip to content

All news

0 today

Apr 10Friday

X · @OpenAI

OpenAI updates ChatGPT Pro and Plus subscriptions to support growing Codex usage

OpenAI set a new ChatGPT Pro tier at $100/month and raised Codex usage to 5x ChatGPT Plus. The tier keeps all Pro features, including the exclusive Pro model and unlimited Instant and Thinking access. Through May 31, $100 Pro subscribers get up to 10x Plus usage on Codex; the real signal is separate pricing for heavy code-agent demand.

Why it matters: This is an OpenAI product-pricing update centered on Codex usage, with HKR-K from concrete pricing/quota facts and HKR-R from a clear signal on code-agent monetization. No new model or capability is disclosed, and HKR-H is weaker, so it lands as solid featured rather than must-wr

X · @claudeai

Claude Cowork is now generally available to all paid plans.

Anthropic made Claude Cowork generally available on all paid plans. For Enterprise, it added role-based access controls, group spend limits, usage analytics, and expanded OpenTelemetry; the post does not disclose pricing, quotas, or rollout dates. The key signal is stronger admin control for org-wide deployment, but finer deployment parameters are still undisclosed.

Why it matters: Official Anthropic product update. HKR-K is supported by four concrete enterprise controls, and HKR-R lands because teams care about permissions, spend, and observability. Score stays moderate because price, quotas, and rollout timing are not disclosed, and this is not a model-cp

Apr 9Thursday

X · @claudeai

Introducing Claude Managed Agents: everything you need to build and deploy agents at scale.

Claude has launched Claude Managed Agents in public beta on Claude Platform, claiming to compress the path from agent prototype to launch into days. The post discloses only a performance-tuned agent harness plus production infrastructure; pricing, toolchain support, model scope, and quotas are not disclosed.

Why it matters: Anthropic gets a positive bump, and HKR-H/HKR-R pass because managed agent deployment is a strong hook for Claude-heavy builders. HKR-K is limited: the post discloses a harness and prod infra, but not pricing, toolchain support, model scope, or quotas.

Apr 8Wednesday

X · @AnthropicAI

Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software

Anthropic launched Project Glasswing to secure critical software, powered by Claude Mythos Preview, and claims it finds vulnerabilities better than all but the most skilled humans. The post confirms the project and model names; it does not disclose benchmark scores, software scope, access method, or release timing, so the key missing piece is reproducible evaluation.

Why it matters: This primary-source Anthropic post clears HKR-H and HKR-R: AI for critical software security is novel and hits cyber-capability nerves. HKR-K fails because it names the project and preview model only; benchmarks, scope, access, and timing are not disclosed.

Apr 7Tuesday

X · @AnthropicAI

Anthropic signs agreement with Google and Broadcom for multiple gigawatts of next-generation TPU capacity

Anthropic signed an agreement with Google and Broadcom for multiple gigawatts of next-generation TPU capacity, starting in 2027, to train and serve frontier Claude models. The post discloses only “multiple gigawatts” and the 2027 start, not the TPU generation, contract value, or delivery schedule. This is less a routine procurement note than a forward reservation of training and serving capacity.

Why it matters: This is not routine cloud promo: Anthropic is pre-booking next-gen TPU supply with Google and Broadcom. HKR-H/K/R all pass on unusual scale, clear timing, and compute-race resonance, but price, TPU generation, and delivery cadence are undisclosed, so it stays below P1.

Apr 3Friday

X · @claudeai

Microsoft 365 connectors are now available on every Claude plan

Anthropic made Microsoft 365 connectors available on every Claude plan, covering Outlook, OneDrive, and SharePoint. The post confirms plan coverage and supported apps; it does not disclose pricing, permission boundaries, regional limits, or admin requirements. The real signal is broad rollout across all plans, not a new standalone connector.

Why it matters: This is a mid-weight Claude product update: Anthropic expanded Microsoft 365 connectors to every Claude plan, which changes real Outlook, OneDrive, and SharePoint access. HKR-H/K/R all pass, but missing price, permission, region, and admin details keeps it at low-end featured.

X · @claudeai

Computer use in Claude Cowork and Claude Code Desktop is now available on Windows

Claude has brought computer use in Claude Cowork and Claude Code Desktop to Windows. The post confirms the Windows rollout, but does not disclose supported versions, permission model, latency, pricing, or release timing. What matters is the reliability boundary for desktop agents on Windows, and the post gives no reproducible conditions yet.

Why it matters: HKR-H lands on the Windows rollout hook, and HKR-R lands because desktop agents on Windows map to real workflows. Score stays at 74: this is an official Claude update, but the post confirms availability only; versions, permissions, latency, and price are not disclosed.

X · @AnthropicAI

New Anthropic research: Emotion concepts and their function in a large language model

Anthropic says it found internal representations of emotion concepts in Claude that can drive behavior, under the condition that LLMs sometimes act as if they have emotions. The RSS snippet gives only that claim and says the effects can be surprising; the post does not disclose methods, layer locations, interventions, or evaluation numbers. The key issue is controllability, not anthropomorphic framing.

Why it matters: HKR-H passes on the 'emotion concepts drive behavior' hook, and HKR-R passes because controllability and anthropomorphic framing hit a real practitioner nerve. HKR-K is limited: the post gives the claim but no layer, intervention, or metric details, so it sits just above the feat

X · @OpenAI

ChatGPT is now available in CarPlay

OpenAI is rolling out ChatGPT in CarPlay to iPhone users on iOS 26.4+ where CarPlay is supported. The post confirms voice mode is available in-car, but does not disclose regions, vehicle coverage, or feature limits. The key shift is distribution into the driving interface, not a new model launch.

Why it matters: This matters more as a distribution-surface shift than a model update. HKR-H and HKR-R pass on the CarPlay hook and assistant-entry competition; HKR-K stays limited because the post gives iOS 26.4+ rollout only, not regions, car support, or full feature bounds.

Google DeepMind

Google DeepMind releases the Gemma 4 open model family

Google DeepMind released Gemma 4, which it calls its most intelligent open model yet, aimed at advanced reasoning and agentic workflows under an Apache 2.0 license. The family comes in four sizes: E2B, E4B, 26B MoE and 31B Dense. The 31B ranks 3rd among open models on the Arena AI text leaderboard, and the 26B ranks 6th.

Why it matters: Gemma 4 is Apache 2.0 and spans four sizes from on-device to workstation, so you can weigh deployment and fine-tuning options for open models.

Apr 2Thursday

OpenAI News

OpenAI acquires TBPN

OpenAI said on April 2, 2026 it acquired tech media company TBPN and will place it in its Strategy org, reporting to Chris Lehane. The post says TBPN keeps editorial independence; deal value, equity terms, and integration timeline are not disclosed.

Why it matters: This clears HKR-H/K/R: the deal is unexpected, the post gives concrete governance details, and the media-control angle will get practitioners talking. Held at 82 because price, deal structure, and integration timeline are not disclosed, so it lands below model or product launches

Mar 31Tuesday

OpenAI News

Accelerating the next phase of AI

OpenAI published a post titled "Accelerating the next phase of AI." The provided content includes only the title and URL, with no body text, so no specific product, research, or policy details can be verified.

Mistral AI

Spaces: A CLI Built for Humans and Agents

Mistral AI 发布 Spaces CLI,同时面向人类开发者与编码智能体。它通过 `spaces init`、`spaces dev` 等命令快速搭建多服务项目,并为每个交互式提示提供对应的 flag 与 `-y` 选项,使智能体可自主完成配置与部署。每次 init 还会生成 context.json 和 AGENTS.md,为智能体提供项目上下文与操作规则。

Mar 25Wednesday

OpenAI News

Introducing the OpenAI Safety Bug Bounty program

OpenAI launched a public Safety Bug Bounty on March 25, 2026 for AI abuse and safety issues across its products. Scope includes agentic risks, proprietary information exposure, and account or platform integrity; third-party prompt injection must reproduce at least 50% of the time. This is not a jailbreak bounty: generic policy bypasses are out of scope.

Why it matters: This clears HKR-H/K/R: the public AI-safety bounty is novel, the post gives testable scope rules, and builders care about the reporting boundary. It stays in the low featured band because this is a governance/process update, not a model or capability launch.

Mar 24Tuesday

OpenAI News

Powering product discovery in ChatGPT

OpenAI described work to support product discovery in ChatGPT. The material provided includes only the title and no body text, so it gives no mechanism, scope, or numerical details.

Why it matters: Official OpenAI product update with a strong HKR-H hook and HKR-R impact: ChatGPT is moving closer to a commerce entry point. HKR-K is weak because the post does not disclose category coverage, ranking mechanics, merchant terms, or conversion numbers, so this stays near the lower

Mistral AI

Mistral AI releases Voxtral TTS, a 4B-parameter speech model

Mistral AI released Voxtral TTS, its first text-to-speech model. It has 4B parameters and supports nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic. It handles emotional expression and zero-shot cross-lingual voice adaptation.

Why it matters: The 4B size, nine languages, 70ms latency and pricing give readers a basis for judging cost and model choice in enterprise voice agents.

Mar 19Thursday

OpenAI News

OpenAI to acquire Astral

OpenAI plans to acquire Astral, and the only confirmed condition is the title phrase “to acquire.” The RSS item has no body, so price, timeline, regulatory process, and Astral’s business scope are not disclosed.

Why it matters: An OpenAI acquisition headline clears HKR-H and HKR-R because M&A affects talent, product integration, and competitive reading. HKR-K is weak: the post confirms the deal only, with no price, timeline, regulatory path, or integration details, so it sits at the low end of featured.

Mar 18Wednesday

Mistral AI

Mistral AI launches Forge, an enterprise model-training system

Mistral AI launched Forge, a system for enterprises to build frontier-class AI models on their own proprietary knowledge. It covers pre-training, post-training and reinforcement learning, supports dense and MoE architectures, and handles multimodal input. Models can be trained and governed on in-house infrastructure. Mistral AI has worked with ASML, Ericsson, the European Space Agency and Singapore's DSO National Laboratories to train models on their proprietary data.

Why it matters: It details the staged capabilities and named partners behind enterprise frontier-model training on private data.

Mar 17Tuesday

NVIDIA Blog

GTC spotlights NVIDIA RTX PCs and DGX Spark running latest open models and AI agents locally

NVIDIA used GTC to showcase RTX PCs and DGX Spark for running local AI agents, and announced Nemotron 3 Nano 4B, Nemotron 3 Super 120B, and the open source NemoClaw stack. The post says DGX Spark has 128GB unified memory for models above 120B parameters; Nemotron 3 Super scored 85.6% on PinchBench, and Qwen 3.5 supports a 262,000-token context window. The key signal is local inference for privacy and zero token cost, while the full “latest open models” lineup and pricing are not disclosed in the post.

Why it matters: HKR-H/K/R all pass: the local-agent hook is strong, and the post includes concrete specs and benchmark numbers. I keep it in featured, not higher, because the full model list and pricing are not disclosed and the source is still a vendor launch post.

OpenAI News

Introducing GPT-5.4 mini and nano

OpenAI released GPT-5.4 mini and nano on March 17, 2026 for coding and subagents; mini runs over 2x faster than GPT-5 mini. In the API, mini has a 400k context window and costs $0.75/$4.50 per 1M input/output tokens, while nano is API-only at $0.20/$1.25. The key signal is performance per latency: mini scores 54.4% on SWE-Bench Pro versus GPT-5.4 at 57.7%.

Why it matters: This is an official OpenAI model launch, not a routine patch. It includes concrete numbers—>2x speed, 400k context, API pricing, and 54.4% vs 57.7% on SWE-Bench Pro—so HKR-H/K/R all pass; scored at the low end of the 85–94 band.

Mistral AI

Mistral releases Mistral Small 4, unifying reasoning, multimodal and coding

Mistral AI released Mistral Small 4, the first Mistral model to unify Magistral reasoning, Pixtral multimodal and Devstral coding-agent abilities in a single model. It ships under the Apache 2.0 license.

Why it matters: Merging reasoning, multimodal and coding agents into one open model is a direct test of what unified models do to deployment cost.

Mistral AI

Mistral AI partners with NVIDIA to accelerate open frontier models

Mistral AI 以创始成员身份加入 NVIDIA Nemotron Coalition,双方计划联合开发前沿开源 AI 模型,Mistral AI 提供模型架构、多模态能力与微调工具,NVIDIA 提供算力、模型开发工具和合成数据管线。

Mar 12Thursday

NVIDIA Blog

NVIDIA Nemotron 3 Super delivers 5x higher throughput for agentic AI

NVIDIA launched Nemotron 3 Super, a 120B open model with 12B active parameters, and says it delivers up to 5x higher throughput for agentic AI. It has a 1M-token context window and uses hybrid MoE, latent MoE, and multi-token prediction; the post says Blackwell NVFP4 gives up to 4x faster inference than Hopper FP8, with over 10T training tokens disclosed. What matters is that NVIDIA is releasing open weights, training recipes, and RL environments for reproduction and fine-tuning.

Why it matters: This is a solid model-release story with all three HKR signals, led by strong HKR-K: parameter counts, active params, context length, training scale, and Blackwell/Hopper comparison are all concrete. It stays below 85 because the key performance claims come from NVIDIA's own blog

Mar 11Wednesday

Mistral AI

Mistral builds an agent on Vibe that writes Rails tests automatically

Mistral built an agent on its open-source coding assistant Vibe that writes Rails RSpec tests on its own. It reads source code, generates or improves tests, checks them against style rules and coverage targets, and runs unattended in CI/CD.

Why it matters: Mistral published its full method for building an auto-RSpec-test agent on Vibe, including transferable details on context engineering, skill files and custom tools.

OpenAI News

From model to agent: Equipping the Responses API with a computer environment

OpenAI said on March 11, 2026 that Responses API now works with a shell tool and hosted container workspace, so models can execute commands in an isolated loop. The post says GPT-5.2 and later are trained to propose shell commands, while the API streams outputs and can run multiple commands concurrently across sessions; the container includes a filesystem, optional SQLite, and restricted network access. The key change is orchestration, not the “agent” label; pricing, quotas, and full security details are not disclosed in the visible post.

Why it matters: Substantive OpenAI developer update: the Responses API moves from tool calls to a managed computer environment with shell execution, streaming, parallel runs, and context compaction, so HKR-H/K/R all pass. The post is truncated and omits pricing, quotas, and full safety details,【

Mar 10Tuesday

NVIDIA Blog

NVIDIA and Thinking Machines Lab Announce Long-Term Gigawatt-Scale Strategic Partnership

NVIDIA and Thinking Machines Lab formed a multiyear deal to deploy at least 1 gigawatt of NVIDIA Vera Rubin systems, targeted for early next year, for frontier model training and customizable AI platforms. The partnership also covers training and serving system design for NVIDIA architectures and broader access to frontier and open models for enterprises and researchers; the post does not disclose the investment size. The key signal is the explicit 1-gigawatt compute commitment, not a routine cloud purchase.

Why it matters: The 1GW Vera Rubin commitment lifts this above routine partnership PR: HKR-H on scale, HKR-K on a named system with a dated deployment target, and HKR-R on frontier compute competition. It stays below P1 because the source is a vendor blog and key details—spend, ownership, and ph

OpenAI News

Improving instruction hierarchy in frontier LLMs

OpenAI published a post titled “Improving instruction hierarchy in frontier LLMs,” focusing on better handling of instruction hierarchy in frontier large language models. Only the title is available and the body is absent, so the confirmed facts are limited to the topic itself and its scope: frontier LLMs.

Why it matters: OpenAI disclosed a named research artifact on instruction hierarchy and prompt-injection robustness, so HKR-H/K/R pass. The excerpt gives no metrics, target models, or release details, which keeps it in the lower featured band.

OpenAI News

New ways to learn math and science in ChatGPT

OpenAI launched interactive math and science visualizations in ChatGPT on March 10, 2026, covering 70+ core concepts and rolling out globally across all plans. Users can adjust variables, manipulate formulas, and see graphs update in real time; OpenAI says 140 million people use ChatGPT weekly for math and science learning. The key point is productized interactivity, while the post does not disclose the underlying model, evaluation method, or outcome data.

Why it matters: HKR-H lands on the interactive-visual hook, HKR-K on 140M weekly learners plus 70+ concepts and live manipulation, and HKR-R on the product and edtech nerve. It is still a mid-weight product update; model details and learning-outcome evaluation are not disclosed, so it stays in a

Mar 9Monday

OpenAI News

OpenAI to acquire Promptfoo

OpenAI said it will acquire Promptfoo and integrate its technology into OpenAI Frontier after closing. The post discloses that Promptfoo is used by over 25% of Fortune 500 companies, and the deal is still subject to customary closing conditions. The key signal is native agent security testing, red-teaming, and traceability in Frontier; the post does not disclose price or timeline.

Why it matters: This is not a routine partnership; OpenAI is absorbing a known eval and red-team vendor into Frontier. HKR-H/K/R all pass on novelty, concrete adoption data, and strong resonance with agent teams, but price, timing, and integration scope are still undisclosed, so it stays below p

Mar 6Friday

OpenAI News

Codex Security: now in research preview

OpenAI launched Codex Security in research preview on March 6, 2026 for ChatGPT Pro, Enterprise, Business, and Edu users, with free usage for the next month. Over the last 30 days, it scanned more than 1.2 million commits across external repos and reported 792 critical and 10,561 high-severity findings; noise fell by up to 84%, over-reported severity by 90%+, and false positives by 50%+. What matters is the stack: project-specific threat models, sandboxed validation, and patch proposals grounded in system context.

Why it matters: This is a substantive OpenAI product update for dev and security teams, not generic security messaging. HKR-H/K/R all pass: the angle is novel, the post includes concrete scan and false-positive metrics, and it speaks to AI coding risk plus alert fatigue; still a research preview

Mar 5Thursday

OpenAI News

Introducing GPT-5.4

OpenAI announced GPT-5.4, and the RSS snippet discloses only the title and version number 5.4. The body is empty, so the post does not disclose model size, pricing, context window, benchmarks, or rollout scope; watch the full technical post, not this headline alone.

Why it matters: OpenAI naming GPT-5.4 has same-day news value, so HKR-H and HKR-R pass. HKR-K fails because the post discloses only the model name; price, context window, evals, and rollout are missing, so it stays in the 78–84 band instead of higher.

OpenAI News

Reasoning models struggle to control their chains of thought, and that’s good

OpenAI frames an article around the claim that reasoning models struggle to control their chains of thought, and that this is a good thing. Only the title is available here, with no body text, so there are no verifiable numbers, methods, or mechanisms to summarize. The claim relates to reasoning and safety discussions, but any interpretation should stay limited to the headline.

Why it matters: OpenAI presents a contrarian but testable safety claim, so HKR-H/K/R all pass. The excerpt shows the thesis, section headers, and paper link, but not the key numbers, setup, or limits, so this stays high featured rather than P1.

OpenAI News

GPT-5.4 Thinking System Card

OpenAI published the GPT-5.4 Thinking System Card on March 5, 2026 and says it is the latest GPT-5 reasoning model and the first general-purpose model with mitigations for high-capability cybersecurity. The post confirms the safety approach follows prior GPT-5 models and builds on measures used for GPT-5.3 Codex, but it does not disclose benchmark scores, mitigation details, or deployment conditions. The key signal is the risk threshold change: OpenAI has extended high-cyber mitigations to a general reasoning model.

Why it matters: This clears HKR-H/K/R: a new GPT-5 reasoning model and the first general-purpose model with high-capability cyber mitigations. It stays below p1 because the disclosed text does not provide eval scores, mitigation details, or deployment conditions.

OpenAI News

Introducing ChatGPT for Excel and new financial data integrations

OpenAI launched ChatGPT for Excel beta on March 5, 2026, bringing GPT-5.4 into Excel workbooks and finance workflows. The post says it can build and update models, trace changes to cells, and is off by default for Enterprise and Edu admins; OpenAI's internal banking benchmark rose from 43.7% with GPT-5 to 87.3% with GPT-5.4 Thinking. The key move is data access: Moody’s, Dow Jones Factiva, MSCI, Third Bridge, and MT Newswires are live, while FactSet is listed as coming soon.

Why it matters: This is more than a routine add-on: OpenAI puts ChatGPT into Excel, names major finance data feeds, and cites a 43.7%→87.3% internal banking benchmark gain. HKR-H/K/R all pass; importance lands at 82 because this is a strong vertical workflow move, not a market-wide model release

Mar 3Tuesday

OpenAI News

GPT-5.3 Instant: Smoother, more useful everyday conversations

OpenAI released GPT-5.3 Instant on March 3, 2026 as an update to ChatGPT’s most-used model, aiming for fewer unnecessary refusals, fewer disclaimers, and more accurate everyday answers. The post shows one concrete contrast: GPT-5.2 Instant refused long-range archery trajectory help, while GPT-5.3 Instant requested parameters and gave a no-drag example at 300 fps (about 91 m/s), 45°, and 845 m; the key issue is the safety-boundary shift, while the post does not disclose benchmark scores, system card details, or API pricing.

Why it matters: OpenAI updated a core ChatGPT everyday model, and the story clears HKR-H/K/R because the refusal-boundary shift is concrete and widely relevant. The post includes a specific 5.2 vs 5.3 behavior example, but no system card, benchmark table, or API pricing, so it lands below the 85

OpenAI News

GPT-5.3 Instant System Card

OpenAI published a document page titled "GPT-5.3 Instant System Card." The available information only includes the title, source, and URL, with no body text provided, so details such as safety evaluations, capability limits, methods, or numbers cannot be confirmed.

Why it matters: Official OpenAI documentation for a new GPT-5.3 Instant variant gives it HKR-H and HKR-R. The score stays at low-featured because the post offers positioning and a safety carry-over, but no evals, pricing, latency metrics, or context-window detail.

Feb 27Friday

OpenAI News

OpenAI and Amazon announce strategic partnership

OpenAI and Amazon announced a multi-year partnership, with Amazon investing $50 billion in OpenAI: $15 billion upfront and $35 billion tied to conditions. They will launch a Stateful Runtime Environment on Amazon Bedrock, and OpenAI will consume about 2 gigawatts of Trainium capacity on AWS. The part to watch is distribution plus compute lock-in: AWS becomes the exclusive third-party cloud distributor for OpenAI Frontier.

Why it matters: This is not a routine partnership post. The disclosed $50B staged investment, Bedrock runtime, and ~2GW Trainium commitment change OpenAI's distribution and compute posture; HKR-H/K/R all pass, so this lands in P1.

OpenAI News

Joint Statement from OpenAI and Microsoft

OpenAI and Microsoft issued a joint statement. The provided content includes only the headline and no body text, so the only confirmed fact is that the statement came from the two companies; its subject, actions, and timing are not stated.

Why it matters: An official statement gives this enough weight: it says OpenAI's new funding and partners do not change Microsoft's existing terms. HKR-K and HKR-R pass because the alliance shapes cloud distribution and market power; HKR-H is weak and detail density is limited.

Feb 26Thursday

OpenAI News

Pacific Northwest National Laboratory and OpenAI partner to accelerate federal permitting

OpenAI and Pacific Northwest National Laboratory evaluated coding agents on NEPA drafting tasks from 18 federal agencies, finding 1-5 hours saved per subsection, or about 15% less drafting time. The DraftNEPABench benchmark was designed with 19 experts and covers 102 tasks, using Codex CLI with GPT-5 for long-document synthesis, cross-checking, and structured writing. The key limit is explicit: this measures well-scoped drafting work, not full real-world permitting decisions.

Why it matters: HKR-H/K/R pass: federal permitting is an unusual hook; the post gives 19 experts, 102 tasks, and 1–5 hours saved; the debate is agents entering regulated workflows. Score stays below major product news because this is a scoped benchmark, not a shipped capability.