Skip to content

OpenAI / ChatGPT

Everything OpenAI: the GPT models, ChatGPT and Sora, company strategy and people moves.

Latest picks

1401–1420 of 1,549

Aug 28, 2025Thursday

OpenAI News

Introducing gpt-realtime and Realtime API updates for production voice agents

OpenAI released the speech-to-speech model gpt-realtime and made the Realtime API generally available, adding remote MCP server support, image input, and SIP phone calling. The post reports 82.8% on Big Bench Audio versus 65.6% for the December 2024 model, and 30.5% on the audio MultiChallenge benchmark versus 20.6%. The key change is that tool access and phone connectivity now ship in the same production API.

Why it matters: This is a substantive OpenAI model + API release, not a minor refresh. HKR-H/K/R all pass: the release has a clear hook, hard benchmark deltas, and direct deployment impact for production voice agents, so it reaches p1.

Aug 27, 2025Wednesday

OpenAI News

Collective alignment: public input on our Model Spec

OpenAI surveyed over 1,000 people worldwide, compared their preferred model behavior with its Model Spec, and adopted some changes from disagreements. The post says participants ranked 4 completions per prompt, OpenAI compared them with a GPT-5 Thinking-based Model Spec Ranker, and released the dataset on HuggingFace. The key issue is default behavior; the captured post does not disclose the full list of adopted changes.

Why it matters: OpenAI turns >1,000 public preference rankings into Model Spec edits and releases the dataset, so HKR-H/K/R all pass. The real signal is default-behavior governance, but the excerpt does not show the full change list, keeping it in the 78–84 band.

OpenAI News

OpenAI and Anthropic share findings from a joint safety evaluation

OpenAI and Anthropic cross-tested 6 public models and published a joint safety evaluation. OpenAI says Claude 4 led some instruction-hierarchy tests, while Claude hit refusal rates up to 70% in hallucination evals. Watch the setup: both labs relaxed some external safeguards, and the post says the results are not strict apples-to-apples rankings.

Why it matters: HKR-H/K/R all pass: rival frontier labs jointly evaluating six public models is inherently clickable, and the post adds five test categories plus a 70% refusal datapoint. This is a strong safety research release, not a model launch or executive move, so it lands in featured, notp

Aug 26, 2025Tuesday

OpenAI News

OpenAI details ChatGPT crisis support and safety improvements

OpenAI says GPT-5, now the default ChatGPT model, cut non-ideal responses in mental health emergencies by over 25% versus 4o. The post says ChatGPT routes suicidal users to 988, Samaritans, or findahelpline.com and that OpenAI works with 90+ physicians across 30+ countries; the post body is truncated, so later plans are not fully disclosed.

Why it matters: HKR-H/K/R all pass: the post gives a concrete 25%+ reduction in non-ideal crisis replies, named referral pathways, and a strong safety-trust angle. I keep it at 82 because this is a focused safety update, not a broad capability launch, and the latter part is truncated.

Aug 7, 2025Thursday

OpenAI News

Introducing GPT-5 for developers

OpenAI released GPT-5 in its API on August 7, 2025, in three sizes: gpt-5, gpt-5-mini, and gpt-5-nano. The post reports 74.9% on SWE-bench Verified, 88% on Aider polyglot, 96.7% on τ2-bench telecom, plus new verbosity, minimal reasoning_effort, and custom tools; pricing and full availability details are not disclosed in the provided text. The real developer signal is the API surface change, not just a model rename.

Why it matters: This is an OpenAI flagship-model API launch, so it belongs in the 95–100 band. HKR-H lands on the GPT-5 debut; HKR-K lands on concrete benchmark scores and new controls; HKR-R lands on immediate developer concerns around migration, tooling, and model comparison; the excerpt omits

OpenAI News

GPT-5 and the new era of work

OpenAI launched GPT-5 on August 7, 2025, started rollout to Team users the same day, said Enterprise and Edu access would follow next week, and made it available in the API immediately. The post gives two hard numbers: 5 million paid ChatGPT business users and nearly 700 million weekly ChatGPT users; it does not disclose benchmark scores, pricing, or context length.

Why it matters: An OpenAI GPT-5 launch is a market-wide event, so HKR-H/K/R all pass. The post gives rollout timing and a 5M paid-business-user datapoint, but it omits benchmark scores, pricing, and context length, so this lands at the low end of the top band.

OpenAI News

Introducing GPT-5

OpenAI launched GPT-5 on August 7, 2025 and made it available to all ChatGPT users. The system combines a base model, GPT-5 thinking, and a real-time router; Plus gets higher limits, while Pro gets GPT-5 pro. The key change is unified routing with built-in reasoning; the post does not disclose pricing, context window, or API specifics.

Why it matters: An OpenAI frontier-model launch is a top-band event on its own. The excerpt confirms a unified system (base model + GPT-5 thinking + router) and rollout to all ChatGPT users; HKR-H/K/R all pass, and missing price/context/API details do not block p1.

OpenAI News

First look at GPT-5

OpenAI published a page titled “GPT-5: First Look” on August 7, 2025, confirming a public reveal for GPT-5. The post only shows the title, date, and site navigation; it does not disclose model size, pricing, context window, benchmarks, or API details. This reads like a placeholder, not a technical brief.

Why it matters: This is a high-attention, low-information official placeholder. HKR-H and HKR-R pass because GPT-5 appears on OpenAI’s site and that alone hits the model-race nerve; HKR-K fails because the body gives no specs, pricing, context window, benchmarks, or API details. Source authority

OpenAI News

GPT-5 System Card

OpenAI published the GPT-5 System Card on Aug. 7, 2025, stating GPT-5 combines gpt-5-main, gpt-5-thinking, and a real-time router, with mini models used after limits are hit. The API exposes gpt-5-thinking, gpt-5-thinking-mini, and gpt-5-thinking-nano, while ChatGPT adds gpt-5-thinking-pro; the post does not disclose pricing, context window, or benchmark scores. The key signal is safety: OpenAI classifies gpt-5-thinking as High capability in biological and chemical domains and applies the related safeguards.

Why it matters: This system card for OpenAI’s flagship model discloses GPT-5’s routed architecture, mini fallback, and direct access to thinking variants. HKR-H/K/R all pass; the High bio/chem capability rating makes this a same-day safety and deployment story, not routine documentation.

OpenAI News

From hard refusals to safe-completions: toward output-centric safety training

OpenAI says GPT-5 uses safe-completion training, shifting safety from binary input refusal to judging whether the output itself stays safe. The post describes two levers: severity-weighted penalties for policy-violating outputs and helpfulness rewards for safe replies; in a fireworks example, o3 gives actionable current and resistance values, while GPT-5 refuses the details and offers compliant alternatives. The key missing piece is the benchmark data: the post claims better safety and helpfulness, but the provided text does not disclose scores, benchmark names, or deltas.

Why it matters: This is a substantive OpenAI GPT-5 safety-training release, and it clears HKR-H/K/R: a real framing shift, concrete mechanisms, and a strong industry nerve. It stops short of p1 because the provided text does not disclose benchmark names, scores, or effect sizes.

Aug 6, 2025Wednesday

OpenAI News

Providing ChatGPT to the Entire U.S. Federal Workforce

OpenAI partnered with the U.S. General Services Administration to offer ChatGPT Enterprise to the full federal executive workforce for $1 per agency for 1 year. Participating agencies also get 60 days of unlimited advanced models and features, including Deep Research and Advanced Voice Mode; federal business data will not be used for training. The key signal is centralized procurement access, while the post does not disclose agency count, budget size, or exact model list.

Why it matters: OpenAI’s GSA deal turns ChatGPT Enterprise into a federal-wide procurement channel at $1 for one year, which is a real distribution signal, not a routine discount. HKR-H/K/R all pass, but the post omits agency count, budget size, and full model scope, so I keep it at 84.

Aug 5, 2025Tuesday

OpenAI News

Open Weights and AI for All

OpenAI said on August 5, 2025 it released its “most capable open-weight reasoning models” and will route them through OpenAI for Countries and its nonprofit grantee programs. The post confirms on-prem deployment and support for data-residency and security-constrained use cases, but does not disclose model names, parameter sizes, licenses, or benchmark results. The key missing piece is distribution detail, not the open-weight claim itself.

Why it matters: OpenAI shipping open-weights reasoning models clears HKR-H/K/R on novelty, a concrete deployment fact, and strategic resonance. Held at 86, not higher, because the post withholds the model name, size, license, and benchmark scores.

OpenAI News

Introducing gpt-oss

OpenAI released gpt-oss-120b and gpt-oss-20b under Apache 2.0, with the 120B model running on one 80GB GPU and the 20B model on devices with 16GB memory. Both are MoE Transformers with 117B and 21B total parameters, 5.1B and 3.6B active params per token, 128k context, and support for the Responses API and Structured Outputs. The part that matters is the lower deployment bar plus open weights; the post excerpt claims strong reasoning, but full benchmark scores are not disclosed here.

Why it matters: Same-day write. OpenAI moving into Apache 2.0 open weights is a strategy story, not a routine update; HKR-H lands on the unexpected move, HKR-K on concrete deployment specs, and HKR-R on cost and open-vs-closed debates. Not 95+ because the excerpt does not disclose full benchmark

OpenAI News

gpt-oss-120b & gpt-oss-20b Model Card

OpenAI released gpt-oss-120b and gpt-oss-20b as open-weight reasoning models under Apache 2.0, with compatibility for the Responses API. They are text-only models with tool use, Structured Outputs, and adjustable reasoning effort; the post does not disclose context length, pricing, or benchmark scores. On safety, OpenAI says gpt-oss-120b stayed below the High threshold in bio, cyber, and AI self-improvement tests, including after adversarial fine-tuning.

Why it matters: This is a same-day write: HKR-H from OpenAI going open-weight, HKR-K from license/mechanism/safety specifics, and HKR-R from the open-vs-closed debate. I kept it below 90 because the post excerpt does not disclose context length, pricing, or full benchmark results.

OpenAI News

Estimating Worst-Case Frontier Risks of Open-Weight LLMs

OpenAI says malicious fine-tuning tests on gpt-oss informed its decision to release the model. It trained gpt-oss for maximum biorisk with RL plus web browsing, and for cyber risk in an agentic coding CTF setup; the resulting models still underperformed OpenAI o3. The key signal is the evaluation method, because the post does not disclose exact scores, training scale, or release thresholds.

Why it matters: HKR-H/K/R all pass: the malicious-fine-tuning setup is novel, the paper gives two concrete eval environments, and the open-weight release debate is a live nerve. It stays at 80 because the post omits scores, training scale, and release thresholds.

Aug 4, 2025Monday

OpenAI News

What OpenAI is optimizing ChatGPT for

OpenAI said on August 4, 2025 that ChatGPT is optimized to help users finish tasks and leave, not maximize time spent. Break reminders are live for long sessions, and new behavior for high-stakes personal decisions is coming soon. Evaluation now includes custom rubrics built with 90+ physicians across 30+ countries.

Why it matters: Official OpenAI guidance on ChatGPT incentives and safety, with concrete facts: rest-break reminders are live and multi-turn evals include 90+ doctors from 30+ countries. HKR-H/K/R all pass, but the high-risk decision behavior lacks shipping scope and trigger details, so this is

Jul 31, 2025Thursday

OpenAI News

Introducing Stargate Norway

OpenAI said Stargate Norway, its first European AI data center project, is planned for 230MW with a further 290MW expansion target. Nscale and Aker will build it in Narvik, aiming for 100,000 NVIDIA GPUs by end-2026, using renewable power and closed-loop direct-to-chip liquid cooling. The key detail is allocation: OpenAI is an initial offtaker, while surplus capacity is intended for users in Norway, the UK, the Nordics, and Northern Europe; the post does not disclose capex or exact GPU models.

Jul 29, 2025Tuesday

OpenAI News

Introducing study mode in ChatGPT

OpenAI launched study mode in ChatGPT on July 29, 2025 for logged-in Free, Plus, Pro, and Team users, with ChatGPT Edu coming in the next few weeks. It uses custom system instructions to deliver Socratic prompts, scaffolded responses, knowledge checks, and on/off toggling instead of direct answers, adapting to skill-level questions and prior chat memory. The key change is interaction design, not a new model; the post does not disclose the underlying model, outcome metrics, or misuse safeguards.

Jul 22, 2025Tuesday

OpenAI News

Pioneering an AI clinical copilot with Penda Health

OpenAI and Penda Health studied 39,849 visits across 15 clinics in Kenya and found clinicians using AI Consult had 16% fewer diagnostic errors and 13% fewer treatment errors. The copilot used GPT-4o from August 2024, was embedded into the EHR in early 2025, and surfaced green/yellow/red alerts, with red alerts requiring review. The key point is deployment design: this is not autonomous care, but a safety net that triggers when an error is likely.

OpenAI News

Stargate advances with 4.5 GW partnership with Oracle

OpenAI and Oracle agreed to add 4.5 GW of Stargate data center capacity in the U.S., bringing capacity under development to over 5 GW and more than 2 million chips. OpenAI says this advances its January pledge to build 10 GW of U.S. AI infrastructure with $500 billion over four years, and it now expects to exceed that target. The concrete signal is deployment: Stargate I in Abilene has started receiving Nvidia GB200 racks and is already running early training and inference workloads.

Why it matters: This clears HKR-H/K/R: the hook is the sheer 4.5GW scale, the post includes concrete capacity numbers, and compute supply is a live industry nerve. At 88, this is a same-day infrastructure story with strategic impact, below only top-tier model or executive news.