Skip to content

Safety & alignment

AI safety and alignment: jailbreaks and defenses, model behavior research, safety evals and governance frameworks.

Latest picks

561–580 of 582

Jan 22, 2025Wednesday

OpenAI News

Trading Inference-Time Compute for Adversarial Robustness

OpenAI reports that o1-preview and o1-mini often drive adversarial attack success rates close to zero as inference-time compute increases. The paper tests math tasks, SimpleQA prompt injection, Attack Bard images, and StrongREJECT misuse prompts; it labels the result as preliminary, and the truncated post does not fully disclose all failure cases. The key point is that this gain comes from longer reasoning at inference, not adversarial training.

Why it matters: Strong HKR-H/K/R: the hook is counterintuitive, the paper proposes a concrete mechanism, and it lands on a real safety/deployment nerve. I kept it at 82, not p1, because the post frames this as initial evidence and the excerpt does not fully disclose failure modes, cost tradeoffs

Dec 27, 2024Friday

OpenAI News

Why OpenAI’s structure must evolve to advance our mission

OpenAI says its board is evaluating changes to its nonprofit/for-profit structure, after estimating in 2019 that AGI would require about $10B. The post cites ChatGPT’s 300M+ weekly users and $137M in 2015 donations, but the specific final structure under consideration is not fully disclosed in the provided text. The key signal is financing pressure: OpenAI says investors at this scale want more conventional equity.

Dec 20, 2024Friday

OpenAI News

Deliberative alignment: reasoning enables safer language models

OpenAI published deliberative alignment on Dec 20, 2024, training o-series models to reason over written safety specs before answering. The post says o1 uses this method and needs no human-labeled CoT or answers; it says o1 beats GPT-4o on internal and external safety benchmarks, but the post does not disclose exact scores.

Why it matters: HKR-H/K/R all land: the angle is novel, the mechanism is concrete, and the topic hits a live industry debate on reasoning-model safety. I keep it at 83 because the post excerpt does not disclose key benchmark scores, so it stays in the high-quality research band, not must-write.

Dec 9, 2024Monday

OpenAI News

Sora is here

OpenAI moved Sora out of research preview on December 9, 2024 and rolled it out to ChatGPT Plus and Pro users. Sora Turbo supports up to 1080p and 20-second videos; Plus includes up to 50 monthly 480p videos or fewer 720p generations. The key detail for practitioners is deployment scope: the UK, Switzerland, and the EEA are excluded, person uploads are limited, and OpenAI says physics and long complex actions remain weak.

Why it matters: OpenAI moved Sora from preview to paid availability, so HKR-H/K/R all pass: high-curiosity launch, concrete specs and limits, and clear impact on creator workflows. I stop below 95 because the post itself notes region blocks, restrictions on uploads with people, and instabilityon

Dec 5, 2024Thursday

OpenAI News

OpenAI o1 System Card

OpenAI published the system card for o1 and o1-mini, with a deployment gate that requires post-mitigation risk scores of medium or lower. The listed Preparedness results are low for cybersecurity, medium for CBRN and persuasion, and low for model autonomy; testing covered o1-near-final-checkpoint and o1-dec5-release. The key point for practitioners is that OpenAI confirms large-scale RL for chain-of-thought reasoning, while the post does not disclose dataset mix or full benchmark scores.

Why it matters: This is a high-signal safety disclosure for a frontier OpenAI reasoning model, not routine collateral. HKR-K is strong because it publishes the deployment threshold, four Preparedness ratings, and test scope; HKR-R lands because practitioners track CoT safety, transparency, and 3

Nov 21, 2024Thursday

OpenAI News

Advancing red teaming with people and AI

OpenAI published 2 papers on Nov 21, 2024, outlining its external human red teaming process and a new automated red teaming method. The post discloses 3 concrete design choices for external testing—threat-model-based team selection, versioned model access, and structured feedback via API or ChatGPT interfaces—but this excerpt does not fully disclose the automated method's metrics or results.

Why it matters: HKR-K carries this story: OpenAI describes 2 papers and at least 3 reusable human red-team design choices. HKR-R also passes because safety and eval teams can apply the workflow; HKR-H is weaker, and the excerpt does not fully disclose automated-red-team results, so this sits at

Oct 30, 2024Wednesday

OpenAI News

Introducing SimpleQA

OpenAI open-sourced SimpleQA, a 4,326-question benchmark for factual short-answer QA and model calibration. Two independent AI trainers verified each item; a 1,000-question audit showed 94.4% agreement and an estimated inherent error rate near 3%. The key signal: it is built to challenge frontier models, and the post says GPT-4o scores below 40%.

Why it matters: This is not a routine paper post. HKR-H comes from the inversion that a 'simple' benchmark stumps frontier models; HKR-K comes from the dataset size, agreement rate, and irreducible-error estimate; HKR-R comes from the ongoing industry fixation on hallucination and calibration,so

Oct 24, 2024Thursday

OpenAI News

OpenAI’s approach to AI and national security

After the White House issued an AI National Security Memorandum on October 24, 2024, OpenAI published a framework for national security partnerships and said each use case goes through formal review by its Product Policy and National Security teams. The post names 3 existing examples: DARPA cyber defense work, USAID using ChatGPT to cut administrative burden, and bioscience collaboration with Los Alamos National Laboratory; it does not disclose pricing, model versions, or contract size. The key signal is the boundary: OpenAI says its policies ban uses that harm people, destroy property, or develop weapons, while it explores research, logistics, translation, summarization, and civilian-harm mitigation use cases with the U.S. and allies.

Why it matters: This is not a product launch, so HKR-H is weak. HKR-K and HKR-R pass on the concrete review process, 3 existing projects, and explicit weapons bans, but missing contract scale, model versions, and outcome data keep it at the low end of featured.

Oct 15, 2024Tuesday

OpenAI News

Evaluating fairness in ChatGPT

OpenAI analyzed millions of ChatGPT requests to test whether user names trigger harmful stereotypes, finding an overall rate of about 0.1%. The study used GPT-4o as a privacy-preserving evaluator; its gender-related judgments matched human raters over 90% of the time, while race and ethnicity agreement was lower. The key signal is model drift across versions: GPT-3.5 Turbo showed the highest task-level bias.

Why it matters: OpenAI provides a rare production-scale fairness audit with concrete rates, evaluator agreement, and a model-comparison result, so HKR-K is strong and HKR-R clears on trust and safety. This is a substantive research release, not a model launch or major product shift, so it lands

Oct 9, 2024Wednesday

OpenAI News

An update on disrupting deceptive uses of AI

OpenAI says it has disrupted more than 20 operations and deceptive networks that tried to abuse its models since the start of 2024. The post ties this to election-related influence campaigns, social-media manipulation, and state-linked actors, and links an October 2024 threat report; the post does not disclose model-level breakdowns or exact enforcement mechanics.

Why it matters: OpenAI clears HKR-H/K/R here: the 20+ takedown count is a real hook, the Oct. 2024 threat-intel update adds a concrete fact, and election-linked deception is highly resonant. It stays in featured, not higher, because operation-level samples, model names, and enforcement mechanics

Sep 26, 2024Thursday

OpenAI News

Upgrading the Moderation API with OpenAI's new multimodal moderation model

OpenAI released omni-moderation-latest on September 26, 2024, a GPT-4o-based Moderation API model for text and image inputs that is free for all developers. It adds illicit and illicit/violent text categories, supports image moderation in 6 subcategories, and improves 42% on an internal 40-language eval, with gains in 98% of languages tested.

Why it matters: Official OpenAI developer product update with strong HKR-K: new moderation classes, image coverage, and a concrete +42% result across 40 languages. HKR-R also lands because moderation and compliance affect shipping teams directly; HKR-H is weak, so this sits at the low end of the

Sep 16, 2024Monday

OpenAI News

An update on OpenAI's safety and security practices

OpenAI said on September 16, 2024 that its Safety and Security Committee will become an independent board oversight committee, chaired by Zico Kolter, for critical safeguards in model development and deployment. The committee can review major model safety evaluations and delay launches until concerns are addressed; the post also cites a 90-day review, evaluation of an AI-sector ISAC, and work with Los Alamos National Laboratory.

Why it matters: HKR-H/K/R all pass. OpenAI says an independent board committee can review major safety evaluations and delay release, which is more concrete than a generic safety post. It stays below P1 because there is no new model, external audit result, or reproducible benchmark data.

Sep 12, 2024Thursday

OpenAI News

Introducing OpenAI o1

OpenAI released o1-preview and o1-mini on Sept. 12, 2024, with access for ChatGPT Plus, Team, and tier-5 API developers. The post cites 83% vs 13% on an IMO qualifier, 84 vs 22 on a jailbreak test, and says o1-mini is 80% cheaper than o1-preview. The tradeoff is clear: the API lacks function calling, streaming, and system messages, and the models do not yet support browsing or file and image uploads.

Why it matters: A major OpenAI reasoning-model launch with all three HKR signals: HKR-H from the new “think before answering” hook, HKR-K from concrete benchmark, safety, and pricing numbers, and HKR-R from the tradeoff practitioners must manage between stronger reasoning and missing API basics.

OpenAI News

OpenAI o1-mini

OpenAI released o1-mini on Sept. 12, 2024 for Tier 5 API users at 80% lower cost than o1-preview. The post reports 70.0% on AIME and 1650 Codeforces Elo, close to o1 at 74.4% and 1673, with about 3-5x faster answers than o1-preview in one word-reasoning example. The key tradeoff is explicit: it targets STEM reasoning, while non-STEM factual knowledge is only comparable to small models like GPT-4o mini.

Why it matters: OpenAI shipped a substantive model release, so this lands in the must-write band. HKR-H comes from the 80%-cheaper/nearly-o1 tradeoff; HKR-K from AIME 70.0 and Codeforces 1650; HKR-R from immediate developer cost/performance implications.

Aug 16, 2024Friday

OpenAI News

Disrupting a covert Iranian influence operation

OpenAI said it banned ChatGPT accounts tied to the Iranian influence operation Storm-2035 in August 2024 after they generated election and geopolitics content for X, Instagram, and five websites. The company identified 12 X accounts and one Instagram account; on Brookings' Breakout Scale, the operation ranked at the low end of Category 2, with most posts getting few or no likes, shares, or comments. What matters is the workflow: the models were used for long articles, comment rewrites, and English-Spanish posting, not for meaningful audience reach.

Why it matters: HKR-H lands on the covert election-influence angle; HKR-K lands on the account counts, sites, languages, and Breakout Scale 2. HKR-R lands via model-abuse governance, but the score stays at 76 because OpenAI reports no meaningful audience reach.

Aug 13, 2024Tuesday

OpenAI News

Introducing SWE-bench Verified

OpenAI released SWE-bench Verified, a human-validated subset built with the benchmark’s authors to assess real software issue resolution more reliably. The post names 3 failure modes in SWE-bench: overly narrow tests, underspecified issue statements, and unreliable environment setup; as of Aug. 5, 2024, top agents scored about 20% on SWE-bench and 43% on SWE-bench Lite. The key point is that the original benchmark can systematically underestimate coding-agent ability.

Why it matters: This is a strong benchmark release, not a routine post: OpenAI re-audited SWE-bench with the original authors, named 3 defect classes, and reported new score ceilings of 20% and 43%. HKR-H/K/R all pass because it changes how builders read code-agent leaderboards.

Aug 8, 2024Thursday

OpenAI News

Zico Kolter Joins OpenAI's Board of Directors

OpenAI appointed Carnegie Mellon professor Zico Kolter to its board on August 8, 2024, and added him to the Safety and Security Committee. The post says he will advise on critical safety and security decisions across all OpenAI projects alongside Bret Taylor, Sam Altman, and other members. The signal here is governance adding AI safety and robustness expertise, not a product launch.

Why it matters: The real signal is governance: OpenAI added a director with AI safety and robustness credentials and placed him on the Safety & Security Committee. HKR-K and HKR-R pass, but HKR-H is limited because this is a straightforward appointment notice, so it lands in low featured.

OpenAI News

GPT-4o System Card

OpenAI published the GPT-4o System Card on August 8, 2024, reporting 3 of 4 Preparedness categories as low risk and persuasion as borderline medium. The post says GPT-4o accepts text, audio, image, and video inputs, responds to audio in as little as 232 ms with a 320 ms average, and is 50% cheaper than GPT-4 Turbo in the API. The key issue for practitioners is voice safety: the card names unauthorized voice generation, speaker identification, and sensitive trait attribution, and says only models with post-mitigation scores at medium or below can be deployed.

Why it matters: This is not a routine post: it adds concrete preparedness ratings, 232ms voice latency, and a clear deployment threshold. HKR-H/K/R all pass, but it is a safety disclosure rather than a new model or major launch, so it lands as featured, not p1.

Jul 24, 2024Wednesday

OpenAI News

Improving Model Safety Behavior with Rule-Based Rewards

OpenAI said on July 24, 2024 it uses Rule-Based Rewards in the RLHF pipeline to reduce repeated human feedback for safety alignment. The post defines three response types—hard refusal, soft refusal, and comply—and says the method has been part of OpenAI’s safety stack since GPT-4, including GPT-4o mini. The key point is maintainability when policies change; the post excerpt does not disclose quantitative gains.

Why it matters: HKR-H/K/R all pass: explicit rules inside RLHF is a strong hook, and the post adds three response modes plus paper/code. I keep it in the 78–84 band because the excerpt does not disclose effect sizes, baselines, or failure-case detail.

Jul 18, 2024Thursday

OpenAI News

GPT-4o mini: advancing cost-efficient intelligence

OpenAI released GPT-4o mini on July 18, 2024 at $0.15 per 1M input tokens and $0.60 per 1M output tokens, replacing GPT-3.5 in ChatGPT. It supports text and vision, offers a 128K context window and 16K max output, scores 82.0% on MMLU and 87.2% on HumanEval. The key detail for builders is that its API version is the first to use instruction hierarchy against jailbreaks and prompt injection.

Why it matters: This is a substantive OpenAI model launch, not a minor refresh: GPT-4o mini adds $0.15/$0.60 pricing, 128K context, 16K max output, benchmark details, and instruction hierarchy, then replaces GPT-3.5 in ChatGPT. HKR-H/K/R all pass, so it lands in P1.