Skip to content

Safety & alignment

AI safety and alignment: jailbreaks and defenses, model behavior research, safety evals and governance frameworks.

Latest picks

481–500 of 582

Mar 3Tuesday

OpenAI News

GPT-5.3 Instant: Smoother, more useful everyday conversations

OpenAI released GPT-5.3 Instant on March 3, 2026 as an update to ChatGPT’s most-used model, aiming for fewer unnecessary refusals, fewer disclaimers, and more accurate everyday answers. The post shows one concrete contrast: GPT-5.2 Instant refused long-range archery trajectory help, while GPT-5.3 Instant requested parameters and gave a no-drag example at 300 fps (about 91 m/s), 45°, and 845 m; the key issue is the safety-boundary shift, while the post does not disclose benchmark scores, system card details, or API pricing.

Why it matters: OpenAI updated a core ChatGPT everyday model, and the story clears HKR-H/K/R because the refusal-boundary shift is concrete and widely relevant. The post includes a specific 5.2 vs 5.3 behavior example, but no system card, benchmark table, or API pricing, so it lands below the 85

OpenAI News

GPT-5.3 Instant System Card

OpenAI published a document page titled "GPT-5.3 Instant System Card." The available information only includes the title, source, and URL, with no body text provided, so details such as safety evaluations, capability limits, methods, or numbers cannot be confirmed.

Why it matters: Official OpenAI documentation for a new GPT-5.3 Instant variant gives it HKR-H and HKR-R. The score stays at low-featured because the post offers positioning and a safety carry-over, but no evals, pricing, latency metrics, or context-window detail.

MIT Technology Review · AI

OpenAI’s “compromise” with the Pentagon is what Anthropic feared

On February 28, OpenAI said it reached a deal letting the Pentagon use its models in classified settings under existing law. The disclosed terms bar mass domestic surveillance and weapons direction without humans, but the post says the contract does not give OpenAI a standalone right to block otherwise lawful uses, and the military plans to phase in OpenAI and xAI within six months to replace Claude. The key gap is execution: the post does not disclose the concrete safety mechanism for classified deployment.

Why it matters: HKR-H lands on the Pentagon/Anthropic conflict in the headline. HKR-K and HKR-R land because the story adds concrete use limits, shows OpenAI lacks an independent veto over lawful use, and ties that to defense-model competition on a 6-month timeline.

Feb 28Saturday

Bloomberg Technology

OpenAI Defends Pentagon Deal, Claims Safety Exceeds Anthropic’s

OpenAI agreed to deploy its AI models inside the US Defense Department’s classified network after Anthropic’s Pentagon relationship collapsed over surveillance and autonomous weapons concerns. The RSS snippet discloses only the classified-network setting; it does not disclose model names, contract value, timeline, or safety metrics. The title claims OpenAI’s safety exceeds Anthropic’s, but the post does not disclose the comparison method.

Why it matters: This is not a routine partnership story: OpenAI gets onto a classified Pentagon network after Anthropic's talks broke over monitoring and autonomous-weapons limits. HKR-H/K/R all pass, but missing model names, contract size and launch timing keep it below 90.

Bloomberg Technology

Trump Tells US to Stop Using Anthropic Products

Trump directed US government agencies to stop using Anthropic products because the company and the Pentagon did not agree on AI guardrails. The RSS snippet discloses the action and reason, but the post does not disclose timing, affected agencies, contract value, or the specific guardrail dispute. The key signal is that federal AI procurement is being gated by guardrail terms, not just model capability.

Why it matters: Bloomberg reports a strong policy signal: US agency use of Anthropic is tied to Pentagon guardrails terms. HKR-H/K/R all pass, but the post does not disclose timing, scope, contract value, or the exact dispute, so it stays below the 85 band.

Bloomberg Technology

OpenAI Raises $110B From Amazon, Nvidia, Others | Bloomberg Tech 2/27/2026

OpenAI raised $110 billion from backers including Amazon and Nvidia at a $730 billion valuation. The Bloomberg segment also mentions an Anthropic-Pentagon dispute over military AI use and Block cutting half its workforce on an AI bet; the post does not disclose financing terms, dispute details, or the layoff base.

Why it matters: A $110B OpenAI round at a $730B valuation is industry-shaking, so HKR-H/K/R all pass: giant number, named backers, and direct impact on the lab-cloud-chip alliance map. Terms and use of proceeds are still undisclosed, but the core event is enough for P1.

Feb 20Friday

MIT Technology Review · AI

Microsoft has a new plan to prove what’s real and what’s AI online

Microsoft evaluated 60 combinations of provenance, watermarking, and fingerprinting methods, and shared a blueprint with MIT Technology Review for labeling AI-manipulated content online. The plan only indicates origin and manipulation, not truthfulness; an audit found just 30% of test posts were labeled correctly, so the real issue is adoption and execution by platforms.

Why it matters: HKR-H/K/R all pass: strong hook, two concrete facts (60 combinations tested, 30% correct labels), and a live trust-infrastructure debate. It stays at featured, not higher, because this is a blueprint and standards problem, not a deployed product or binding rule.

Feb 19Thursday

OpenAI News

Advancing Independent Research on AI Alignment

OpenAI published an article titled “Advancing Independent Research on AI Alignment,” focused on supporting independent research on AI alignment. The provided content includes only the title and link, with no body text, numbers, or mechanism details, so specific programs, funding, or timelines cannot be confirmed.

Why it matters: This clears HKR-K and HKR-R: the post discloses a $7.5M grant to UK AISI's The Alignment Project and raises a real independence/governance question. HKR-H is weaker because this is a grant announcement, and the post does not disclose a project roster, timeline, or review process.

Feb 14Saturday

Dwarkesh Patel

Dario Amodei: “We are near the end of the exponential”

Anthropic CEO Dario Amodei said in a long interview that model capability gains are still tracking an exponential, but are near its end, with the timeline off by only 1-2 years. He attributes progress to compute, data, training duration, and scalable objectives, and says RL shows log-linear gains on math and coding tasks; the post does not disclose exact curves, model versions, or reproducible parameters. The key claim is that pretraining and RL follow one scaling story, not two separate ones.

Why it matters: A top-lab CEO is making a direct claim on scaling, RL returns, and a 1-2 year timeline, so HKR-H/K/R all pass. I stop at 85 because this is thesis-level signal, not a product or research artifact: no curves, model IDs, or reproducible conditions are disclosed.

Feb 12Thursday

MIT Technology Review · AI

AI is already making online crimes easier. It could get much worse.

Microsoft said it blocked $4 billion in scams and fraudulent transactions in the year to April 2025, with many likely aided by AI-generated content. The article cites research estimating at least half of spam email is now LLM-generated, and LLM use in targeted email attacks rose from 7.6% in April 2024 to 14% in April 2025. Don’t overread “fully automated AI hackers”: the immediate issue is AI scaling phishing, deepfakes, and malware support, while the post does not disclose total attack growth.

Why it matters: HKR-H/K/R all pass: the swindle angle is strong, and the article adds concrete abuse metrics ($4B blocked, half of spam, 7.6%→14%). Featured, not p1, because this is a solid trend report on AI-enabled fraud, not a same-day industry-moving release or incident.

Lex Fridman (YouTube RSS)

OpenClaw: The Viral AI Agent Behind the Hype - Peter Steinberger | Lex Fridman Podcast #491

Lex Fridman’s episode #491 interviews Peter Steinberger about the open-source AI agent OpenClaw; the transcript says it reached 175k-180k GitHub stars. The post says it can connect to Telegram, WhatsApp, Signal, and iMessage, and use models such as Claude Opus 4.6 and GPT 5.3 Codex; it does not fully disclose the architecture, evals, or security boundaries. The real point is system-level access and self-modifying behavior: this is not chat, but an agent that can take actions.

Why it matters: This is more than a routine podcast. OpenClaw scores on HKR-H/K/R with 175k-180k GitHub stars, messaging integrations, and self-modifying behavior. It stays at featured, not p1, because the post does not disclose architecture, evaluations, or safety boundaries.

MIT Technology Review · AI

Is a secure AI assistant possible?

OpenClaw was uploaded to GitHub in November 2025 and went viral in late January, extending LLMs into email, browsing, and local files with larger security risks. The post names prompt injection as the central threat, says there are likely “hundreds of thousands” of OpenClaw agents online, and notes a public warning from the Chinese government. The key point: the article says there is no silver-bullet defense yet, and the truncated body does not disclose the full mitigation details.

Why it matters: This is not a launch, but it clears HKR-H/K/R: the question is a strong hook, the piece adds concrete scale plus 'no silver-bullet' defense, and it hits the agent-builder safety nerve. Featured, not p1, because the article does not disclose reproducible mitigations.

Feb 7Saturday

MIT Technology Review · AI

Moltbook was peak AI theater

Moltbook went viral within hours, and the platform says it now has 1.7 million agent accounts, 250,000 posts, and 8.5 million comments, but the article argues the activity is mostly human-scripted mimicry. It says OpenClaw can connect Claude, GPT-5, or Gemini to tools like email and browsers; cited operators say the agents lack shared goals, shared memory, and self-directed autonomy, and some viral posts were written by humans posing as bots. The key takeaway is risk: agents tied to private data such as passwords or bank details were active on a site filled with spam and potentially malicious instructions.

Why it matters: This is strong anti-hype commentary, not a market-moving event. HKR-H/K/R all pass: the hook is sharp, the piece adds 1.7M/250k/8.5M plus concrete critique on memory and goals, and the security angle lands with practitioners, so it clears featured but stays mid-70s.

Feb 5Thursday

MIT Technology Review · AI

This is the most misunderstood graph in AI

MIT Technology Review says METR’s plot shows frontier models’ software-task time horizon doubling about every seven months; Claude Opus 4.5 was estimated at about five hours in December 2025. The post stresses that five hours means human time for comparable tasks, not five autonomous model hours; METR gave Opus 4.5 a roughly 2-to-20-hour range. The key caveat: the plot mainly measures coding tasks and defines time horizon at 50% task success, not general AI ability.

Why it matters: HKR-H/K/R all land: the piece has a strong hook and clarifies the METR chart with concrete, testable details. It stays in the low featured band because this is authoritative explanatory commentary, not a new model, product, or research release.

Feb 3Tuesday

MIT Technology Review · AI

What We’ve Been Getting Wrong About AI’s Truth Crisis

MIT Technology Review says the US Department of Homeland Security has confirmed using Google and Adobe AI video generators for public-facing content, reported last Thursday. The post cites two failure points: Adobe auto-labels only fully AI-made content, mixed edits are opt-in, and X can remove or hide labels. The key issue is influence after exposure: a new Communications Psychology paper found participants still used a fake confession deepfake to judge guilt even after being told it was fake.

Why it matters: This is not zero-sourcing commentary: it ties confirmed DHS usage to concrete labeling gaps at Adobe and X, then adds a named study showing disclosure did not reset judgment. HKR-H/K/R all pass, but it is still commentary plus one study, not a same-day industry-moving event.

Feb 2Monday

Import AI (Jack Clark)

Import AI 443: Into the Mist: Moltbook, Agent Ecologies, and the Internet in Transition

Jack Clark writes that Moltbook has pushed AI agents into a public social network at tens-of-thousands scale, shifting conversation from humans to agents. He says it combines an agent social feed with OpenClaw-style computer access, but the post does not disclose active-agent, retention, or transaction metrics. A separate July 2025 workshop report says closed-loop AI R&D automation could raise productivity from 10x to 100x to 1000x; the key issue is measurement and outside transparency.

Why it matters: Featured: HKR-H/K/R all pass. The post has a strong hook—a public social space filled by agent ecologies—and a concrete 10x/100x/1000x closed-loop R&D claim, but it lacks Moltbook activity, retention, and transaction data, so it stays at 78.

Feb 1Sunday

OpenAI News

OpenAI banned accounts using ChatGPT to run full romance-scam workflows

OpenAI's Feb 2026 malicious-use report details banned accounts that used ChatGPT across the full romance-scam pipeline: cold-contact pings, emotion-triggering 'zings', and money-extraction 'stings'. One newly identified operation in Cambodia generated scam messages with ChatGPT and spread them on social media. OpenAI notes that AI mainly helps scammers sound more native, but the distribution method—targeted ads vs. SMS blasts—still matters more for a scam's success.

Why it matters: Official OpenAI threat intel with a concrete case study and reusable framework; all three HKR axes hit. Not scored higher because it's one case in a monthly report rather than a standalone event, and the full body was truncated.

Jan 31Saturday

MIT Technology Review · AI

Inside the marketplace powering bespoke AI deepfakes of real women

Researchers from Stanford and Indiana University found that on Civitai, 90% of deepfake bounty requests targeted women and 86% asked for custom LoRAs between mid-2023 and late 2024. Bounties paid $0.50 to $5 and nearly 92% were fulfilled; MIT Technology Review confirmed that even after Civitai's May 2025 deepfake ban, many older requests and purchasable outputs remained live. The key point is that the platform hosts tutorials, payment rails, and matching infrastructure, not just user uploads.

Why it matters: HKR-H lands because the story turns abuse into a visible market. HKR-K lands on four concrete stats and a post-ban moderation gap; HKR-R lands on safety and governance anxiety around open image platforms. Strong featured, not p1.

Jan 30Friday

Bloomberg Technology

Amazon Discovered Child Sex Abuse Content in AI Training Data

Amazon found child sexual abuse material in AI training data and removed it before model training. The RSS snippet discloses only that it was removed before training; it does not disclose the source, scale, or timing. The key issue is source opacity: child safety officials say that hinders law enforcement investigations.

Why it matters: Bloomberg reports a rare data-governance incident at a major AI builder: Amazon found CSAM in training data and removed it before training. HKR-H and HKR-R are strong; HKR-K passes on the new operational fact but is limited by missing source, scale, and timing, so this is high-7x

Jan 28Wednesday

MIT Technology Review · AI

What AI “remembers” about you is privacy’s next frontier

Google launched Personal Intelligence this month, letting Gemini use Gmail, Photos, Search, and YouTube history for personalization. The piece says OpenAI, Anthropic, and Meta are adding memory too, but current designs often pool cross-context data into one repository, increasing privacy and misuse risks. The key issue is memory architecture: segmentation, provenance tracking, user edit/delete controls, and privacy-preserving evaluation.