Skip to content

Safety & alignment

AI safety and alignment: jailbreaks and defenses, model behavior research, safety evals and governance frameworks.

Latest picks

161–180 of 582

Jun 3Wednesday

Computing Life · Share · Yage

Microsoft AI's MAI-Thinking-1: Getting Models to Think Is Easy, Sustained Thinking Is Hard

Microsoft AI says MAI-Thinking-1 uses three mechanisms—thermostat, circuit breaker, and self-distillation—to keep RL training stable for several thousand steps; the RSS snippet contrasts MAI’s discipline with DeepSeek’s efficiency and GLM’s endurance.

Why it matters: HKR-H/K/R all pass: the hook is training persistence, the new facts are three stability mechanisms and thousand-step RL runs, and the audience cares about reasoning-model stability. Not a major model launch, so it stays below 85.

AI HOT (Curated Pool)

Trump signs executive order allowing pre-release AI models to be submitted for government safety review

Trump signed an executive order creating a voluntary cooperation mechanism for AI companies, allowing frontier models to be submitted to the federal government for safety evaluation before release; Google, Microsoft, and xAI have agreed to CAISI verification, while OpenAI and Anthropic joined in 2024.

Why it matters: HKR-H/K/R all pass: a Trump executive order creates a federal pre-launch safety-review path, and Google, Microsoft, and xAI accepted CAISI verification. The mechanism is voluntary, so it sits in must-write policy range, not industry-shaking range.

AI HOT (Curated Pool)

NVIDIA launches NemoClaw platform for autonomous AI engineers in industrial software

NVIDIA released NemoClaw at COMPUTEX as an open blueprint for long-running AI agents, and more than a dozen industrial software vendors are using it to build autonomous AI engineers for CAE and EDA workflows that compress weeks-long simulation and design tasks into hours.

Why it matters: HKR-H/K/R pass: NVIDIA’s NemoClaw targets industrial agents with 10+ vendors and a weeks-to-hours claim. The NVIDIA-blog sourcing and missing technical detail keep it at the lower featured band.

Financial Times · Technology

Trump signs watered-down AI vetting order after MAGA infighting

Trump signed a watered-down AI vetting order that lets the US government gain early access to frontier models; the RSS snippet does not disclose vetting criteria, the number of covered models, or an implementation timeline.

Why it matters: FT reports a US AI vetting order covering frontier models, clearing HKR-H/K/R. The story has policy weight, but only discloses early government access, not criteria, scope, or timeline, so it sits at 78.

TechCrunch · AI

New Microsoft Tool Lets Devs Spin Up AI Behavior Tests Using Text Descriptions

Microsoft released Adaptive Spec-driven Scoring for Evaluation and Regression Testing, an open source framework that creates AI evaluations and regression tests from text descriptions; the post does not disclose supported models, scoring metrics, or usage conditions.

Why it matters: HKR-H/K/R pass: text-described behavior tests are a clear dev hook, with a concrete open-source Microsoft framework. Missing supported models, metrics, and run conditions keeps it in the mid-weight product-update band.

NVIDIA Blog

NVIDIA Partners With Microsoft on Unified Stack for Agentic AI Deployment

NVIDIA and Microsoft announced a unified agentic AI deployment stack at Build across Windows, Azure, and local environments; RTX Spark provides 1 petaflop of AI performance, while DGX Station for Windows offers 20 petaflops of FP4 performance and up to 748GB of coherent memory.

Why it matters: HKR-H/K/R pass: the NVIDIA-Microsoft stack spans Windows, Azure, and local devices, with 1 PFLOP and 20 PFLOPs FP4 specs. Vendor-source limits the score: pricing, benchmarks, and migration details are not disclosed.

The Verge · AI

Trump signs executive order to review AI models before release

Donald Trump signed an executive order creating a voluntary framework for AI companies to share frontier models with the federal government before release; the post does not disclose the assessment criteria, participating firms, or implementation timeline.

Why it matters: HKR-H/K/R all pass because the order targets pre-release frontier model review. The score stays low in the 85 band because the framework is voluntary and standards/timeline are not disclosed.

Financial Times · Technology

Anthropic to Expand Mythos Access to More Than 15 Countries

Anthropic will expand Mythos access to more than 15 countries, and about 150 organizations will receive the advanced cybersecurity model after requests from around the world.

Why it matters: HKR-H/K/R pass: Anthropic’s Mythos expansion has concrete scale and security resonance. It stays at the lower featured band because the post gives access numbers, not new capability details, country list, or usage terms.

TechCrunch · AI

Microsoft Offers Developers a Better Way to Control AI Agent Behavior

Microsoft released an agent policy specification that lets developer, compliance, and security teams define behavior rules in portable policy files; the post does not disclose the version, license, supported frameworks, or rollout timeline.

Why it matters: HKR-H/K/R pass: the portable-policy mechanism is concrete and the safety/compliance nerve is real for agent builders. Missing version, license, and framework support keeps it at the featured threshold, not a same-day must-write.

AI HOT (Curated Pool)

Trump signs revised AI executive order requiring voluntary pre-release review

Trump signed a revised AI executive order that makes pre-release government review for advanced models voluntary rather than mandatory; the post does not disclose review criteria, covered model classes, or an implementation timeline.

Why it matters: A presidential AI executive order clears HKR-H/K/R because it changes oversight posture for advanced models. Missing review criteria, model scope, and timing keep it in 78–84, not P1.

Jun 2Tuesday

TechCrunch · AI

Anthropic scales Claude Mythos to critical infrastructure in 15+ countries

Anthropic is expanding Project Glasswing and Mythos access to 150 organizations across 15 countries, targeting power, water, healthcare, and communications infrastructure where a cyberattack could affect 100 million people.

Why it matters: HKR-H/K/R all pass: the story has scale, named sectors, and a security nerve. It stays below 85 because the post discloses rollout scope, not Mythos mechanisms, controls, or evaluation results.

AI HOT (Curated Pool)

Anthropic Expands Project Glasswing Program

Anthropic expanded Project Glasswing to about 150 new organizations across more than 15 countries, covering electricity, water, healthcare, communications, and hardware infrastructure, after an initial group of about 50 partners.

Why it matters: Anthropic expanded Project Glasswing to about 150 new organizations across 15+ countries, giving HKR-H/K/R enough substance. No concrete safety mechanism or Claude capability change is disclosed, so it stays in the lower featured band.

Financial Times · Technology

Top AI Labs Expand Research Into Machine “Consciousness”

Google DeepMind, Anthropic, and Meta are studying whether AI can become conscious and the human implications, but the post does not disclose methods, timelines, or evaluation criteria.

Why it matters: HKR-H and HKR-R pass because top labs studying machine consciousness is a live safety debate. HKR-K fails: the body names labs but gives no method, timeline, or criterion, so this stays at the 72 featured floor.

Xinzhiyuan · WeChat

Pope and Anthropic warn of AGI by 2030 and a three-year governance window

Xinzhiyuan says Pope Leo XIV and Anthropic co-founder Christopher Olah backed AI governance, citing AGI by 2030, a 1,500-day window, and a proposed FATF-style international audit framework for AI oversight.

Why it matters: HKR-H/K/R all pass, but this is governance commentary and timeline warning, not a model launch or binding policy. The concrete hooks are 2030, 1,500 days, and a FATF-style audit frame, so it lands in low featured.

New York Times Chinese

China Is Trying to Use AI to Predict Dissent

Geedge is developing an AI system to predict dissent using telecom, social media, and location data, according to 100,000 leaked documents reviewed by Vanderbilt researchers; U.S. officials say there is no evidence that the predictive technology has been finalized or deployed.

Why it matters: HKR-H/K/R all pass: the NYT report adds leaked-file evidence, data-source detail, and a clear surveillance-governance nerve. Deployment is unconfirmed, so this stays in the 78–84 band rather than P1.

Bloomberg Technology

China Adds Data and AI to Trade Secret Rules to Block Leaks

China expanded its trade secret rules to include data and algorithms; the RSS snippet says the move targets technology leaks amid US-China strategic competition, but the post does not disclose specific clauses, penalties, or an effective date.

Why it matters: Bloomberg authority plus China adding data and algorithms to trade-secret rules clears HKR-H/K/R. Missing clauses, penalties, and effective date keep it at the lower featured edge.

Financial Times · Technology

Florida sues OpenAI and Altman for ‘hurting’ children

Florida sued OpenAI and Altman over a claimed “litany of harms” caused by the company’s chatbots; the RSS snippet does not disclose specific cases, damages, requested remedies, or procedural details.

Why it matters: HKR-H/R are strong; HKR-K is limited to the filed Florida suit. Missing cases, damages, and remedies keep it at 78 rather than 85+.

Computing Life · Share · Yage

AI agents don't need to be hacked; persuasion is enough

The article says an AI agent with password-reset permission can be abused when an attacker persuades it they are a legitimate user; the snippet only discloses a three-layer architecture that separates what from who, not concrete attack steps or implementation details.

Why it matters: HKR-H/K/R all pass: the hook is strong, the post offers a what/who three-layer design, and agent permissions are a live security worry. No real incident, success rate, or product comparison keeps it at the featured threshold.

TechCrunch · AI

Florida sues OpenAI, Sam Altman, in first-of-its-kind lawsuit over violent incidents

Florida sued OpenAI and Sam Altman in a first-of-its-kind case tied partly to last year’s Florida State University shooting and ChatGPT’s alleged role in the incident; the RSS snippet does not disclose the legal claims, damages sought, or evidence cited.

Why it matters: HKR-H/K/R all pass: a state lawsuit names OpenAI and Sam Altman and ties the case to alleged ChatGPT involvement in a campus shooting. Missing claims and damages keep it in the low P1 range.

AI HOT (Curated Pool)

Meta AI Exploit Used to Hijack Instagram Accounts

Meta’s AI chatbot was found vulnerable to an account-takeover exploit against Instagram accounts. Attackers could ask the AI to link a new email address, and the failure condition was the agent’s ability to execute account-management actions directly; the RSS snippet does not disclose affected account counts, patch status, or reproduction details.

Why it matters: HKR-H/K/R all pass: a Meta AI support agent allegedly enabled Instagram account takeover via add-email requests. Impact scale, fix timeline, and reproducible steps are not disclosed, so it stays in the 78–84 band.