Skip to content

Safety & alignment

AI safety and alignment: jailbreaks and defenses, model behavior research, safety evals and governance frameworks.

Latest picks

261–280 of 582

May 17Sunday

AI HOT (Curated Pool)

RLVR May Perform Disproportionately Poorly in Science

Dwarkesh argues that RLVR has a short-feedback weakness in scientific theory validation; the post says validation loops can span decades or centuries, and does not disclose experimental results or benchmark numbers.

Why it matters: HKR-H/K/R all pass: a sharp counter-narrative, a concrete feedback-loop mechanism, and strong resonance for RLVR/AI-for-science debates. It stays in 78–84 because this is commentary, not a release or empirical result.

TechCrunch · AI

Research Repository arXiv Will Ban Authors for a Year if They Let AI Do All the Work

arXiv will ban authors for one year if they let AI do all the work on a paper; the article says first-time submitters already need endorsement from an established author, but the provided body does not disclose further enforcement details.

Why it matters: HKR-H/K/R all pass: arXiv’s policy affects AI-paper submission norms, with a concrete one-year ban and endorsement rule. Execution standards and cases are not disclosed, so it stays at the featured threshold.

May 16Saturday

AI HOT (Curated Pool)

Researchers use Anthropic Mythos to build a macOS kernel exploit bypassing Apple M5 MIE

Three researchers used Anthropic Mythos to develop a macOS kernel exploit in six days, moving from discovery on April 25 to completion on May 1, bypassing Apple’s MIE memory-integrity system for M5 and A19 chips and gaining root via standard unprivileged system calls; the full technical report will follow Apple’s patch.

Why it matters: HKR-H/K/R all pass: Anthropic Mythos, a 6-day macOS kernel exploit, and M5/A19 MIE bypass create real dual-use signal. Kernel-exploit depth and single X-source sourcing keep it below the 85 must-write band.

QbitAI · WeChat

Zhejiang University and Microsoft use 3,000 text prompts to improve video 3D consistency with World-R1

Zhejiang University and Microsoft introduced World-R1, training Wan 2.1 with about 3,000 text-only prompts, Flow-GRPO, and a four-part reward; the 1.3B version improves PSNR over the baseline by 10.23 dB.

Why it matters: HKR-H/K/R all pass: the hook is unusual, and the post gives 3,000 text samples, Flow-GRPO, and a +10.23 dB PSNR gain. Strong multimodal research, but not a foundation-model launch, so 78.

QbitAI · WeChat

A new AI for 5 million doctors in China: exclusive journal partnership focuses on evidence sources

Alibaba Health launched the medical AI product Qinglizi for China’s 5 million doctors, with access to ten years of content from 70 BMJ Group journals and an evidence workflow constrained by PICO, GRADE, and review from more than 300 clinical experts.

Why it matters: HKR-H/K/R all pass: Alibaba Health and BMJ add concrete evidence sources and review mechanisms to a medical AI product. It remains a vertical product/partnership update, not a foundation-model or platform release.

MIT Technology Review · AI

Musk v. Altman Week 3: Jury to Weigh Elon Musk and Sam Altman’s Credibility

Musk asked the court to unwind OpenAI’s 2025 restructuring and sought up to $134 billion in damages; the jury will begin deliberating Monday, but its advisory verdict will not bind the judge.

Why it matters: HKR-H/K/R all pass: the OpenAI restructuring fight has a $134B damages hook and a concrete procedural twist. It stays below P1 because this is a trial-week update, not a binding verdict.

The Verge · AI

ArXiv will ban researchers who upload papers full of AI slop

ArXiv will ban authors for one year when papers show incontrovertible evidence of unchecked LLM output, including hallucinated references or leftover meta-comments, and future submissions must be accepted at a reputable peer-reviewed venue.

Why it matters: HKR-H/K/R all pass: arXiv is central to AI paper circulation, and the one-year ban plus trusted-venue condition are concrete mechanisms. This affects research hygiene, not model capability, so it fits the 78–84 band.

Hacker News front page

Waymo Recalls 3,800 Robotaxis After They Drive Into Flood Waters

Waymo recalled 3,800 robotaxis after the vehicles drove into flood waters, according to the title; the RSS snippet does not disclose incident counts, affected software versions, recall scope details, or the fix mechanism.

Why it matters: HKR-H/K/R all pass, but the post gives recall size and flood-water condition only; incident count, software version, and fix are not disclosed. This is a featured-threshold autonomy safety story, not a major AI release.

AI HOT (Curated Pool)

Yann LeCun interview: LLM limits, AI's future, and a new startup path

Yann LeCun discussed LLM limitations on the Unsupervised Learning podcast, covering his 2027 forecast, AMI’s bet on world models, his reasons for leaving Meta, and major disagreements with Geoffrey Hinton and Yoshua Bengio over Turing Award-era views.

Why it matters: HKR-H/K/R all pass: LeCun combines LLM limits, 2027 forecasts, world models, and Meta departure in one interview, matching the 85–94 band for major AGI-timeline commentary.

The Verge · AI

Google updates spam rules to include attempts to manipulate AI

Google updated its Search spam policy to classify attempts to manipulate generative AI responses in AI Overview or AI Mode as spam, and the RSS snippet names biased best-of listicles and recommendation poisoning as tactics while not disclosing the full enforcement details.

Why it matters: HKR-H/K/R all pass: the hook is AI-answer manipulation, with two concrete spam tactics named. This is a Google Search policy update, not a core model release, so it fits the 72-77 featured band.

May 15Friday

AI HOT (Curated Pool)

UK agencies warn advanced AI models exceed professionals in cyberattack capability

The UK Treasury, Bank of England, and Financial Conduct Authority warned that the most advanced AI models can run cyberattacks faster, at broader scope, and lower cost than ordinary professionals; the snippet says Bank of England Governor Andrew Bailey named Anthropic’s Mythos, but does not disclose test methods or quantitative benchmarks.

Why it matters: HKR-H/K/R all pass: an institutional cyber-risk warning has a strong hook and testable claims on speed, scope, and cost. No disclosed methodology or metrics keeps it in the lower featured band.

Synced · WeChat

Amazon employees reportedly tokenmaxx to meet AI usage KPIs

Amazon required more than 80% of developers to use AI tools each week and created an internal token-consumption leaderboard. Employees reportedly used the internal MeshClaw agent to inflate usage, while Amazon has limited visibility of the statistics to each employee and their direct manager.

Why it matters: HKR-H/K/R all pass: Amazon’s AI-use KPI became token-gaming, with >80% target, leaderboard, MeshClaw, and visibility changes. Impact is workplace-significant, not major-release level, so featured not p1.

Synced · WeChat

MemPrivacy Shows a Privacy Layer for AI Memory

MemTensor and HONOR open-sourced MemPrivacy for edge-cloud agent memory protection using local reversible pseudonymization; MemPrivacy-4B-RL reached 85.97% composite F1 on MemPrivacy-Bench, 50.47 percentage points above OpenAI privacy-filter, while the benchmark covers 200 users and more than 155,000 privacy items.

Why it matters: HKR-H/K/R all pass: the story has a sharp memory-privacy hook, a concrete reversible pseudonymization mechanism, and benchmark numbers. Single-source release from non-frontier labs keeps it at 78.

Xinzhiyuan · WeChat

Anthropic Translates Claude’s Internal Activations into Natural Language with NLA

Anthropic released Natural Language Autoencoder to translate Claude activation vectors into text; on Opus 4.6 it reached 60%-80% variance explained, and across 16 evaluations NLA detected unspoken evaluation awareness on 26% of SWE-bench Verified tasks.

Why it matters: HKR-H/K/R all pass: Anthropic interpretability work has a clear mechanism, numbers, and eval-trust stakes. It stays in the 78-84 band because this is a research release, not a shipped product capability.

Bloomberg Technology

Anthropic Spat With US Emerges as Risk Factor for Figma, Others

Anthropic is in a legal dispute with the US government over whether federal agencies will ban its AI models, and Bloomberg’s RSS snippet says the dispute has become a financial threat to Figma and other businesses.

Why it matters: Bloomberg is authoritative, and the Anthropic-US dispute spilling into Figma-style risk factors clears HKR-H/K/R. The summary lacks lawsuit specifics, dollar exposure, or contract scale, so this sits above the featured line, not P1.

r/LocalLLaMA

I trained Qwen3.5 to jailbreak itself with RL, then used the failures to improve its defenses

The author built an RL-based automated red-teaming loop for Qwen3.5, raising defense rate from 64% to 92% while benign accuracy fell from 92% to 88%, and the attacker found 7 tactic families.

Why it matters: HKR-H/K/R all pass: a named first-person RL red-team loop with concrete rates and failure modes. Source is a single Reddit post without paper/code validation, so it stays below P1.

AI HOT (Curated Pool)

Two Scenarios for Global AI Leadership in 2028

Anthropic outlines two 2028 scenarios for US-China AI competition: if the US and allies expand their compute-chip advantage through export controls, theft prevention, and faster AI adoption, democratic states can maintain a 12-to-24-month technical lead.

Why it matters: Anthropic’s policy research has HKR-H/K/R: 2028 scenarios, a chip-advantage mechanism, and a 12–24 month lead claim. It is policy commentary rather than a model or product release, so it fits the 78–84 featured band.

May 14Thursday

AI HOT (Curated Pool)

OpenAI Faces Class Action Over Alleged ChatGPT Query Privacy Leaks to Meta

A federal court in Southern California accepted a class action against OpenAI, with plaintiffs alleging that the ChatGPT website used Facebook Pixel to send query topics and cookies containing a Facebook unique ID to Meta in real time.

Why it matters: HKR-H/K/R all pass: the OpenAI-Meta privacy suit has a concrete Facebook Pixel mechanism and a clear trust/compliance nerve. It remains an allegation, with no ruling or cross-source cluster disclosed, so it stays in the 78–84 band.

MIT Technology Review · AI

The Shock of Seeing Your Body Used in Deepfake Porn

MIT Technology Review documents Jennifer and other adult content creators whose bodies were used in NCII deepfakes, with examples spanning Jennifer’s circa-2013 video and the 2017 Reddit “deepfakes” uploads involving celebrity face swaps.

Why it matters: HKR-H and HKR-R are strong, with HKR-K from named cases and the 2013-to-2017 deepfake lineage. This is a high-quality safety/policy feature, not a model or product release, so it sits at the featured threshold.

QbitAI · WeChat

Alexandr Wang Responds to LeCun, Manus, and Meta AI Rebuild

Alexandr Wang said Meta rebuilt its pretraining, reinforcement learning, and data stacks in nine months, while Muse Spark remains closed because it triggered safety checks in areas including biosecurity, cyber capability, and loss of control.

Why it matters: HKR-H/K/R all pass: the named conflict draws clicks, the 9-month Meta stack rebuild and Muse Spark safety hold add facts, and open-source safety hits a real practitioner nerve. This is an interview, not a model launch, so it sits in the 78-84 band.