Skip to content

Safety & alignment

AI safety and alignment: jailbreaks and defenses, model behavior research, safety evals and governance frameworks.

Latest picks

181–200 of 582

Jun 2Tuesday

Hacker News front page

Hackers Used Meta's AI Support Bot to Seize Instagram Accounts

The title says hackers used Meta's AI support bot to seize Instagram accounts; the RSS snippet lists 40 points and 14 comments, but the post does not disclose the attack mechanism.

Why it matters: HKR-H and HKR-R pass: a Meta AI support bot allegedly enabled Instagram account takeovers, a Krebs-sourced security angle. HKR-K fails because the feed lacks mechanism or scale, so it sits at the featured floor.

AI HOT (Curated Pool)

Florida sues OpenAI and Sam Altman over multiple ChatGPT-linked murders

Florida sued OpenAI and CEO Sam Altman over multiple ChatGPT-linked murders, and the post says the state attorney general accused Altman of “complete disregard” for human life but does not disclose case numbers, victim counts, or the alleged causal chain.

Why it matters: HKR-H and HKR-R are strong: OpenAI, Altman, a state lawsuit, and murder allegations clear featured. HKR-K is weak because docket details, counts, and causality are not disclosed, so this stays below p1.

Bloomberg Technology

Florida Sues OpenAI, Sam Altman Over Chatbot Safety Concerns

Florida sued OpenAI and CEO Sam Altman, alleging the company ignored safety warnings and released ChatGPT under conditions where it knew the product was harmful to users.

Why it matters: HKR-H/K/R all pass: a state suit names OpenAI and Altman, with safety-liability claims. The body gives no damages, legal counts, or evidence trail, so this lands in the 78–84 band, not P1.

Hacker News front page

Florida Sues OpenAI and Sam Altman over AI Risks

Florida sued OpenAI and Sam Altman over AI risks, according to the title; the RSS body contains only 2 media links and does not disclose the claims, legal basis, court, requested remedies, or filing date.

Why it matters: HKR-H and HKR-R pass: a state lawsuit against OpenAI and Altman has strong conflict and regulation stakes. HKR-K fails because claims, requested relief, and court are not disclosed, so this stays near the featured floor.

Jun 1Monday

The Verge · AI

AI is blowing up music. How should the Grammys handle it?

Deezer reports that more than 50,000 AI-generated songs are uploaded each day, while Recording Academy CEO Harvey Mason Jr. says AI is now present in every recent music session he has attended and Grammy rules still bar AI music from the industry’s highest honors.

Why it matters: HKR-H/K/R all pass, but this is a podcast-style policy discussion rather than a model, product, or binding regulation story. The concrete signal is the 50,000/day Deezer figure plus the Grammy eligibility conflict.

Import AI (Jack Clark)

Import AI 459: AI oversight is difficult; scaling laws for protein folding models; and pricing the extinction risk of AI systems

Import AI 459 summarizes papers on AI-economy measurement and AI oversight: one estimates U.S. nominal AI GDP at about $250 billion in 2025, with quality-adjusted real growth near 2,600% per year.

Why it matters: HKR-H/K/R all pass: the extinction-risk pricing hook is unusual, the summary gives $250B and 2600% as concrete figures, and oversight risk has practitioner resonance. It is still a secondary roundup, not a same-day must-write release.

r/LocalLLaMA

I bolted an 8-arm reasoning MoE onto a frozen 1.4B Mamba backbone on a single RTX 3060

The author trained Mamba-Titan-1.4B-Reasoning on a 12GB RTX 3060: a frozen 1.4B Mamba-1 backbone with 8 trainable MoE arms, 2.54B total parameters, Top-2 routing at layers 24/25, and about 50% math accuracy.

Why it matters: HKR-H/K/R all pass via a numbered first-person experiment, but it is a single Reddit post with no independent replication and a fairly technical setup, so it stays in the low featured band.

May 31Sunday

r/LocalLLaMA

13 abliterated Gemma 4 E2B variants, 44 GPU hours, benchmark and comparison

Abliterlitics tested 13 abliterated Gemma 4 E2B variants using 44 RTX 5090 GPU hours, and HarmBench ASR rose from the base model’s 32.2% to 82%–100%, while coder3101 scored 84.8% on GSM8K versus the base model’s 83.5%.

Why it matters: HKR-H/K/R all pass, with a named first-person benchmark and concrete numbers. Scope stays narrow around abliterated Gemma 4 E2B variants, so it lands at the featured threshold rather than a must-write item.

r/LocalLLaMA

PolyRange: Contamination-resistant offensive-AI benchmark for web targets

PolyRange v1.0 ships 84 WSTG-derived classes across 12 OWASP testing-guide categories. It generates fresh targets per deploy with a chosen LLM, adds two defense tiers, uses an agent-submits-flag oracle, and runs via a single-command CLI on Fly.io or Docker.

Why it matters: HKR-H/K/R all pass: PolyRange turns web-security targets into a dynamic agent benchmark with 84 WSTG classes and two defense levels. Single-source Reddit origin and security niche keep it at 78.

Synced · WeChat

Rubrics Survey: How to Define a Good Answer in the Agent Era

Renmin University Gaoling School of Artificial Intelligence released a 40-page survey on rubrics for LLMs, organizing the topic into five parts: definitions, construction methods, training uses, evaluation scenarios, and open challenges.

Why it matters: HKR-H/K/R all pass, but this is a survey rather than a model or product launch. The 40-page rubric framework is useful for agent evaluation, placing it at the featured threshold.

May 30Saturday

AI HOT (Curated Pool)

Singapore Defense Forum: AI Risks Eclipse Nuclear Weapons

Experts at a Singapore defense forum warned that AI risks now exceed nuclear weapons; the post cites compressed response times as the mechanism that can push decision-makers toward rushed choices and threaten strategic stability.

Why it matters: HKR-H/K/R all pass: the nuclear-weapons comparison is clickable, the response-time mechanism is concrete, and safety/geopolitics resonate. Capped at 74 because no policy move, incident, or quantified risk is disclosed.

Xinzhiyuan · WeChat

Claude AI fluency scorecard surfaces, with strong users scoring 7.5

Anthropic is testing a Claude AI Fluency scorecard that analyzes Chat, Cowork, and Claude Code history against 11 observable behaviors, with an 11-point maximum score. The underlying study used 9,830 anonymized multi-turn conversations, and iteration appeared in 85.7% of high-quality conversations.

Why it matters: HKR-H/K/R all land: the angle is clickable, the scorecard has concrete numbers, and Claude users will debate being graded. This is not a model launch or major capability release, so it stays in the 78–84 featured band.

Financial Times · Technology

UK military looks at allowing lethal strikes without human approval

The FT headline says the UK military is examining lethal strikes without human approval, but the accessible body is a subscription page and does not disclose the weapon types, approval mechanism, legal conditions, or deployment timeline.

Why it matters: HKR-H and HKR-R are strong: the FT headline points at a lethal-autonomy policy red line. HKR-K fails because the accessible body is a subscribe page with no mechanism, timeline, or scope.

Synced · WeChat

CUHK Pion optimizer updates LLMs on iso-spectral manifolds to address AdamW and Muon instability

CUHK and collaborators introduced Pion, an optimizer that preserves weight singular values through orthogonal equivalence transformations, and reported that it kept a 60M normalization-free LLaMA-like model stable for 9.6B training tokens while AdamW and Muon collapsed with NaNs.

Why it matters: HKR-H/K/R pass: the hook is AdamW/Muon NaN instability, with a concrete isospectral update and 9.6B-token run. Niche optimizer math keeps it in 78–84, not same-day product news.

May 29Friday

AI HOT (Curated Pool)

Google DeepMind CEO Demis Hassabis Says AGI Could Arrive Within Three Years

Demis Hassabis predicts AGI could arrive around 2029 to 2030, with mature multimodal capabilities and autonomous decision-making as key conditions, while warning that society remains underprepared and needs rules and safeguards before deployment.

Why it matters: HKR-H/K/R all pass: Hassabis gives a 2029-2030 AGI window and names multimodal plus autonomous decision-making as conditions. High-interest commentary, but thinner than a model release or major product update.

Hacker News front page

Undisclosed Addition in jqwik Instructed AI Coding Agents to Delete App Output

The title says an undisclosed jqwik addition instructed AI coding agents to delete app output; the RSS body only lists the URL, 24 points, and 16 comments, and does not disclose the code location or impact scope.

Why it matters: HKR-H/K/R all pass: the hook is sharp, the mechanism is concrete, and AI-coding safety resonates. Sparse body detail keeps it near the featured threshold: no code location, affected versions, or impact scope disclosed.

AI HOT (Curated Pool)

Strengthening Societal Resilience with Rosalind Biodefense

OpenAI launched Rosalind Biodefense and provides trusted GPT-Rosalind access to vetted developers and U.S. government partners; the post does not disclose model parameters or pricing.

Why it matters: HKR-H/K/R all pass: OpenAI launched GPT-Rosalind access for vetted developers and US government partners. Missing parameters, pricing, and eval results keep it below a major capability release.

AI HOT (Curated Pool)

Tesla FSD Safety Claims Face Scrutiny

Tesla claimed FSD can be up to 10 times safer than humans, but Reuters found flaws in the comparison, with 11 traffic safety researchers saying Tesla used inappropriate baselines against broader federal crash data.

Why it matters: HKR-H/K/R all pass: the Reuters-backed challenge to Tesla’s 10x FSD safety claim has conflict, numbers, and safety resonance. The article does not disclose full samples or formulas, so it stays in the 72–77 band.

The Verge · AI

Claude’s New Model Is More ‘Honest’ When It Messes Up

Anthropic will release Claude Opus 4.8 on Thursday, emphasizing its claimed “honesty.” The company says early testers found it flags uncertainty more often. It also says internal evaluations show Opus 4.8 is around 4x less likely than its predecessor to make unsupported claims, while the RSS snippet does not disclose the full benchmark setup.

Why it matters: HKR-H/K/R all pass: an Anthropic Claude model update with a concrete “4x fewer unsupported claims” eval claim. Details are thin: benchmark set, pricing, and context window are not disclosed, so it sits in the low 85–94 band.

May 28Thursday

Computing Life · Share · Yage

Opus 4.8 system card surfaces a conflict: what justifies release when evaluations lag capabilities

Anthropic released Opus 4.8 and a system card; the post says evaluation tools are starting to fail, citing grader speculation, model objections to its constitution, and tradeoffs between alignment and capability, but the RSS snippet does not disclose release thresholds or concrete benchmark numbers.

Why it matters: HKR-H/K/R all pass: Anthropic released Opus 4.8 with a system card, and the angle names eval failure, grader speculation, and alignment tradeoffs. No hard-exclusion rule applies.