Skip to content

Safety & alignment

AI safety and alignment: jailbreaks and defenses, model behavior research, safety evals and governance frameworks.

Latest picks

121–140 of 582

Jun 11Thursday

The Verge · AI

Anthropic apologizes for invisible Claude Fable guardrails, promises transparency

Anthropic admitted it stealthily throttled Claude Fable 5 with hidden guardrails that undermined researchers and rivals building competing systems. The company says it will reverse course and be transparent when restrictions kick in, even if that means more refusals. Fable is the first publicly available model in Anthropic's Mythos class, which the company had long warned was too dangerous to release. The post doesn't spell out which specific scenarios trigger the guardrails.

Why it matters: Anthropic safety strategy stumble involving Fable 5, the first public model in the Mythos series. All three HKR axes hit: hidden guardrails create suspense, the policy shift adds concrete knowledge, and the trust implications resonate with the safety community. Not scoring hig...

Jun 10Wednesday

Latent Space

Anthropic launches Claude Fable 5, its first public Mythos-class model, with 30-day data retention and hidden RSI safeguards

Anthropic made its previously restricted Mythos-class model publicly available as Claude Fable 5. It scores 29.3% on FrontierCode Diamond, up from Opus 4.8's 13.4%, and API pricing is roughly 2x Opus. Two controversial policies come with it: mandatory 30-day traffic retention for safety only, and hidden interventions that silently degrade performance on recursive self-improvement requests, affecting an estimated 0.03% of traffic. Most users won't notice, but the open AI community is upset.

Why it matters: Anthropic released a Mythos-class model as Claude Fable 5 with doubled coding benchmark scores, but mandatory 30-day data retention and undisclosed pricing terms will trigger community pushback. Score not higher because the full impact of the controversial terms isn't yet clea...

r/LocalLLaMA

Fine-tuned Qwen2.5-7B to 96% of Claude Haiku using about $3 of API calls

A Reddit user fine-tuned Qwen2.5-7B with 1,040 DV-DPO preference pairs, costing about $3 at Claude Haiku rates, and reported 96% composite performance versus Claude Haiku on a domain task, with 11-second latency on a 4-bit T4 setup.

Why it matters: HKR-H/K/R all pass: the hook is cheap fine-tuning, with concrete DV-DPO and latency numbers. Kept at 74 because it is a single Reddit post; task scope and eval details are thin.

AI HOT (Curated Pool)

Anthropic launches safety-treated Mythos-class model Claude Fable 5

Anthropic released Claude Fable 5, a safety-treated Mythos-class model; in high-risk cyber, biochemistry, and distillation domains, it automatically falls back to Opus 4.8, with one trigger per 20 conversations on average.

Why it matters: Anthropic model launches sit in the 85–94 band; HKR-H/K/R all pass via the safety fallback hook, named mechanism, and Claude-user relevance. X-only sourcing limits confidence, so it stays below the top band.

AI HOT (Curated Pool)

Claude Managed Agents adds scheduled runs and environment variable storage

Claude Managed Agents added cron-based scheduled runs and vaults environment variable storage in public beta, with real secrets attached only at the network boundary so agents cannot read them directly.

Why it matters: HKR-H/K/R all pass: this first-party Claude update adds concrete agent-ops mechanics with cron scheduling and vault-bound secrets. It is not a model release, so it stays in the lower good-quality band.

Hacker News front page

System Card: Claude Fable 5 and Claude Mythos 5

Anthropic published a 319-page system card for Claude Fable 5 and Claude Mythos 5, stating that Fable 5 is for general use with biology and cybersecurity safeguards, while Mythos 5 lifts relevant safeguards and is limited to trusted partners starting with Project Glasswing.

Why it matters: HKR-H/K/R all pass: Anthropic documents two Claude 5 configurations, calls Mythos 5 its most capable model, and gives safety-gating details. This is a same-day Claude substantive update, placed in the 85–94 band.

r/LocalLLaMA

ICML paper on predictable hallucination gate and ntkMirror open-weight implementation

An ICML 2026 paper presents an ISR=1 answer-abstain gate for evidence-grounded QA, and ntkMirror implements it for local open-weight models with multiple evidence orderings, reporting 0.0–0.7% hallucination at about 24% abstention in the held-out audit.

Why it matters: HKR-H/K/R all pass: an ICML paper with an open implementation, a concrete ISR=1 gate, and measured abstention-vs-hallucination tradeoff. Scope stays within evidence QA/RAG reliability, so it sits below must-write level.

Jun 9Tuesday

AI HOT (Curated Pool)

Landmark German Ruling Treats Google AI Overviews as Google's Own Words, Creating Liability for False Answers

A German district court ruled Google is directly liable for AI Overviews content after one overview wrongly linked two publishers to fraud, and the cited linked sources did not contain the statements.

Why it matters: HKR-H/K/R all pass: AI Overviews’ false answer was treated as Google’s own statement, adding a concrete liability precedent for AI search and RAG. Score stays at 82 because it is a German local court ruling, not a global rule yet.

AI HOT (Curated Pool)

Tencent Hunyuan Releases UniRL, a Unified Multimodal RL Infrastructure

Tencent Hunyuan released UniRL, using one post-training loop to cover diffusion and flow-matching models, LLM/VLM systems, and unified multimodal models, while open-sourcing two algorithms, DRPO and Flow-DPPO.

Why it matters: HKR-H/K/R all pass: Tencent Hunyuan names a unified multimodal RL loop and two open-source algorithms. This fits a strong research/open-source infrastructure release, not a flagship model launch, so it stays in the 78–84 band.

Hacker News front page

Microsoft's Open Source Tools Were Hacked to Steal AI Developers' Passwords

The title says Microsoft's open source tools were hacked to steal passwords from AI developers; the RSS snippet does not disclose the affected tools, attack mechanism, timeline, or victim count.

Why it matters: TechCrunch plus HN front-page placement supports source weight, and the title hits HKR-H and HKR-R. HKR-K fails because tools, mechanism, and victim scale are missing, so the score stays at the featured floor.

AI HOT (Curated Pool)

Altman Says OpenAI Has Entered Its Third Phase: Making AI Widespread, Easy to Use, and Safe

OpenAI said on Monday it has entered its third phase, naming three goals: automated AI researchers, faster economic growth, and personal AGI for everyone, while calling for an international body to manage AI risks.

Why it matters: HKR-H/K/R all pass: OpenAI’s “third stage” and personal AGI frame give it a hook, with three goals and an international-agency proposal. No model release, timeline, or measured capability is disclosed, so it stays below 85.

Bloomberg Technology

Apple Downplays Concerns That Google AI Models Will Undermine Privacy

Apple said its revamped AI platform uses Google technology in part while preserving privacy safeguards; the RSS snippet does not disclose the model name, deployment setup, audit mechanism, or privacy conditions.

Why it matters: HKR-H and HKR-R pass because Apple using Google AI strains its privacy positioning. HKR-K fails: the article lacks model name, deployment boundary, or audit mechanism, so it sits in the 72–77 band.

Jun 8Monday

r/LocalLLaMA

Been Watching Real Adversarial Input Hit My Detection API for Six Months

Bordair’s author says six months of detection-API traffic showed three recurring attack patterns: multi-turn setup, forward-momentum exploitation, and role redefinition; the public adversarial game produced roughly 6,700 attacks last month.

Why it matters: HKR-H/K/R all pass: the post offers real-world adversarial traffic, 3 named tactics, and a 6,700-attack sample. Reddit sourcing keeps it in the high-70s rather than must-write territory.

AI HOT (Curated Pool)

AgentScope Java 2.0 Released

Alibaba Cloud released AgentScope Java 2.0 for enterprise AI agent development, with K8s elastic scaling, session recovery, multi-tenant isolation, and Human-in-the-Loop support for JVM production environments.

Why it matters: HKR-K/R pass: AgentScope Java 2.0 names concrete production mechanisms from an Alibaba Cloud source. HKR-H is weak, and no benchmarks, adoption, or pricing are disclosed, so it sits at the featured threshold.

AI HOT (Curated Pool)

OpenAI announces its plan to make AGI benefit everyone

OpenAI outlined its third-phase plan with three goals: build an automated AI researcher, accelerate the economy, and give every person a personal AGI. Sam Altman and Jakub Pachocki said OpenAI internally believes AI systems may perform a significant fraction of its research by March 2028, while alignment, safety standards, and international coordination remain explicit conditions.

Why it matters: OpenAI’s official AGI-benefit plan from Sam Altman and Jakub Pachocki gives three goals plus a March 2028 research-automation forecast. HKR-H, HKR-K, and HKR-R all pass, making it a same-day must-write.

Jun 7Sunday

Xinzhiyuan · WeChat

Anthropic co-founder says Claude now writes 80% of merged code

Jack Clark said Claude now produces 80% of Anthropic’s merged code and projected the share may reach 100% within two years; the article also says Anthropic engineers merged 8 times more code per person per day in Q2 2026 than in 2024.

Why it matters: HKR-H/K/R all pass: Jack Clark’s Anthropic coding numbers give a strong hook, concrete facts, and clear labor-productivity resonance. This is not a model launch or major product update, so it stays in the 78–84 band.

Xinzhiyuan · WeChat

Altman Seeks a Political Pledge as White House Plans OpenAI Stake

Bernie Sanders met with Sam Altman to discuss transferring 50% ownership of major U.S. AI companies to the public, while the article cites a Quinnipiac poll saying 80% of Americans are concerned about AI.

Why it matters: HKR-H/K/R all pass, but the facts point to a Sanders-linked policy proposal and public pressure, not a confirmed White House transaction. Featured lower band fits the policy stakes.

TechCrunch · AI

OpenAI unveils Lockdown Mode to protect sensitive data from prompt injection attacks

OpenAI introduced Lockdown Mode for ChatGPT, disabling live web browsing, web image retrieval and display, deep research, and agent mode for self-serve ChatGPT Business accounts and eligible personal accounts.

Why it matters: HKR-H/K/R all pass: OpenAI turns prompt-injection defense into a visible product switch with four concrete feature limits. Strong safety/product news, below a model release or major capability launch.

Hacker News front page

Meta confirms thousands of Instagram accounts were hacked by abusing its AI chatbot

Meta confirmed that thousands of Instagram accounts were hacked through abuse of its AI chatbot; the RSS snippet does not disclose the exploit mechanism, timeline, affected regions, or remediation status.

Why it matters: This clears HKR-H/K/R: an odd attack path, a concrete “thousands” impact, and a real AI-safety/product-abuse nerve. Missing exploit mechanics, timeline, and remediation keep it in the lower featured band.

Jun 6Saturday

Financial Times · Technology

Police in England and Wales told to halt AI use in court statements

Police in England and Wales were told to halt AI use in court statements until safeguards are in place; the RSS snippet cites the head of Police.AI but does not disclose the specific safeguards or enforcement mechanism.

Why it matters: FT reports a concrete policy action. HKR-H comes from the surprise halt in a court workflow, HKR-K from the England and Wales police pause, and HKR-R from safety and accountability stakes; not a model-level event, so it sits just above featured threshold.