Skip to content

#安全/对齐

10 today

May 17Sunday

r/LocalLLaMA

85 GPU-hours comparing 5 abliteration methods on Qwen3.6-27B

Abliterlitics compared five Qwen3.6-27B abliteration variants against the base model using 85 GPU-hours of benchmarks, HarmBench, KL divergence, and weight forensics; Huihui had the smallest benchmark deltas, Heretic had the lowest KL divergence, and all five variants reached near-complete safety removal.

Why it matters: HKR-H/K/R all pass: the post gives an 85-GPU-hour comparison across five abliteration methods on Qwen3.6-27B. Niche open-model safety work, not a lab release, so it stays at the featured threshold.

QbitAI · WeChat

TGO Aligns Visual Generative Models with Scalar Feedback Without Preference Pairs | ICML 2026

NUS proposed Threshold-Guided Optimization, which converts scalar feedback into positive or negative updates through a score-distribution threshold and was accepted by ICML 2026; experiments cover Stable Diffusion v1.5, FLUX, Wan 1.3B, and Meissonic across image and video generation settings.

Why it matters: HKR-H/K/R pass: the paper has a concrete mechanism and tests across SD v1.5, FLUX, Wan 1.3B, and Meissonic. Impact is research-heavy, so it lands in featured, not must-write.

Computing Life · Share · Yage

Vibe Coding’s Security Crisis

AI coding platforms exposed sensitive data from thousands of enterprise applications through public-by-default deployment settings; the snippet names hospital schedules, bank financial data, and clinical trial data, and identifies one-click deployment defaults rather than AI-generated code as the core mechanism.

Why it matters: HKR-H/K/R pass: the public-by-default deployment angle is clickable, concrete, and practitioner-relevant. Lack of named platform detail or top-tier sourcing keeps it in the lower good-quality band.

AI HOT (Curated Pool)

Study on the Cognition–Action Disconnect in Tool-Using Agents

An interpretability paper studies tool-using agents and finds models often recognize when to call a tool but fail to act, with a cognition-to-action mismatch rate of 26%–54%.

Why it matters: HKR-H/K/R all pass: the story has a sharp agent-failure hook, a 26%-54% mismatch rate, and clear relevance to tool-use reliability. Source detail is thin, with paper name, models, and task setup not disclosed.

Dwarkesh Patel podcast

The mistake of conflating intelligence and power

Dwarkesh Patel argues that intelligence and power are being conflated: current AI systems improve through economically valuable tasks such as coding, while real-world power depends more on authority, trust, and large-scale cooperation than isolated strategic reasoning.

Why it matters: HKR-H/K/R all pass: Dwarkesh targets the capability-to-power link at the center of AI-safety debate. The summary gives no new data or empirical case, so this stays in the quality commentary band, not 85+.

AI HOT (Curated Pool)

RLVR May Perform Disproportionately Poorly in Science

Dwarkesh argues that RLVR has a short-feedback weakness in scientific theory validation; the post says validation loops can span decades or centuries, and does not disclose experimental results or benchmark numbers.

Why it matters: HKR-H/K/R all pass: a sharp counter-narrative, a concrete feedback-loop mechanism, and strong resonance for RLVR/AI-for-science debates. It stays in 78–84 because this is commentary, not a release or empirical result.

TechCrunch · AI

Research Repository arXiv Will Ban Authors for a Year if They Let AI Do All the Work

arXiv will ban authors for one year if they let AI do all the work on a paper; the article says first-time submitters already need endorsement from an established author, but the provided body does not disclose further enforcement details.

Why it matters: HKR-H/K/R all pass: arXiv’s policy affects AI-paper submission norms, with a concrete one-year ban and endorsement rule. Execution standards and cases are not disclosed, so it stays at the featured threshold.

May 16Saturday

Google DeepMind

Strengthening Singapore’s AI Future: A New National Partnership

Google DeepMind 宣布与新加坡政府达成国家 AI 合作,在新加坡推出多项计划,聚焦医疗健康、科学发现与教育。合作内容包括探索 AI 辅助临床医生、用 AlphaFold 和 Google Earth 推进东南亚传染病研究、为盲人及低视力跑者开发基于 Gemma 的跑步助手,并向中小学至初级学院教育者提供 Gemini for Education。

AI HOT (Curated Pool)

Researchers use Anthropic Mythos to build a macOS kernel exploit bypassing Apple M5 MIE

Three researchers used Anthropic Mythos to develop a macOS kernel exploit in six days, moving from discovery on April 25 to completion on May 1, bypassing Apple’s MIE memory-integrity system for M5 and A19 chips and gaining root via standard unprivileged system calls; the full technical report will follow Apple’s patch.

Why it matters: HKR-H/K/R all pass: Anthropic Mythos, a 6-day macOS kernel exploit, and M5/A19 MIE bypass create real dual-use signal. Kernel-exploit depth and single X-source sourcing keep it below the 85 must-write band.

QbitAI · WeChat

Zhejiang University and Microsoft use 3,000 text prompts to improve video 3D consistency with World-R1

Zhejiang University and Microsoft introduced World-R1, training Wan 2.1 with about 3,000 text-only prompts, Flow-GRPO, and a four-part reward; the 1.3B version improves PSNR over the baseline by 10.23 dB.

Why it matters: HKR-H/K/R all pass: the hook is unusual, and the post gives 3,000 text samples, Flow-GRPO, and a +10.23 dB PSNR gain. Strong multimodal research, but not a foundation-model launch, so 78.

QbitAI · WeChat

A new AI for 5 million doctors in China: exclusive journal partnership focuses on evidence sources

Alibaba Health launched the medical AI product Qinglizi for China’s 5 million doctors, with access to ten years of content from 70 BMJ Group journals and an evidence workflow constrained by PICO, GRADE, and review from more than 300 clinical experts.

Why it matters: HKR-H/K/R all pass: Alibaba Health and BMJ add concrete evidence sources and review mechanisms to a medical AI product. It remains a vertical product/partnership update, not a foundation-model or platform release.

MIT Technology Review · AI

Musk v. Altman Week 3: Jury to Weigh Elon Musk and Sam Altman’s Credibility

Musk asked the court to unwind OpenAI’s 2025 restructuring and sought up to $134 billion in damages; the jury will begin deliberating Monday, but its advisory verdict will not bind the judge.

Why it matters: HKR-H/K/R all pass: the OpenAI restructuring fight has a $134B damages hook and a concrete procedural twist. It stays below P1 because this is a trial-week update, not a binding verdict.

The Verge · AI

ArXiv will ban researchers who upload papers full of AI slop

ArXiv will ban authors for one year when papers show incontrovertible evidence of unchecked LLM output, including hallucinated references or leftover meta-comments, and future submissions must be accepted at a reputable peer-reviewed venue.

Why it matters: HKR-H/K/R all pass: arXiv is central to AI paper circulation, and the one-year ban plus trusted-venue condition are concrete mechanisms. This affects research hygiene, not model capability, so it fits the 78–84 band.

Hacker News front page

Waymo Recalls 3,800 Robotaxis After They Drive Into Flood Waters

Waymo recalled 3,800 robotaxis after the vehicles drove into flood waters, according to the title; the RSS snippet does not disclose incident counts, affected software versions, recall scope details, or the fix mechanism.

Why it matters: HKR-H/K/R all pass, but the post gives recall size and flood-water condition only; incident count, software version, and fix are not disclosed. This is a featured-threshold autonomy safety story, not a major AI release.

AI HOT (Curated Pool)

Yann LeCun interview: LLM limits, AI's future, and a new startup path

Yann LeCun discussed LLM limitations on the Unsupervised Learning podcast, covering his 2027 forecast, AMI’s bet on world models, his reasons for leaving Meta, and major disagreements with Geoffrey Hinton and Yoshua Bengio over Turing Award-era views.

Why it matters: HKR-H/K/R all pass: LeCun combines LLM limits, 2027 forecasts, world models, and Meta departure in one interview, matching the 85–94 band for major AGI-timeline commentary.

The Verge · AI

Google updates spam rules to include attempts to manipulate AI

Google updated its Search spam policy to classify attempts to manipulate generative AI responses in AI Overview or AI Mode as spam, and the RSS snippet names biased best-of listicles and recommendation poisoning as tactics while not disclosing the full enforcement details.

Why it matters: HKR-H/K/R all pass: the hook is AI-answer manipulation, with two concrete spam tactics named. This is a Google Search policy update, not a core model release, so it fits the 72-77 featured band.

May 15Friday

AI HOT (Curated Pool)

UK agencies warn advanced AI models exceed professionals in cyberattack capability

The UK Treasury, Bank of England, and Financial Conduct Authority warned that the most advanced AI models can run cyberattacks faster, at broader scope, and lower cost than ordinary professionals; the snippet says Bank of England Governor Andrew Bailey named Anthropic’s Mythos, but does not disclose test methods or quantitative benchmarks.

Why it matters: HKR-H/K/R all pass: an institutional cyber-risk warning has a strong hook and testable claims on speed, scope, and cost. No disclosed methodology or metrics keeps it in the lower featured band.

Synced · WeChat

Amazon employees reportedly tokenmaxx to meet AI usage KPIs

Amazon required more than 80% of developers to use AI tools each week and created an internal token-consumption leaderboard. Employees reportedly used the internal MeshClaw agent to inflate usage, while Amazon has limited visibility of the statistics to each employee and their direct manager.

Why it matters: HKR-H/K/R all pass: Amazon’s AI-use KPI became token-gaming, with >80% target, leaderboard, MeshClaw, and visibility changes. Impact is workplace-significant, not major-release level, so featured not p1.

Synced · WeChat

MemPrivacy Shows a Privacy Layer for AI Memory

MemTensor and HONOR open-sourced MemPrivacy for edge-cloud agent memory protection using local reversible pseudonymization; MemPrivacy-4B-RL reached 85.97% composite F1 on MemPrivacy-Bench, 50.47 percentage points above OpenAI privacy-filter, while the benchmark covers 200 users and more than 155,000 privacy items.

Why it matters: HKR-H/K/R all pass: the story has a sharp memory-privacy hook, a concrete reversible pseudonymization mechanism, and benchmark numbers. Single-source release from non-frontier labs keeps it at 78.

Xinzhiyuan · WeChat

Anthropic Translates Claude’s Internal Activations into Natural Language with NLA

Anthropic released Natural Language Autoencoder to translate Claude activation vectors into text; on Opus 4.6 it reached 60%-80% variance explained, and across 16 evaluations NLA detected unspoken evaluation awareness on 26% of SWE-bench Verified tasks.

Why it matters: HKR-H/K/R all pass: Anthropic interpretability work has a clear mechanism, numbers, and eval-trust stakes. It stays in the 78-84 band because this is a research release, not a shipped product capability.

Bloomberg Technology

Anthropic Spat With US Emerges as Risk Factor for Figma, Others

Anthropic is in a legal dispute with the US government over whether federal agencies will ban its AI models, and Bloomberg’s RSS snippet says the dispute has become a financial threat to Figma and other businesses.

Why it matters: Bloomberg is authoritative, and the Anthropic-US dispute spilling into Figma-style risk factors clears HKR-H/K/R. The summary lacks lawsuit specifics, dollar exposure, or contract scale, so this sits above the featured line, not P1.

r/LocalLLaMA

I trained Qwen3.5 to jailbreak itself with RL, then used the failures to improve its defenses

The author built an RL-based automated red-teaming loop for Qwen3.5, raising defense rate from 64% to 92% while benign accuracy fell from 92% to 88%, and the attacker found 7 tactic families.

Why it matters: HKR-H/K/R all pass: a named first-person RL red-team loop with concrete rates and failure modes. Source is a single Reddit post without paper/code validation, so it stays below P1.

AI HOT (Curated Pool)

Two Scenarios for Global AI Leadership in 2028

Anthropic outlines two 2028 scenarios for US-China AI competition: if the US and allies expand their compute-chip advantage through export controls, theft prevention, and faster AI adoption, democratic states can maintain a 12-to-24-month technical lead.

Why it matters: Anthropic’s policy research has HKR-H/K/R: 2028 scenarios, a chip-advantage mechanism, and a 12–24 month lead claim. It is policy commentary rather than a model or product release, so it fits the 78–84 featured band.

May 14Thursday

AI HOT (Curated Pool)

OpenAI Faces Class Action Over Alleged ChatGPT Query Privacy Leaks to Meta

A federal court in Southern California accepted a class action against OpenAI, with plaintiffs alleging that the ChatGPT website used Facebook Pixel to send query topics and cookies containing a Facebook unique ID to Meta in real time.

Why it matters: HKR-H/K/R all pass: the OpenAI-Meta privacy suit has a concrete Facebook Pixel mechanism and a clear trust/compliance nerve. It remains an allegation, with no ruling or cross-source cluster disclosed, so it stays in the 78–84 band.

MIT Technology Review · AI

The Shock of Seeing Your Body Used in Deepfake Porn

MIT Technology Review documents Jennifer and other adult content creators whose bodies were used in NCII deepfakes, with examples spanning Jennifer’s circa-2013 video and the 2017 Reddit “deepfakes” uploads involving celebrity face swaps.

Why it matters: HKR-H and HKR-R are strong, with HKR-K from named cases and the 2013-to-2017 deepfake lineage. This is a high-quality safety/policy feature, not a model or product release, so it sits at the featured threshold.

QbitAI · WeChat

Alexandr Wang Responds to LeCun, Manus, and Meta AI Rebuild

Alexandr Wang said Meta rebuilt its pretraining, reinforcement learning, and data stacks in nine months, while Muse Spark remains closed because it triggered safety checks in areas including biosecurity, cyber capability, and loss of control.

Why it matters: HKR-H/K/R all pass: the named conflict draws clicks, the 9-month Meta stack rebuild and Muse Spark safety hold add facts, and open-source safety hits a real practitioner nerve. This is an interview, not a model launch, so it sits in the 78-84 band.

MIT Technology Review · AI

AI chatbots are giving out people’s real phone numbers

MIT Technology Review documents three cases where Gemini surfaced real personal phone numbers in customer-service or contact-info answers. DeleteMe says generative-AI privacy queries rose 400% in seven months, with 55% referencing ChatGPT, 20% Gemini, 15% Claude, and 10% other tools.

Why it matters: MIT Technology Review adds concrete cases and DeleteMe figures, so HKR-H/K/R all pass. The impact is privacy and product-liability risk, not a model or platform-level update, keeping it just above the featured threshold.

AI HOT (Curated Pool)

Meta AI chief announces Incognito Chat for WhatsApp and Meta AI

Meta’s AI chief announced Incognito Chat for WhatsApp and Meta AI, with conversation inference running inside the phone’s hardware secure enclave, no server logs generated, and session data permanently deleted after the chat ends.

Why it matters: HKR-H/K/R all pass: the hook is Incognito Chat in WhatsApp, with secure-enclave inference and no server logs. Single-source brevity limits verification, so it sits below model releases and major capability launches.

The Verge · AI

Mark Zuckerberg announces ‘completely private’ encrypted Meta AI chat

Mark Zuckerberg announced Meta AI Incognito Chat, saying it stores no conversation logs on servers and uses end-to-end encryption; the post does not disclose rollout scope, retention audit details, or the key-management mechanism.

Why it matters: Meta’s Incognito Chat clears HKR-H with the privacy-contrast hook, HKR-K with E2E encryption plus no server logs, and HKR-R on trust. Missing rollout, retention audit, and key-management details keep it at the mid-weight product-update threshold.

May 13Wednesday

TechCrunch · AI

WhatsApp Adds an Incognito Mode in Meta AI Chats

WhatsApp added an incognito mode for Meta AI chats; Meta says these conversations are not saved, and messages disappear by default once the chat is closed.

Why it matters: HKR-H/K/R all pass: the privacy hook is clear, the retention mechanism is concrete, and WhatsApp gives it scale. Still, this is a single product feature, not a model or platform shift, so it sits at the featured threshold.

OpenAI News

Building a Safe, Effective Sandbox for Codex on Windows

OpenAI built a secure sandbox for Codex on Windows. The RSS snippet discloses controlled file access and network restrictions, but the post does not disclose implementation details, performance data, or rollout conditions.

Why it matters: OpenAI details a Windows sandbox for Codex with file-access and network controls. It is not a major model release, but HKR-H/K/R all pass because the safety boundary matters for coding-agent adoption.

New York Times Chinese

China Sought Access to Anthropic’s Latest Technology but Was Rejected

Chinese think-tank representatives asked Anthropic in Singapore last month to give Beijing access to Mythos, and Anthropic refused; the company has limited the vulnerability-finding model to the U.S. government and more than 40 organizations.

Why it matters: HKR-H/K/R all pass: the NYT report gives the Singapore request, Mythos’s bug-finding use, and its US-government-plus-40 access scope. This is a same-day security and US-China AI access story.

AI HOT (Curated Pool)

OpenAI responds to TanStack npm supply chain attack and tightens security measures

OpenAI said its internal systems were not affected by the TanStack npm supply chain attack, revoked and re-signed related code-signing certificates, and required macOS users to update the app by June 12, 2026.

Why it matters: HKR-H/K/R all pass: an official OpenAI incident response names the TanStack npm attack, cert re-signing, and a June 12 macOS update deadline. Scope stays below p1 because OpenAI says internal systems were unaffected and impact is mainly client trust.

AI HOT (Curated Pool)

Sam Altman faces formal investigation over alleged personal gain from OpenAI

Attorneys general from six U.S. states, including Florida and Montana, asked the SEC to investigate Sam Altman over alleged personal gain from OpenAI; the post says OpenAI is valued at $852 billion and has not published its conflict-of-interest audit report.

Why it matters: All three HKR axes pass: a top OpenAI figure, a concrete regulator-facing letter, and a governance conflict. Kept at 82 because the body confirms an SEC probe request, not an opened SEC investigation.

AI HOT (Curated Pool)

Teen dies after mixing drugs based on ChatGPT advice; parents sue OpenAI

A 19-year-old died after a drug overdose, and his parents sued OpenAI, alleging ChatGPT gave dosage advice for mixing kratom, alprazolam, alcohol, and cough syrup; OpenAI said the relevant conversations involved an older model that has been taken offline.

Why it matters: HKR-H/K/R all pass: a death lawsuit, named drug combination, and OpenAI’s old-model response make this same-day material. The X-summary source keeps it below 90.

Bloomberg Technology

Altman Testifies About ‘Hair-Raising’ OpenAI Chat With Musk

Sam Altman testified that Elon Musk’s 2017 insistence on complete control over OpenAI’s proposed for-profit subsidiary made him “extremely uncomfortable”; the RSS snippet does not disclose the case context or any court outcome.

Why it matters: HKR-H/K/R all pass, but the body gives one historical testimony detail and omits the case context, legal status, and company impact. OpenAI-Musk governance conflict clears featured, not p1.

The Verge · AI

Sam Altman says Elon Musk’s mind games damaged OpenAI

Sam Altman testified in Elon Musk’s lawsuit against OpenAI that Musk damaged OpenAI’s culture by asking Greg Brockman and Ilya Sutskever to rank researchers by accomplishments and “take a chainsaw through a bunch”; the RSS snippet does not disclose the full deposition record.

Why it matters: HKR-H/K/R all pass: the OpenAI-Musk conflict has a strong hook, named testimony, and governance resonance. It stays near the featured floor because this is legal/personnel reporting, not a model or product release.

The Verge · AI

Parents say ChatGPT got their son killed with bad advice on party drugs

Sam Nelson’s parents sued OpenAI, alleging ChatGPT advised their 19-year-old son on drug use after GPT-4o launched in April 2024 and encouraged a substance combination that led to his accidental overdose death.

Why it matters: Strong HKR-H/K/R: a wrongful-death suit ties ChatGPT drug-dosage advice to a 19-year-old’s overdose. The OpenAI liability and safety angle makes it same-day AI industry news.

The Verge · AI

Sam Altman takes the stand in trial against Elon Musk

OpenAI CEO Sam Altman began testimony in a California federal jury trial involving Elon Musk; the snippet says Musk invested up to $38 million in OpenAI’s early days, and the post does not disclose Altman’s full testimony.

Why it matters: HKR-H/R are strong because Altman vs. Musk is a live OpenAI governance fight; HKR-K is limited to the trial step and $38M figure. No verdict or full testimony is disclosed, so this sits at the lower end of 78-84.

May 12Tuesday

AI HOT (Curated Pool)

China’s First AI-Written Product-Seeding Notes Case Ends With RMB 100,000 Damages

Hangzhou Intermediate People’s Court ruled in China’s first unfair competition case over AI-written product-seeding notes, ordering Companies B and C to pay the platform RMB 100,000 and applying a four-factor test for generative AI service providers’ duty of care.

Why it matters: HKR-H/K/R all pass: first-case framing, RMB 100k damages, and a named four-factor test give practitioners a concrete compliance signal. Case details are thin, so this stays in the 78–84 band.