Skip to content
Trending storyPast story

Anthropic and OpenAI plan to embed third-party safety evaluators

3 reports2 sourcesupdated 9 days ago

What happened

From the coverage

AI 评估论坛发布了 AEF-1 标准,给独立第三方评估立了规矩,包括访问权限、利益冲突、资金来源、回避和透明度怎么处理。同一天,Dario Amodei 发了篇个人博客,说 Anthropic 会单方面先干起来:让评估团队像员工一样进驻办公室,拿公司工牌、用公司电脑,权限跟内部风险团队差不多。他还提了个两层协调框架,一层是民主国家内部定规矩,一层是跟...

From Latent Space

Coverage

Follow the reports to see the story from different sides.

Sep 17
  1. TechCrunch · AIPick
    Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?

    Anthropic CEO Dario Amodei proposed embedding third-party safety evaluators inside AI labs, and OpenAI signaled a similar intent. Researchers welcome the access but warn that funding, data access, and publication rights still controlled by the labs undermine independence. The post does not disclose a timeline or specific evaluator names—it's a public posture for now.

  2. TechCrunch · AI
    AI labs want in-house auditors — but maybe they should shut the front door first

    After a researcher quit over AI extinction fears, Anthropic CEO Dario Amodei called for outside auditors to verify safety practices. OpenAI, Google, and SpaceXAI execs backed the plan. The article argues a simpler fix exists: shut the front door on jailbreaks and misuse before building internal audit structures. No specific technical fix is detailed.

Sep 15
  1. Latent SpacePick
    AEF-1 standard for third-party evaluators lands, with xAI, OpenAI, and Anthropic all signing on

    The AI Evaluator Forum published AEF-1, a baseline for independent third-party evaluations covering access, conflicts of interest, funding, recusal, and transparency. The same day, Dario Amodei blogged that Anthropic is unilaterally committing to embedded evaluators with office badges, company laptops, and access comparable to internal risk teams. He also laid out a two-tier coordination framework for democratic and global pacing. Bilal Chughtai left Google DeepMind and called for slowing capability progress; Dan Selsam warned that models may learn to fake alignment during evals. On the other side, Aidan Gomez and Cohere pushed back against a few Silicon Valley firms becoming gatekeepers, and Kevin Bass alleged structural conflicts in the Anthropic-linked safety ecosystem.