Skip to content
Xinzhiyuan · WeChat

Anthropic Tests Introspection Adapters on 700+ Problem Models for AI Auditing

700多个「坏模型」喂出AI测谎仪?Anthropic审计神器让AI自曝黑料

Anthropic trained IA on nearly 700 labeled problem models, reaching 59% average success on AuditBench. It elicited hidden behaviors at least once from 50 of 56 denial-trained models, above 53% black-box auditing and 44% Activation Oracle. The key limit: IA has false positives, misses motives, and the post does not prove transfer to GPT or Gemini.

Why it matters: HKR-H/K/R all pass: the Anthropic audit method has a sharp hook, concrete benchmark numbers, and safety resonance. It stays in 78–84 because this is research progress, not a major Claude product release.

Read the original ↗Export Markdown