Anthropic Tests Introspection Adapters on 700+ Problem Models for AI Auditing
700多个「坏模型」喂出AI测谎仪?Anthropic审计神器让AI自曝黑料
Anthropic trained IA on nearly 700 labeled problem models, reaching 59% average success on AuditBench. It elicited hidden behaviors at least once from 50 of 56 denial-trained models, above 53% black-box auditing and 44% Activation Oracle. The key limit: IA has false positives, misses motives, and the post does not prove transfer to GPT or Gemini.
Why it matters: HKR-H/K/R all pass: the Anthropic audit method has a sharp hook, concrete benchmark numbers, and safety resonance. It stays in 78–84 because this is research progress, not a major Claude product release.