Skip to content

Safety & alignment

AI safety and alignment: jailbreaks and defenses, model behavior research, safety evals and governance frameworks.

Latest picks

581–582 of 582

Jul 17, 2024Wednesday

OpenAI News

Prover-Verifier Games improve legibility of language model outputs

OpenAI trained GPT-4-family prover-verifier games so stronger models write solutions weaker models can verify; under time-limited human review, correctness-only optimization led to nearly 2x more evaluation errors. The post says the large and small models differ by about 3 orders of magnitude in pretraining compute, and checkability training recovers about half the performance gain of correctness-only optimization; the full experimental numbers are not fully disclosed in the provided text.

Why it matters: This is a substantive OpenAI research release with HKR-H/K/R all present: novel setup, clear mechanism, and strong relevance to scalable oversight. The excerpt confirms the method and the human-evaluation effect, but not the full experimental tables, so it fits the 78–84 band, نه

May 28, 2024Tuesday

OpenAI News

OpenAI Board Forms Safety and Security Committee

OpenAI's board formed a Safety and Security Committee; that action is the only confirmed fact so far. The source provides only a title, and the post does not disclose members, authority, reporting lines, or timing. Watch governance power, not the committee name.

Why it matters: This is an official board-level OpenAI governance move with HKR-H and HKR-R. It stays in the low featured band because HKR-K is weak: the post confirms the committee exists, but gives no members, remit, reporting line, or effective date.