Skip to content

Research & technical reports

Official research posts and technical reports from labs: architectures, training methods, measurement and safety research. Purely academic papers are not collected.

Latest picks

261–262 of 262

Jul 24, 2024Wednesday

OpenAI News

Improving Model Safety Behavior with Rule-Based Rewards

OpenAI said on July 24, 2024 it uses Rule-Based Rewards in the RLHF pipeline to reduce repeated human feedback for safety alignment. The post defines three response types—hard refusal, soft refusal, and comply—and says the method has been part of OpenAI’s safety stack since GPT-4, including GPT-4o mini. The key point is maintainability when policies change; the post excerpt does not disclose quantitative gains.

Why it matters: HKR-H/K/R all pass: explicit rules inside RLHF is a strong hook, and the post adds three response modes plus paper/code. I keep it in the 78–84 band because the excerpt does not disclose effect sizes, baselines, or failure-case detail.

Jul 17, 2024Wednesday

OpenAI News

Prover-Verifier Games improve legibility of language model outputs

OpenAI trained GPT-4-family prover-verifier games so stronger models write solutions weaker models can verify; under time-limited human review, correctness-only optimization led to nearly 2x more evaluation errors. The post says the large and small models differ by about 3 orders of magnitude in pretraining compute, and checkability training recovers about half the performance gain of correctness-only optimization; the full experimental numbers are not fully disclosed in the provided text.

Why it matters: This is a substantive OpenAI research release with HKR-H/K/R all present: novel setup, clear mechanism, and strong relevance to scalable oversight. The excerpt confirms the method and the human-evaluation effect, but not the full experimental tables, so it fits the 78–84 band, نه