Prover-Verifier Games improve legibility of language model outputs
OpenAI trained GPT-4-family prover-verifier games so stronger models write solutions weaker models can verify; under time-limited human review, correctness-only optimization led to nearly 2x more evaluation errors. The post says the large and small models differ by about 3 orders of magnitude in pretraining compute, and checkability training recovers about half the performance gain of correctness-only optimization; the full experimental numbers are not fully disclosed in the provided text.
Why it matters: Training a large prover against a small verifier cut human evaluation error from double to half, offering a training pattern that transfers to other scalable oversight problems.