Skip to content

Reasoning

Progress in model reasoning: chain of thought, reasoning models, math and logic benchmarks and the debates around them.

Latest picks

601–603 of 603

Jul 17, 2024Wednesday

OpenAI News

Prover-Verifier Games improve legibility of language model outputs

OpenAI trained GPT-4-family prover-verifier games so stronger models write solutions weaker models can verify; under time-limited human review, correctness-only optimization led to nearly 2x more evaluation errors. The post says the large and small models differ by about 3 orders of magnitude in pretraining compute, and checkability training recovers about half the performance gain of correctness-only optimization; the full experimental numbers are not fully disclosed in the provided text.

Why it matters: Training a large prover against a small verifier cut human evaluation error from double to half, offering a training pattern that transfers to other scalable oversight problems.

Jun 28, 2024Friday

DeepSeek · API updates

deepseek-chat upgraded to DeepSeek-V2-0628, with better reasoning and role-play

The deepseek-chat model has been upgraded to DeepSeek-V2-0628. Reasoning improved, and role-play got a clear boost. HumanEval Pass@1 rose from 79.88% to 84.76%, MATH ACC@1 from 55.02% to 71.02%, and BBH from 78.56% to 83.40%.

Why it matters: The math benchmark gain is larger than the code and reasoning gains, and the win rate against GPT-4-0314 on Arena-Hard also improved.

Jun 14, 2024Friday

DeepSeek · API updates

deepseek-coder upgraded to DeepSeek-Coder-V2-0614

The deepseek-coder model has been upgraded to DeepSeek-Coder-V2-0614, with a clear gain in coding ability. The company says its code generation, code understanding, code repair and code completion reach the level of GPT-4-Turbo-0409, with strong math and reasoning, while general ability matches DeepSeek-V2-0517.

Why it matters: DeepSeek upgrades its code model and benchmarks code generation, understanding, repair and completion against GPT-4-Turbo-0409.