Skip to content

Benchmarks

Who is actually stronger: benchmark results, methodology disputes and leaderboard changes.

Latest picks

441–453 of 453

Mar 20, 2025Thursday

OpenAI News

Introducing next-generation audio models in the API

OpenAI released three API audio models on March 20, 2025: gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-mini-tts. The post says the STT models beat Whisper v2 and v3 on FLEURS and other benchmarks across 100+ languages, while the TTS model adds style control but stays limited to monitored preset synthetic voices. The key shift is controllable TTS plus lower WER; the post does not disclose pricing or latency figures.

Why it matters: OpenAI shipped 3 API audio models with concrete benchmark and mechanism details, so HKR-H/K/R all pass and it clears featured. I kept it at 84, not 85+, because price, latency, and a fuller benchmark table are not disclosed.

Jan 22, 2025Wednesday

OpenAI News

Trading Inference-Time Compute for Adversarial Robustness

OpenAI reports that o1-preview and o1-mini often drive adversarial attack success rates close to zero as inference-time compute increases. The paper tests math tasks, SimpleQA prompt injection, Attack Bard images, and StrongREJECT misuse prompts; it labels the result as preliminary, and the truncated post does not fully disclose all failure cases. The key point is that this gain comes from longer reasoning at inference, not adversarial training.

Why it matters: Strong HKR-H/K/R: the hook is counterintuitive, the paper proposes a concrete mechanism, and it lands on a real safety/deployment nerve. I kept it at 82, not p1, because the post frames this as initial evidence and the excerpt does not fully disclose failure modes, cost tradeoffs

Dec 5, 2024Thursday

OpenAI News

Introducing ChatGPT Pro

OpenAI launched ChatGPT Pro at $200 per month, with unlimited access to OpenAI o1, o1-mini, GPT-4o, Advanced Voice, and a higher-compute o1 pro mode. The post specifies a stricter 4/4 reliability metric, where a question counts only if the model answers correctly in all four attempts, but it does not disclose concrete quotas or latency figures. The key signal is compute tiering: longer reasoning time is now a paid product feature.

Nov 21, 2024Thursday

OpenAI News

Advancing red teaming with people and AI

OpenAI published 2 papers on Nov 21, 2024, outlining its external human red teaming process and a new automated red teaming method. The post discloses 3 concrete design choices for external testing—threat-model-based team selection, versioned model access, and structured feedback via API or ChatGPT interfaces—but this excerpt does not fully disclose the automated method's metrics or results.

Why it matters: HKR-K carries this story: OpenAI describes 2 papers and at least 3 reusable human red-team design choices. HKR-R also passes because safety and eval teams can apply the workflow; HKR-H is weaker, and the excerpt does not fully disclose automated-red-team results, so this sits at

Oct 30, 2024Wednesday

OpenAI News

Introducing SimpleQA

OpenAI open-sourced SimpleQA, a 4,326-question benchmark for factual short-answer QA and model calibration. Two independent AI trainers verified each item; a 1,000-question audit showed 94.4% agreement and an estimated inherent error rate near 3%. The key signal: it is built to challenge frontier models, and the post says GPT-4o scores below 40%.

Why it matters: This is not a routine paper post. HKR-H comes from the inversion that a 'simple' benchmark stumps frontier models; HKR-K comes from the dataset size, agreement rate, and irreducible-error estimate; HKR-R comes from the ongoing industry fixation on hallucination and calibration,so

Oct 23, 2024Wednesday

OpenAI News

Simplifying, stabilizing, and scaling continuous-time consistency models

OpenAI introduced sCM and scaled continuous-time consistency models to 1.5B parameters on ImageNet at 512×512. The post says sCM reaches sample quality comparable to leading diffusion models in 2 sampling steps, with about 50x wall-clock speedup. Its largest model generates one sample in 0.11s on a single A100 at batch size 1 without inference optimization.

Why it matters: This clears HKR-H/K/R: the hook is 2-step sampling with diffusion-like quality, and the paper gives concrete numbers—1.5B params, ImageNet 512x512, ~50x wall-clock speed, and 0.11s per sample on one A100. Strong research release, but not a shipped product, so featured fits better

Oct 15, 2024Tuesday

OpenAI News

Evaluating fairness in ChatGPT

OpenAI analyzed millions of ChatGPT requests to test whether user names trigger harmful stereotypes, finding an overall rate of about 0.1%. The study used GPT-4o as a privacy-preserving evaluator; its gender-related judgments matched human raters over 90% of the time, while race and ethnicity agreement was lower. The key signal is model drift across versions: GPT-3.5 Turbo showed the highest task-level bias.

Why it matters: OpenAI provides a rare production-scale fairness audit with concrete rates, evaluator agreement, and a model-comparison result, so HKR-K is strong and HKR-R clears on trust and safety. This is a substantive research release, not a model launch or major product shift, so it lands

Oct 10, 2024Thursday

OpenAI News

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

OpenAI released MLE-bench, a benchmark built from 75 Kaggle competitions to measure ML engineering ability in AI agents. The best setup, o1-preview with AIDE scaffolding, reached at least Kaggle bronze-medal level on 16.9% of tasks; the benchmark code is open-source.

Why it matters: Strong HKR-H/K/R: OpenAI moves evaluation from exam-style tasks to real ML engineering, anchored by 75 Kaggle competitions and a 16.9% bronze-level result. Important as a benchmark release with concrete numbers, but still research rather than a major product launch, so featured,

Oct 1, 2024Tuesday

OpenAI News

Model Distillation in the API

OpenAI launched an API distillation workflow on October 1, 2024, letting developers use outputs from GPT-4o and o1-preview to fine-tune cheaper models such as GPT-4o mini. The suite includes Stored Completions, Evals in beta, and fine-tuning; setting store:true auto-saves input-output pairs with no added latency, per the post. Pricing includes 2M free GPT-4o mini training tokens per day and 1M for GPT-4o through October 31; Evals are free up to 7 runs per week through year-end if shared with OpenAI.

Sep 12, 2024Thursday

OpenAI News

Learning to reason with LLMs

OpenAI released o1-preview and reported 74% single-sample accuracy on AIME 2024, versus 12% for GPT-4o. The post says o1 reached the 89th percentile on Codeforces and exceeded human PhD experts on GPQA Diamond; it attributes this to large-scale RL and gains from both train-time and test-time compute. The key signal is scaling reasoning with compute, not just pretraining a larger base model.

Why it matters: This is a substantive OpenAI research release with product implications. HKR-H lands on the new reasoning line, HKR-K on the disclosed benchmark jumps and compute-scaling mechanism, and HKR-R on the direct impact to model strategy and inference economics; strong 90s, not 95+.

Aug 20, 2024Tuesday

OpenAI News

Fine-tuning now available for GPT-4o

OpenAI has opened GPT-4o fine-tuning to developers on all paid tiers, with 1M free training tokens per org per day through September 23. Training costs $25 per 1M tokens, and inference costs $3.75 per 1M input tokens and $15 per 1M output tokens on gpt-4o-2024-08-06. The signal for practitioners: partners reported 43.8% on SWE-bench Verified and 71.83% on BIRD-SQL with fine-tuned GPT-4o.

Why it matters: This is a substantive OpenAI developer release with concrete details: temporary free training quota, train/inference prices, base model version, and two benchmark datapoints. HKR-H/K/R all pass, but this is an API capability expansion, not a new frontier-model launch or platform-

Aug 13, 2024Tuesday

OpenAI News

Introducing SWE-bench Verified

OpenAI released SWE-bench Verified, a human-validated subset built with the benchmark’s authors to assess real software issue resolution more reliably. The post names 3 failure modes in SWE-bench: overly narrow tests, underspecified issue statements, and unreliable environment setup; as of Aug. 5, 2024, top agents scored about 20% on SWE-bench and 43% on SWE-bench Lite. The key point is that the original benchmark can systematically underestimate coding-agent ability.

Why it matters: This is a strong benchmark release, not a routine post: OpenAI re-audited SWE-bench with the original authors, named 3 defect classes, and reported new score ceilings of 20% and 43%. HKR-H/K/R all pass because it changes how builders read code-agent leaderboards.

Jul 17, 2024Wednesday

OpenAI News

Prover-Verifier Games improve legibility of language model outputs

OpenAI trained GPT-4-family prover-verifier games so stronger models write solutions weaker models can verify; under time-limited human review, correctness-only optimization led to nearly 2x more evaluation errors. The post says the large and small models differ by about 3 orders of magnitude in pretraining compute, and checkability training recovers about half the performance gain of correctness-only optimization; the full experimental numbers are not fully disclosed in the provided text.

Why it matters: This is a substantive OpenAI research release with HKR-H/K/R all present: novel setup, clear mechanism, and strong relevance to scalable oversight. The excerpt confirms the method and the human-evaluation effect, but not the full experimental tables, so it fits the 78–84 band, نه