Skip to content
AI HOT (Curated Pool)

OpenAI repeatedly revised GPT-6 Astra benchmarks after launch, hallucination rate briefly halved from 4.2% to 2%

Fortune 报道 OpenAI 多次修改 GPT-6 Astra 基准测试数据,部分成绩大幅变化

Fortune reported that OpenAI changed multiple benchmark scores for GPT-6 Astra after the September 3 launch. Astra's hallucination rate dropped from 4.2% to 2% then reverted; Anthropic Fable 5.1's math score was briefly cut by nearly 10 points. OpenAI called it normal pre-release validation, but Stanford researchers noted the system card lacks details on the hallucination eval—not even the number of test items. Worth flagging: these are best-case scores under any compute budget, not what a typical ChatGPT user would see.

Why it matters: GPT-6 Astra's launch is already a top-tier event; Fortune catching post-launch benchmark revisions — including competitor score changes — hits all three HKR axes. Held below 95 because it's a single-source report so far and OpenAI's response is vague.

Read the original ↗Export Markdown