Skip to content
AI HOT (Curated Pool)

OpenAI releases GPT-5.6 medical evaluation: smallest Luna variant beats GPT-5.5 at lowest reasoning strength, 25× cheaper

OpenAI 发布 GPT-5.6 系列医疗评估结果

OpenAI had specialists write answers with unlimited time and web access, then other doctors blind-rated them against GPT-5.6 across 20,000 scores on accuracy, communication, completeness, instruction-following, and health-decision helpfulness. All GPT-5.6 models outperformed doctors significantly, and doctors found fewer flaws in GPT-5.6 answers than in peer-written ones. The smallest variant, GPT-5.6 Luna, surpassed the highest-reasoning GPT-5.5 at its lowest reasoning strength while costing 25× less; the largest variant, GPT-5.6 Sol, set a new high bar. The post doesn't disclose the disease mix or specialist composition tested.

Why it matters: OpenAI ran 20,000 blind ratings pitting GPT-5.6 models against specialist physicians across five dimensions. The smallest Luna model at minimum reasoning effort already beat GPT-5.5 at max effort, and doctors flagged more issues in peer-written answers than in GPT-5.6's. The e...

Read the original ↗Export Markdown