OpenAI o1 reached 67% diagnostic accuracy in a Harvard ER triage trial, versus 50–55% for triage doctors; that gap sells the headline, not deployment.
My first read is not “AI beats doctors.” The test shape likely matters more than the raw model score. ER triage is not a clean diagnosis exam. It has time pressure, missing data, noisy patient descriptions, unstable vitals, and handoff liability. The article title gives Harvard and 67% versus 50–55%. The captured body does not disclose sample size, case source, inclusion rules, final-diagnosis definition, whether doctors saw identical inputs, or whether o1 could ask follow-up questions. Remove those details, and the 67% number becomes very elastic.
Practitioners should be unusually careful with this class of medical benchmark. o1-style reasoning models are good at turning structured symptoms into differential diagnoses. That is especially true when the case has already been cleaned into a vignette. NEJM, JAMA, and Nature Medicine have carried many adjacent results since GPT-4. GPT-4 was already strong on USMLE-style tasks in 2023, and Med-PaLM 2 crossed expert-acceptable thresholds on medical QA. Those setups usually start after the hard operational work has been done. Real ER triage starts earlier: decide what to ask, decide what to measure, decide what cannot wait. If o1 received curated case text, it won at synthesis and ranking. That is useful, but it is not the whole ER workflow.
The 50–55% doctor baseline also needs dissection. Triage doctors in real care do not optimize for “first guess equals final discharge diagnosis.” They optimize for risk sorting. They try not to miss the patient who deteriorates in the waiting room. If the endpoint was “final diagnosis ranked first,” humans are being judged against the wrong local objective. If the endpoint was high-risk miss rate, resource use, time to escalation, or wrong-pathway routing, the 67% may carry less clinical weight. The article does not disclose the endpoint, and that omission matters more than the model name.
I also do not buy the simple “doctor versus AI” framing. The credible product is a second reader or triage copilot. Give o1 the chief complaint, vitals, history, meds, and early labs within three minutes of intake. Ask it for five diagnoses that must not be missed, plus the missing facts that would change the ranking. The value is not moving 67% to 75%. The value is reducing rare but catastrophic misses: aortic dissection, sepsis, pulmonary embolism, ectopic pregnancy. The article gives no disease-level breakdown, so we cannot tell whether o1 won on common cases or on dangerous tails.
Data contamination is another live concern. A Harvard trial sounds serious, but if cases came from teaching files, exam banks, or published case summaries, o1 may have seen similar patterns during training. OpenAI cannot easily prove absence of near-neighbor exposure. A clean design would use unpublished, consecutive, real-world cases collected after the model’s training cutoff. It would lock prompts, temperature, information order, and adjudication rules. The captured article gives none of that. So the honest read is simple: the title gives the win rate; the body does not give reproducible conditions.
Regulatory friction is not a footnote here. The FDA’s treatment of clinical decision support software depends heavily on whether clinicians can independently review the basis for a recommendation. An LLM-generated differential diagnosis without inspectable evidence and calibrated uncertainty is hard to place inside high-risk triage. Epic, Abridge, and Nuance DAX have moved faster in documentation and summarization because the liability path is shorter. ER diagnosis is a different category. One miss carries legal, workflow, and trust costs.
So my call is deliberately restrained: o1 may already be a strong diagnostic ranker, but this article does not show it is a strong emergency triage system. The 67% versus 50–55% gap is a promising signal and an easy PR chart. I would wait for the paper or trial protocol, then inspect four fields first: sample size, consecutive enrollment, identical information for model and doctors, and high-risk miss rate. Without those, the headline has more force than the engineering evidence.