A study led by Harvard University researchers evaluated how OpenAI models perform in medical scenarios, with a focus on emergency department cases. The research team ran multiple experiments comparing diagnoses produced by artificial intelligence with those made by human clinicians.
Study design
In one experiment, 76 patients who presented to a hospital emergency department were included. The researchers compared diagnoses from two internal medicine physicians with the outputs of OpenAI’s o1 and 4o models. Two independent physicians reviewed the diagnoses in a blinded fashion — they did not know whether a given diagnosis came from an AI or a human — as noted in TechCrunch’s coverage.
Results
According to the paper published in Science, the o1 model performed better or at the same level as the two physicians in several key diagnostic measures, and it often outperformed the 4o model. Overall, the AI produced a "accurate or very close" diagnosis in 67 percent of cases based on the information provided, compared with 50–55 percent for the two human physicians under the same criteria.
The performance gap was most pronounced when limited patient information was available — precisely the situations where rapid and reliable decision-making is most critical.
Implications and caution
The authors and commentators stress that despite encouraging results, the AI is not yet ready to make life-or-death decisions unassisted in emergency departments. The findings indicate that further research into integrating AI as a diagnostic support tool is warranted. The study’s results do not imply that artificial intelligence will fully replace human physicians; rather, the technology may serve a complementary role in diagnostics, especially in information-scarce, time-sensitive settings.


