Performance of a large language model on the reasoning tasks of a physician
成果类型:
Article
署名作者:
Brodeur, Peter G.; Buckley, Thomas A.; Kanjee, Zahir; Goh, Ethan; Ling, Evelyn Bin; Jain, Priyank; Cabral, Stephanie; Abdulnour, Raja-Elie; Haimovich, Adrian D.; Freed, Jason A.; Olson, Andrew; Morgan, Daniel J.; Hom, Jason; Gallo, Robert; McCoy, Liam G.; Mombini, Haadi; Lucas, Christopher; Fotoohi, Misha; Gwiazdon, Matthew; Restifo, Daniele; Restrepo, Daniel; Horvitz, Eric; Chen, Jonathan; Manrai, Arjun K.; Rodman, Adam
署名单位:
Harvard University; Harvard University Medical Affiliates; Beth Israel Deaconess Medical Center; Harvard University; Harvard Medical School; Stanford University; Stanford University; Stanford University; Harvard University; Harvard University Medical Affiliates; Massachusetts General Hospital; Harvard University; Harvard University Medical Affiliates; Brigham & Women's Hospital; Harvard University; Harvard University Medical Affiliates; Beth Israel Deaconess Medical Center; Harvard University; Harvard University Medical Affiliates; Beth Israel Deaconess Medical Center; University of Minnesota System; University of Minnesota Twin Cities; University System of Maryland; University of Maryland Baltimore; US Department of Veterans Affairs; Veterans Health Administration (VHA); VA Palo Alto Health Care System; University of California System; University of California San Francisco; Massachusetts Institute of Technology (MIT); University of Alberta; Harvard University; Harvard University Medical Affiliates; Massachusetts General Hospital; Microsoft; Stanford University
刊物名称:
SCIENCE
ISSN/ISSBN:
0036-8075; 1095-9203
DOI:
10.1126/science.adz4433
发表日期:
2026-04-30
页码:
524-527
关键词:
diagnosis
摘要:
More than 65 years ago, complex clinical diagnostic reasoning cases were introduced as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. In this study, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases across five experiments with a baseline of hundreds of physicians. We then report a real-world study comparing human expert and artificial intelligence (AI) second opinions in randomly selected patients in the emergency room of a major tertiary academic medical center. In all experiments, the LLM outperformed physician baselines and displayed continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have eclipsed most benchmarks of clinical reasoning, motivating the urgent need for prospective trials.
来源URL: