As artificial intelligence (AI) becomes increasingly capable of supplying medical information, reasoning through clinical problems, and potentially performing medical tasks, physicians still need clinical judgment to determine when and how to use those capabilities, according to Jonathan H. Chen, MD, PhD, director for medical education in artificial intelligence and associate professor of medicine (biomedical informatics) of biomedical data science at Stanford University.
Speaking at Cleveland Clinic’s AI Summit for Healthcare Professionals, Dr. Chen used a series of examples to explore what advances in AI mean for physicians and their role in patient care. Taken together, they underscored that technical performance alone does not determine AI’s value in practice. Physicians still must assess when to trust or verify its output, how to incorporate it into clinical reasoning, and when reliance on the technology could affect how they practice.
“More than knowledge, you need judgment to tell what is real, when not everything is what it appears to be,” said Dr. Chen.
Using AI for clinical problems
AI’s ability to answer medical questions has rapidly improved, according to Dr. Chen. He noted that only a few years ago, large language models that his students evaluated scored around 33% on medical examination questions. GPT-3.5 later reached about 60%, which he characterized as barely passing, while more recent frontier models can score 90% or higher.
But examination performance is only a proxy for clinical practice, Dr. Chen said, raising a more consequential question: What happens when physicians use these increasingly capable systems to reason through actual clinical problems?
Dr. Chen shared a randomized clinical trial of 50 physicians in which access to GPT-4 in addition to conventional resources did not significantly improve diagnostic reasoning compared with conventional resources alone (76% vs 74%). In an exploratory analysis, however, GPT-4 used alone scored 16% higher than physicians using conventional resources. The finding highlighted a disconnect between strong standalone AI performance and the performance of physicians using it as a diagnostic aid. “The human-in-the-loop in the middle kind of seems to be slowing the computer down,” he said.
However, Dr. Chen emphasized that the effect of AI assistance can differ by clinical task. In a second randomized trial of 92 physicians, those using GPT-4 in addition to conventional resources scored 6.5% higher on complex patient-management reasoning tasks than physicians using conventional resources alone. Their performance did not significantly differ from GPT-4 when used alone. The contrasting trials illustrated that the effect of AI assistance on physician performance can differ by clinical task, suggested Dr. Chen.
When AI misses what matters
Dr. Chen also emphasized that strong AI performance does not necessarily translate into safe or effective clinical use. He described research currently under review that used hundreds of real-world electronic consultations in which primary care physicians sought advice from specialists. In one case, a 57-year-old patient with a smoking history developed multiple seborrheic keratoses, a presentation that the specialist said warranted further evaluation for possible malignancy. According to Dr. Chen, 20% of the frontier models recognized something was wrong, while none of 10 general internal medicine physicians identified the concern.
Physicians performed better when given AI assistance, he added, but their performance still fell short of what some AI systems achieved when used independently. The results also demonstrated “the potential of what could have happened vs the reality where you just drop technology in somebody’s lap that does not actually solve problems.” For Dr. Chen, the discrepancy highlighted why simply giving physicians access to a capable AI model is insufficient. Successful use also requires attention to change management, integration, and workflow, he said.
The potential consequences of getting that implementation wrong can be substantial. Dr. Chen noted that potentially serious errors were often omissions rather than dramatic hallucinations—for example, failing to recommend appropriate counseling, referral, treatment, or follow-up.
Verifying what AI presents
Clinical judgment also matters when AI is used to retrieve or synthesize the information on which clinical decisions may be based. Dr. Chen pointed to retrieval-based tools that search medical literature and generate answers linked to supporting sources. “It will read those articles and answer based on what they say with links back to the sources,” he said. “You can trust but verify.”
Summarizing an entire medical record also carries risks. Dr. Chen said large language models can struggle with numbers and temporal relationships, potentially confusing whether an event occurred before or after surgery or whether it occurred days vs months earlier.
They can also perpetuate what Dr. Chen called “chart lore,” an unsupported characterization copied repeatedly through the medical record that may be summarized by the AI as established fact simply because previous clinicians repeated it. “The AI will happily summarize what the last few doctors were saying and very confidently say the same thing,” he said.
When assistance becomes dependence
The need for independent judgment also extends to how reliance on AI may affect physicians’ own performance. Dr. Chen described a study involving endoscopists using computer vision during colonoscopy. According to him, AI assistance initially improved polyp detection, but after the technology was removed, physicians performed worse than they had before using it. “They got used to it. They started to depend [on it]. They had automation bias,” said Dr. Chen.
The stakes of determining when to rely on AI may increase as the technology moves beyond providing information and begins performing tasks, according to Dr. Chen. He described an AI agent attempting to complete a consultation in a simulated electronic health record, an application he said was “clearly not good enough for real medical practice yet.” He also pointed to emerging systems intended to perform more limited medical tasks, including medication refills without human double-checking.
As AI takes on a larger role in clinical care, questions about what the technology can do are increasingly accompanied by questions about how physicians should use it.
“These are very powerful, really good tools, but they are not an oracle,” he said. “You have to really be careful about how you use them in real practice.”
