When artificial intelligence (AI) is applied to medical images, it can be used to improve the diagnostic capability of physicians as well as to identify many features that humans are unable to see. From identifying biomarkers to guiding treatment decisions, AI in medical imaging has already demonstrated many use cases in research settings, yet few of these tools have yet to reach the clinic.
In a session at the Cleveland Clinic A.I. Summit for Healthcare Professionals, Benjamin H. Kann, MD, PhD, Assistant Professor of Radiation Oncology at Harvard Medical School, explained some of the ways AI has transformed diagnostics and further potential impacts of AI on medical imaging.
Dr. Kann, who also directs the Kann Lab in the Artificial Intelligence in Medicine (AIM) Program at Mass General Brigham and Harvard Medical School, described how to approach different types of AI applications in health care to determine how each will impact patient management.
Tiers of AI applications
Dr. Kann laid out three tiers of AI applications for diagnostic purposes in health-care systems.
The first tier encompasses solutions that identify features that humans can see well. These applications can often simply be deployed and monitored with engineering and governance guidelines.
He explained that most tier one applications that are currently cleared by the U.S. Food and Drug Administration (FDA) for radiology are automation tools, and that there are even fewer cleared tools overall for pathology.
In tier two, the AI identifies features that humans see poorly on their own. As an example of a tier two application, he described a study of a real-world deployment for a fine-tuned pathology foundation model that could detect biomarkers in lung cancer based on histopathology alone.
Tier three, which represents AI applications that can identify features that humans cannot see at all, includes decision-making support tools. He explained that tier three algorithms have the highest actionability, but also have the biggest challenges associated with them.
He noted that there are currently no FDA-cleared tier three applications in the radiology setting. However, he did point to the predictive Artera AI assay as a successful tier three application in oncology. The multimodal AI score can predict the benefit of hormonal therapy with radiotherapy for patients with localized prostate cancer and is based on large-scale evaluations pooled from five phase III trials.
Verifiability and other challenges with high-tier applications
In 2021, Dr. Kann and colleagues described an AI translational gap defined by missing validity, usability, and utility. Five years later, he said the field has underestimated a fourth factor. “What also really matters is how verifiable these tools are,” he said. “The more AI changes what we do, and by that I mean the decisions we’re making, the less able we are to check it.”
He explained that tier one applications are easily verifiable with each case, but the verifiability decreases with the higher tiers. Tier two applications are only quantifiable in aggregate through local validation. Although this is feasible to assess in real time, it is resource-intensive. Tier three applications, on the other hand, are only auditable months or years later to determine if the AI did what it was promised to do.
This difference in verifiability drives the translational gap and the acceptance of more tier one models in health care.
Additionally, tier three applications are often subject to data set drift and shift, whereby changes in the patient population, demographics, or clinical practice arise that influence the data, making it unclear if the original training data still apply. This can cause prognostic models to degrade over time. He stressed that after models are deployed, someone has to monitor the AI.
He also explained that as biologic understanding of disease increases, statistical and training stability decreases. Without further training once a subtype is discovered, for example, the model is potentially unable to accurately predict how each subtype will respond.
“As we learn more and more, this becomes statistically a much harder problem to solve, so this gives me a lot of skepticism about some of these prognostic and predictive approaches,” Dr. Kann commented.
Foundation models, on the other hand, can reduce the need for labeled data with fine tuning, but it does not solve other issues of event scarcity, confounding, data set drift, follow-up time, or signal ceilings.
With multimodal AI models, the multimodal input can raise the signal ceiling but requires even more data to reach the ceiling, introducing more potential for confounding data, redundancy, missing data, estimation variance, and measurement noise across the different platforms for extracting data.
Opportunities with tier one applications
Dr. Kann noted that a lot of success can come from tier one applications that are linked to actionable impacts. These tools can even be used to uncover features that were not previously being measured.
For example, Dr. Kann’s team used a segmentation algorithm to measure the temporalis muscle on the routine brain MRIs of pediatric patients with brain tumors to assess body composition for prognostication. They found that 73% of survivors were sarcopenic at one or more time points, and 84% had a normal or overweight body mass index at the time. Detecting sarcopenia enabled the researchers to predict endocrine disorders, poor physical functioning, and potential mortality. The group is now planning a prospective trial (STRIVE-AI) to screen the temporalis muscle to triage patients who may need exercise and nutritional intervention.
Colleagues applied the same idea to the thymus, visible on every chest CT obtained for patients with lung cancer. An AI score of thymic health proved to be a significant determinant of immunotherapy response, and trials are being designed to use it in treatment decisions.
What to know about an AI model before acting
Dr. Kann explained that a model used in his own lab for detecting extranodal extension in CT scans of patients with head and neck cancer once output that extranodal extension was detected on a scan when it was wrong.
Before a model is deployed, Dr. Kann stressed, several aspects must be accessible or understood in order to understand how to use the output in the management of a patient. Each model should include raw output data, calibration to understand the certainty of the model, conformal prediction to allow the model to abstain on unfamiliar cases, and training data information to understand its limitations.
Clinicians should also know whether a model fails predictably. “If you know where your model is likely to fail, you can build out the specification of the model for the user,” he said. “[Unpredictable failures are] more concerning, because then you don’t really know when it’s failing and when it’s not.”
He added that clinicians should also understand the potential risks of applying the AI algorithm. “Nobody would be comfortable with an AI guiding where to make an incision if it was only correct 80% of the time,” he explained.
The tier system helps him feel comfortable deciding when to use a model, or not, for patients in the clinic. “Knowing which tier you’re evaluating will drive the questions you should ask and the resources you need for safe and effective clinical use," Dr. Kann explained.
He reiterated that tier one models are safe to use with monitoring and governance guardrails, as long as health systems budget for the monitoring as well as the tool’s license. For tier two applications, health-care professionals should validate the model on their local population and allow for uncertainty quantification. “A model that declines beats one that’s confidently wrong,” he said. And for tier three applications, focus first on cohort development before building the model and require both prognostic and predictive clinical evidence.
“The future of cancer AI diagnostics depends less on whether models can predict—and more on whether clinicians and health-care systems can use predictions responsibly,” Dr. Kann concluded. “We have a lot of models coming out. The power is there, the capability is there, but we really need to think about whether health-care clinicians and systems can use these conditions responsibly and thoughtfully. That’s really the only way it will actually help our patients.”
