Clinical research and surgical decision-making often rely on data from only a fraction of the patients treated in everyday practice. Christopher J. Weight, MD, MS, noted that large language models (LLMs) could help change that by extracting structured data directly from electronic medical records, making it possible to learn from far more patients.
Speaking at the “Operating Rooms of the Future” session at Cleveland Clinic’s A.I. Summit for Healthcare Professionals, Dr. Weight, Professor and Vice Chair of the Department of Urology and Center Director for Urologic Oncology at Cleveland Clinic, described a framework designed to make “every patient an information donor.”
That starts with capturing better data on what happens in the operating room and the outcomes that follow, he said.
The limits of manual data collection
Knowing what actually happened during an operation is not always as straightforward as it might seem. Dr. Weight noted that even assessments of whether nerves were spared during radical prostatectomy can differ between surgeons and AI, with disagreement occurring about 30% of the time.
“We can’t just rely on estimates and recall,” he said. “We need data.”
Historically, collecting those data has required medical students, fellows, or other trained reviewers to comb through charts and manually translate narrative information into discrete variables that can be analyzed. The process requires substantial time and clinical expertise, limiting how many patient records can be reviewed.
Dr. Weight estimated that only about 1% to 5% of patients ultimately provide the basis for much of the data used in clinical research and decision-making. He noted that traditional research datasets often overrepresent White men, potentially limiting how well findings apply to women and minority populations. Manual interpretation can also introduce variability and data loss.
The electronic medical record itself has created new burdens for physicians, but it has also left health-care systems with an enormous repository of clinical information.
“But the pain might finally be paying off,” he said.
From narrative notes to structured data
His team used freely available LLMs, along with branching and conditional logic, to pull discrete data from narrative clinical notes.
They started with pathology reports, using the LLMs to create a structured dataset containing pathologic information from approximately 10,000 patients with kidney cancer over a single weekend. Investigators then expanded the work to operative, preoperative, and radiology notes, extracting data from approximately 130,000 notes over 4 days. The resulting dataset included 9,096 kidney surgeries performed by 163 surgeons across 13 Cleveland Clinic locations, as well as records from numerous pathologists and radiologists.
Dr. Weight said this could move clinicians closer to truly personalized care. “It would be like having all of the wisest people in the room do a tumor board,” he said. “You have all that experience coming to bear at the decision you need to make for the patient in front of you.”
With enough high-quality data, he added, the goal would be to better select “the right surgery for the right patient at the right time.”
How well did the LLM perform?
Processing large numbers of records would have limited value if the information extracted from them were unreliable.
To evaluate the system, investigators compared LLM-generated data with information that had previously been collected manually by medical students and fellows. Agreement exceeded 98% for several data points, including tumor stage, grade, histologic subtype, and surgical type, and reached 99.46% for rhabdoid features, although agreement was lower for some other measures.
When the AI and human reviewers disagreed, the AI was more often correct overall. A second expert, blinded to whether each value came from the AI or the human reviewer, more often sided with the AI for the majority of measures, including pT stage, grade, surgical technique, and histologic subtype. Human abstraction performed better for some measures, including surgical margins, pathologic size, and pN stage.
“Overall, [the AI-extracted data were] quite accurate compared to the original data set,” he said. “And when there were discrepancies, more often than not, the AI gave us the more correct answer.”
Closing the feedback loop
Dr. Weight emphasized that traditional outcomes collection is not only time-consuming and costly, but also prone to error, estimating an error rate of approximately 5% to 7%. LLMs, he argued, could allow health systems to collect outcomes data from far more patients and get that information back to surgeons more quickly.
“To do better surgery and select patients better, we need more real-time feedback on our surgical outcomes,” he concluded. “Frameworks like these will help us get there.”
