Artificial intelligence could make surgical video useful for far more than documenting a procedure after the fact, according to research presented at Cleveland Clinic’s AI Summit for Healthcare Professionals.
During the “Operating Rooms of the Future” session, Abhinav Khana, MD, a urologic oncologist and robotic surgeon at Mayo Clinic in Rochester, Minnesota, described several ways computer vision could extract information from surgical video, including generating operative reports, helping surgeons assess bladder lesions, and automating instrument counting.
“Video is a window into the operating room and has been for a very long time,” Dr. Khana said. “Modern advances in computer vision may allow clinicians to begin unlocking some of the data insights that we can gain from that video.”
Teaching AI to understand surgery
Dr. Khana first described computer vision models designed to identify individual phases of an operation from surgical video. In robotic prostatectomy, he said, a model identified surgical steps with accuracy in the low- to mid-90% range, approximately the level of agreement that might be expected between two urologists identifying which step of a procedure was underway.
That ability could provide the foundation for other applications, including surgical documentation.
In a study involving 172 robotic prostatectomy videos, investigators used automated step detection to generate operative reports. As the AI identified each surgical step, predetermined narrative text associated with that step was added to the report. The resulting reports were then evaluated against the surgical video itself rather than against existing surgeon documentation.
“We actually used the video as the ground truth because that is probably the clearest source of what actually happened,” Dr. Khana explained.
Overall accuracy was 87.3% for AI-generated reports compared with 72.8% for surgeon-written reports (P = .001). Meaningful discrepancies were identified in 12.7% and 27.2% of reports, respectively.
“Essentially, the topline study result was that AI achieves higher accuracy than surgeons at writing operative reports,” he said. “And this was exciting from an academic exercise perspective because this is the first time that AI-generated operative reports had ever really been explored, not just in urology or robotic surgery, but really in any aspect of any surgical domain.”
Dr. Khana characterized the findings as an early proof of concept, but said the approach raises another possibility: operative reports that remain connected to the video from which they were generated.
Instead of a static document in the medical record, a video-linked report could allow clinicians to return to the relevant footage for individual portions of an operation. A statement documenting a bladder-neck dissection, anastomosis, lymph-node dissection, or unusual finding, for example, could link back to the corresponding portion of the surgical video.
He compared the concept with radiology, where clinicians can review the images underlying a written report.
“We’re not relying on a secondary source, which is a human interpretation of what happened,” he said. “We have the ability to go back to the original source.”
Augmenting surgical decision-making
Dr. Khana also described preliminary work using computer vision to evaluate bladder lesions during endoscopy. Some lesions, he noted, are neither clearly malignant nor clearly benign, leaving urologists to decide whether a finding warrants further evaluation and biopsy.
Investigators trained a computer vision model to predict tumor histology from video obtained during transurethral resection of bladder tumor. The model achieved an area under the curve (AUC) of 0.829 in the Mayo Clinic test cohort and 0.867 in an external validation cohort from Tel Aviv Sourasky Medical Center.
He emphasized that the work remains preliminary and has not been deployed in clinical practice, but said the findings offer proof of concept that computer vision could eventually augment surgeons’ decision-making.
Automating routine tasks
Computer vision could also be used to automate some of the routine processes surrounding surgery.
“As surgeons, we all know one of the most tedious parts, or parts where we get most impatient, is at the end of surgery when we’re waiting for the team to count the instruments to make sure we didn’t lose one,” he noted. “At that point all progress stops, and we can’t finish the surgery until that process is done.”
In one experimental proof-of-concept project, researchers trained a computer vision model to detect and count surgical instruments. The work used approximately 1,000 images taken in a laboratory rather than during actual surgeries, including images in which instruments overlapped to simulate conditions at the end of a procedure.
The model was able to detect instruments even when they overlapped and could track an instrument as it was moved in real time. He said computer vision could eventually help automate such routine tasks, potentially reducing human labor requirements and augmenting safety.
Challenges to clinical use
Dr. Khana cautioned that technical feasibility does not necessarily mean a model is ready for clinical practice.
He noted that AI models can be difficult to train for uncommon situations and that subjective data can introduce noise. Models can also overfit to the data on which they were trained, making validation across diverse settings particularly important.
“It’s really important to show these models data from very diverse settings,” Dr. Khana said, pointing to the need for multi-institutional studies.
Still, he sees applications for surgical video extending beyond the examples he presented, including surgeon assessment, education, quality improvement, billing, operations, and logistics.
“I think we’re just seeing the tip of the iceberg,” he concluded.
