Artificial intelligence–assisted lesion measurements shortened reading time and increased classification agreement among readers in follow-up computed tomography examinations of patients with cancer. However, artificial intelligence assistance was associated with greater variability in individual lesion measurements compared with unassisted assessment.
In a retrospective study, researchers included follow-up chest, abdomen, and pelvis computed tomography (CT) examinations from 212 patients, encompassing 539 target lesions. Fifteen radiologists and 8 radiology residents evaluated examinations under unassisted, artificial intelligence (AI)-assisted, and expert-assisted conditions. The readers remeasured predefined target lesions on follow-up CT according to Response Evaluation Criteria in Solid Tumors (RECIST) 1.1.
In the assisted conditions, the readers received proposed follow-up measurements that they could accept, modify, or replace; AI proposals came from an automated segmentation system; whereas expert proposals were previous unassisted measurements made by participating radiologists. The primary outcomes were reading time and interobserver measurement variability. Secondary outcomes included proposal acceptance and modification, patient-level change in the sum of longest diameters (SLD), and agreement on RECIST response classification.
Mean expected reading time was 105 seconds per patient without assistance and 71 seconds with AI assistance, a reduction of 34% (36 seconds). Expert-assisted reading was faster at 56 seconds per patient but required an initial unassisted measurement, resulting in an estimated cumulative reading time of 151 seconds for the double-reading workflow.
AI assistance also increased interreader agreement on RECIST response classification by approximately 8 percentage points compared with unassisted assessment. Expert assistance increased agreement by approximately 13 percentage points. Differences in patient-level SLD change were small and showed no meaningful effect of assistance on estimated tumor burden change.
At the lesion level, AI-assisted measurements showed greater deviation from the expert-derived reference compared with unassisted measurements. Mean absolute error was 4.35 mm with AI assistance vs. 3.08 mm without assistance and 1.51 mm with expert assistance. The researchers noted that the expert-assisted condition was expected to have the lowest error because measurements from that group contributed to the reference standard.
The readers were also less likely to accept AI-generated proposals without modification. Radiologists accepted 61% of AI proposals compared with 77% of expert proposals, while residents accepted 62% compared with 81%, respectively. Substantial modifications occurred in 23% of AI-assisted vs. 11% of expert-assisted measurements among radiologists and 21% vs. 9% among residents.
In analyses by reader experience, AI assistance reduced reading time by 34 seconds and 39 seconds among radiologists and residents, respectively, compared with unassisted assessment. AI assistance increased RECIST outcome agreement by approximately 11 percentage points among radiologists and 4 percentage points among residents.
The readers used study-specific software rather than their usual picture archiving and communication systems, and patient data came from 2 institutions. Baseline target lesions were predefined, preventing assessment of interobserver differences in baseline lesion selection. The researchers also did not include new or nontarget lesion assessment, both of which could affect RECIST outcomes. Because those assessments required additional reading time and were not automated, the researchers noted that the observed time savings would likely be smaller in a complete RECIST 1.1 assessment.
The measurement variability analysis used the mean measurement of radiologists in the expert-assisted group as the reference standard, which inherently biased the analysis in favor of expert assistance. Additionally, measurements for a given patient across the three conditions came from different readers, introducing between-reader variability as a potential confounder.
The investigators concluded that AI assistance could accelerate RECIST assessment and improve consistency while avoiding the additional human workload required for a second reader, although expert assistance produced greater agreement and time efficiency.
“AI-assisted RECIST assessment may improve workflow and response classification consistency, while also providing a benchmark from expert-assisted reading for future AI development,” wrote lead study author Max J. J. de Grauw, MSc, of the Department of Medical Imaging at Radboud University Medical Center in The Netherlands, and colleagues.
The study authors reported no conflicts of interest.
Source: Radiology Advances
