Radiologists correctly identified about three-quarters of artificial intelligence-generated radiological images, with accuracy varying by imaging modality and relevant subspecialty expertise.
Robert A. Cronshaw and Michelle C. Williams, of the University of Edinburgh, conducted a United Kingdom-based online survey in which 182 radiologists reviewed 30 images and classified each as real or artificial intelligence (AI)-generated. The test included 20 AI-generated images produced by fine-tuning Stable Diffusion version 2.1 using the DreamBooth approach and 10 real images. Images included computed tomography (CT), radiography, magnetic resonance imaging (MRI), and ultrasound across several anatomical regions. The models were fine-tuned using 20 to 30 example images for each modality or target region, with training images depicting normal anatomy except for one cardiac CT case.
The primary assessment was correct classification of real and AI-generated images. The researchers also assessed performance by imaging modality, years of radiology experience, self-reported familiarity with AI, relevant subspecialty interest, and radiologists’ confidence in each classification. Overall, the median proportion of correctly classified images per respondent was 78%. Across 3,640 assessments of AI-generated images, 75% were correctly identified as synthetic; across 1,820 assessments of real images, 83% were correctly identified as real. The researchers found no statistically significant difference between classification accuracy for AI-generated and real images.
Classification performance varied by imaging modality. CT images had the lowest median accuracy at 70%, followed by MRI at 77%. Accuracy was higher for ultrasound at 88% and radiographs at 91%. Among AI-generated images specifically, radiologists correctly identified 66% of CT images, 75% of MRI images, 85% of ultrasound images, and 89% of radiographs as synthetic.
Some synthetic images were substantially more difficult to identify than others. An AI-generated head CT was correctly identified as synthetic by only 24% of respondents, while another synthetic head CT was correctly identified by 43%. In contrast, several AI-generated chest radiographs were correctly classified by more than 90% of respondents.
Relevant subspecialty expertise was associated with greater classification accuracy. Radiologists with a declared subspecialty interest correctly classified 81% of images pertaining to that subspecialty, compared with 77% among those without the relevant interest. Accuracy among radiologists with vs without relevant subspecialty expertise was 89% vs 83% for chest imaging and 79% vs 71% for cardiac imaging. The researchers did not find statistically significant differences for the other subspecialties assessed.
Years of radiology experience and self-reported familiarity with AI were not associated with classification performance. Radiologists also reported similar confidence when classifying AI-generated and real images. Greater confidence was strongly correlated with more frequent correct classification, and confidence varied by imaging modality, with higher scores for radiographs.
The study had several limitations. Images were presented in an online form, preventing radiologists from windowing or scrolling through image datasets as they could in clinical practice. The survey did not include equal numbers of AI-generated and real images across imaging modalities and anatomical regions. Respondents also knew that some images were AI-generated, which may have increased their awareness of potentially synthetic images.
The AI models were developed using images of normal anatomy except for 1 case, limiting conclusions about their ability to generate images depicting different pathologies. About half of participants had fewer than 5 years of radiology experience, and the researchers acknowledged potential selection bias related to participants’ experience, subspecialty, and previous experience with AI.
Fine-tuning a generative AI model could produce radiological images that radiologists did not consistently distinguish from real images, although performance varied by modality and subspecialty expertise.
Disclosures: Williams reported relationships with Canon Medical Systems, Siemens Healthineers, Novartis, GE Healthcare, and FEOPS and support from the British Heart Foundation. Cronshaw reported no competing interests.
Source: Clinical Radiology
