Image-conditioned generative artificial intelligence produced synthetic chest radiographs that radiologists were less likely to identify as artificial than images generated from text alone, according to a blinded reader study published in Frontiers in Medicine. Radiologists identified 34% of image-conditioned synthetic radiographs as artificial compared with 56% of text-only synthetic radiographs.
The single-center retrospective study included 320 real disease-positive frontal chest radiographs and 320 age- and sex-matched normal conditioning radiographs retrieved from the institution’s radiology picture archiving and communication system. The disease-positive set included 80 radiographs each representing cardiomegaly, pneumothorax, pleural effusion, and pneumonia. All patients were aged 20 years or older.
Investigators evaluated OpenAI’s gpt-image-2 and Google’s gemini-3-pro-image-preview, referred to in the study as the GPT-image and Gemini-image models. Patient age, sex, and imaging findings extracted from original radiology reports were converted into standardized prompts. Each model generated disease-positive frontal chest radiographs using either the text prompt alone or the same prompt combined with an age- and sex-matched normal chest radiograph. Four board-certified radiologists with four to 20 years of experience independently reviewed randomized real and synthetic images and classified each as real or AI-generated. Readers were blinded to image provenance, the proportion of real and synthetic images, and the model and synthesis method used.
The lower detection rate with image conditioning was observed with both models. For GPT-image, radiologists identified 30% of image-conditioned radiographs as artificial intelligence (AI)-generated compared with 49% of text-only images, an absolute reduction of 19 percentage points. For Gemini-image, detection decreased from 64% with text-only generation to 38% with image conditioning, an absolute reduction of 26 percentage points. The same pattern was observed across all four disease categories.
The models also differed in detectability. Across generation methods, radiologists identified Gemini-image radiographs as synthetic 51% of the time compared with 39% for GPT-image radiographs. With text-only generation, detection was 64% for Gemini-image vs 49% for GPT-image; with image-conditioned generation, detection was 38% vs 30%, respectively.
Among disease categories, pneumothorax had the highest pooled AI detection rate at 55%, compared with 45% for pleural effusion, 41% for cardiomegaly, and 39% for pneumonia. The researchers suggested that the higher detection rate for pneumothorax could reflect difficulty generating convincing pleural lines and adjacent lung texture. However, the study did not prospectively record the visual cues underlying each reader judgment, and the researchers characterized these observations as qualitative rather than a formal image-level error analysis.
Quantitative image analyses complemented the reader findings. Among image-conditioned radiographs, GPT-image generated images with greater structural and perceptual similarity to the corresponding normal conditioning radiographs than Gemini-image. GPT-image also produced lower Fréchet Inception Distance values with both generation methods, indicating closer distributional similarity to the real disease-positive radiograph sets.
The researchers emphasized that these measures assessed image-level similarity rather than clinical validity. The metrics did not independently validate diagnostic correctness, disease severity, anatomical extent, or pathological realism.
Researchers also conducted a paired assessment of image-conditioned synthetic radiographs that each reader had initially failed to identify. When these images were subsequently displayed alongside the corresponding real disease-positive radiograph, radiologists correctly identified the synthetic image in 79% of GPT-image comparisons and 84% of Gemini-image comparisons. Detection varied among readers, ranging from 61% to 98% for GPT-image and 69% to 97% for Gemini-image.
The researchers characterized the paired assessment as an evaluation of artifact awareness rather than clinical diagnostic accuracy. They also noted that the setting differed from routine health care workflows, in which a matched real radiograph is typically unavailable for comparison.
Several factors limited the generalizability of the findings. The study was conducted at a single center with limited case and imaging-equipment diversity and evaluated only four common chest radiographic abnormalities. Images were exported, cropped, and standardized to a uniform aspect ratio rather than viewed in native clinical Digital Imaging and Communications in Medicine format.
Synthetic radiographs accounted for 80% of images in the reader study, which did not reflect routine clinical practice and may have affected reader vigilance and decision thresholds. The researchers therefore cautioned that the results represented relative detectability under controlled, enriched conditions rather than real-world screening performance.
The study also did not independently validate diagnostic correctness, consistency of disease severity, anatomical extent, downstream educational value, or real-world misuse. The radiologist reader study was focused on provenance and should not be interpreted as a de novo diagnostic task. Because commercial generative models evolve rapidly, the findings also reflected only the model versions, application programming interface settings, prompt structure, preprocessing steps, and generation workflow tested in the study.
Overall, image-conditioned generation was associated with lower detection of synthetic disease-positive chest radiographs during isolated review, but the findings did not establish that the generated images were diagnostically accurate or clinically interchangeable with real radiographs. Side-by-side comparison improved detection, suggesting that synthetic-image artifacts may become more apparent when a real comparator is available.
“These findings highlight the need for transparent labeling, provenance tracking, expert review, and controlled use of AI-generated medical images in healthcare, education, and research,” wrote lead study author Jinghang Wang, of the Department of Radiology at The Second Xiangya Hospital of Central South University, and colleagues.
Disclosures: The authors reported no conflicts of interest.
Source: Frontiers in Medicine
