An artificial intelligence model may have comparable accuracy to experienced radiologists in estimating pleural effusion volume on chest radiography, although both approaches showed only moderate agreement with computed tomography and substantial estimation errors.
Investigators conducted a retrospective observational study of 88 adult patients who underwent chest radiography and computed tomography (CT) on the same day at a single tertiary institution in 2025. Patients were included if CT confirmed a pleural effusion of at least 20 mL in 1 hemithorax. The analysis encompassed 176 hemithoraces, and the median CT-derived volume among pleural spaces with effusions was 389 mL. Two board-certified radiologists with 19 and 9 years of thoracic imaging experience independently estimated pleural fluid volume on radiographs while blinded to the CT measurements and each other's assessments. The investigators used a commercial general-purpose multimodal artificial intelligence (AI) model, ChatGPT 5.2, analyzed the same radiographs using a standardized zero-shot prompt. Manually segmented CT volumes served as the reference standard.
The investigators evaluated agreement between CT-derived pleural effusion volumes and estimates generated from chest radiographs by the radiologists and AI as well as repeatability, factors associated with estimation error, and the ability of each approach to identify clinically significant pleural effusions.
Agreement with CT was moderate across all 3 approaches. Mean estimation errors were 300 mL, 250 mL, and 264 mL, respectively. All 3 approaches showed positive bias, with mean overestimation of 103 mL for radiologist 1, 146 mL for radiologist 2, and 57 mL for AI. The wide limits of agreement indicated substantial variability in individual volume estimates.
For identifying pleural effusions of at least 300 mL, the area under the curve (AUC) was 0.84 for radiologist 1, 0.90 for radiologist 2, and 0.85 for AI. Using an AI-estimated volume threshold of 380 mL, sensitivity and specificity were both 78%. Radiologist 2 had a higher AUC compared with AI, whereas performance did not differ statistically between radiologist 1 and AI.
Estimation error varied according to several imaging and patient characteristics. For AI, CT-derived effusion volume, left-sided effusions, and patient position were independently associated with error. Upright radiographs were associated with less AI estimation error compared with supine radiographs. Fluid in supine patients may appear as diffuse opacity rather than discrete layering, while the cardiac silhouette may complicate assessment of left-sided effusions. Confidence scores were not associated with estimation accuracy for either AI or the radiologists.
AI showed high within-session consistency, but between-session consistency was lower. Repeatability was 0.87 for radiologist 1 and 0.91 for radiologist 2, while agreement between the radiologists was 0.88.
The study had several limitations. It was retrospective and conducted at a single tertiary institution with a relatively small adult cohort. Selection enriched the cohort for positive cases and introduced spectrum bias, limiting generalizability, particularly to patients with very small or absent effusions. Radiography and CT were performed on the same day but not simultaneously, allowing pleural fluid volume to change between examinations. CT volumetry, rather than actual drained fluid volume, served as the reference standard. The investigators also evaluated just 1 general-purpose commercial AI model using a specific zero-shot prompting strategy, and the nonlocked model may change with software updates.
“The wide [limits of agreement] observed across all evaluators in this retrospective cohort confirms that chest radiography provides only a coarse approximation of pleural fluid volume, irrespective of the interpreter,” wrote lead study author Dominik Schiller, of the Department of Medical Imaging at the Motol and Homolka University Hospital and Second Faculty of Medicine at Charles University in the Czech Republic, and colleagues.
The study authors reported no conflicts of interest.
Source: European Radiology Experimental
