A previously trained artificial intelligence model designed to estimate gestational age from blind sweep ultrasonography may have demonstrated accuracy in two new clinical settings, achieving performance comparable to the clinical standard.
Researchers conducted a prospective multicenter diagnostic study to evaluate whether the artificial intelligence (AI) model could generalize to clinical environments beyond those in which it had been trained. They enrolled 2,043 patients at urban medical centers in Chicago and Nairobi, Kenya, 385 of whom with pregnancies between 16 and 36 weeks were included in the primary analysis. Five novice operators acquired blind sweep ultrasonography examinations using Clarius probes, and the model’s gestational age estimates were compared with those obtained using standard clinical ultrasonography performed by expert sonographers.
The primary outcome was the mean absolute error (MAE) of gestational age estimation compared with the clinical standard. Secondary analyses evaluated mean error and assessed model performance by study site, timing of the reference dating scan, gestational age beyond 36 weeks, and operator.
Across both sites, the adapted AI model achieved an MAE of 4.2 days vs. 4.5 days for the clinical standard, meeting the study’s prespecified noninferiority criterion. Performance was consistent across both clinical settings, with MAEs of 4.1 and 4.3 days in Chicago and Nairobi, respectively. The AI model did not demonstrate systematic bias in gestational age estimation.
Performance remained stable across different reference dating windows. Among patients in Nairobi whose dating scans were performed prior to 14 weeks’ gestation, both the AI model and the clinical standard achieved an MAE of 4.2 days. Among those whose dating scans occurred between 14 and 22 weeks, both approaches achieved an MAE of 4.4 days, indicating similar performance despite differences in scan timing.
Exploratory analyses suggested comparable performance between the AI model and clinical standard. The researchers reported improved accuracy among fetuses above the 90th weight percentile compared with standard fetal biometry formulas. Performance remained generally consistent across novice operators, although the lower sweep rejection rate in Nairobi suggested that structured hands-on instruction may improve image acquisition.
Performance was more variable beyond 36 weeks’ gestation. In Nairobi, the clinical standard outperformed the AI model in late-term pregnancies, whereas the Chicago analysis included relatively few patients. The researchers noted that differences in clinical workflow and the lack of blinding to reference gestational age at one site may have contributed to the findings, while also suggesting that additional model refinement may be needed for later gestation.
The researchers acknowledged several limitations. The model was adapted using data collected only in Chicago prior to external validation, and generalizability was evaluated using a single new ultrasound manufacturer. Differences in operator training between sites also may have influenced image acquisition, and subgroup analyses involving later gestation, multiple gestations, fetal weight extremes, and individual operators were based on relatively small sample sizes.
“In this diagnostic study of AI for [gestational age] estimation using blind sweep ultrasonography, we effectively demonstrated generalization to new clinical settings with new novice operators and different probes, and we achieved noninferior performance to the clinical standard, thereby highlighting its potential for expanding access to essential prenatal care,” wrote lead study author Angelica Willis, MS, of Google, and colleagues.
Full disclosures of the study authors can be found in the study.
Source: JAMA Network Open
