Patients considering surgery for medically refractory chronic rhinosinusitis or nasal obstruction may have reduced decisional conflict following use of either ChatGPT or Google.
In a prospective, randomized controlled pilot study at a single academic rhinology clinic between June and December 2024, researchers enrolled English-speaking adult patients with medically refractory chronic rhinosinusitis and/or nasal obstruction who qualified for surgery, had received one-on-one counseling with a rhinologist, and were deciding whether to pursue surgery. The patients were randomized 1:1 to use ChatGPT-4 (n = 29) or Google with disabled generative artificial intelligence (AI) features (n = 28) and had 15 minutes to ask at least three treatment-related questions without standardized prompts.
The primary outcome was change in the 16-item Decisional Conflict Scale (DCS), which ranges from 0 to 100, with higher scores indicating greater decisional conflict. Secondary outcomes included changes in treatment-related knowledge, system usability scale (SUS) scores, and rhinologist-rated accuracy of ChatGPT responses.
Mean DCS scores decreased from 17.5 to 13.2 in the ChatGPT group and from 14.1 to 11.4 in the Google group. Relative to baseline, total DCS scores decreased 36% with ChatGPT and 20% with Google. However, the degree of improvement was not statistically significantly different between the two groups.
Within-group analyses showed improvements in the informed and values-clarity domains with both platforms. The uncertainty subscore improved by 7.2 points in the ChatGPT group, whereas the Google group did not have a statistically significant change in that domain. However, the between-group analysis did not establish an advantage for ChatGPT in reducing uncertainty.
Neither intervention changed patients’ treatment preferences or improved treatment-related knowledge. Usability was also similar between the platforms, with mean SUS scores of 67.3 for ChatGPT and 66.4 for Google. Fellowship-trained rhinologists independently and blindly evaluated 56 ChatGPT responses and assigned them a mean accuracy rating of 7.2 on a 10-point scale.
An exploratory subgroup analysis showed that patients who were unsure about treatment at baseline had greater decisional conflict compared with those who already preferred surgery, declining by a respective 4% vs. 31%. Among 13 patients who were unsure, mean DCS scores were 30.8 at baseline and 29.3 following the intervention. Among 36 patients who preferred surgery, scores decreased from 9.8 to 7.0. The researchers suggested that exposure to information without sufficient guidance may heighten confusion or reinforce ambivalence among patients who remain uncertain about treatment. Informational tools alone may also be insufficient in altering clinical decision-making without personalized counseling and identifying patients with high baseline uncertainty as a potential population for more structured physician guidance or hybrid decision-making approaches.
The single-center pilot design and small sample size limited generalizability and the ability to detect smaller differences between ChatGPT and Google. Several measures of how patients used the platforms were not captured, including the content of queries, number of links accessed, and how patients engaged with the information.
Patients completed the intervention and follow-up assessments during a single clinic encounter following one-on-one counseling with a rhinologist. The researchers noted that prior counseling likely reduced baseline decisional conflict and may have limited the ability to detect between-group differences. Conducting the intervention in the medical office also limited the time patients had to process information, reflect on their preferences, or discuss treatment options with others. The researchers did not assess patients’ perceptions of information clarity or trustworthiness beyond usability scores.
The researchers suggested that ChatGPT may support shared decision-making among patients considering rhinologic surgery, but its reliability in clinical practice relative to standalone physician counseling or other decision-support tools remains uncertain.
“Further research across diverse patient populations is warranted,” wrote lead study author Omer Baker, of the University of California, San Diego School of Medicine, and colleagues.
The study authors reported no conflicts of interest.
