ChatGPT-5 may provide useful introductory information about acne, but dermatologist ratings suggested its answers were less consistently appropriate when questions involved diagnosis or treatment.
Investigators conducted a cross-sectional observational study using 25 questions commonly asked by patients with acne in outpatient practice. Questions were divided into 4 categories: skin care routine, nutrition and lifestyle, diagnosis and general information, and treatment and medications. Each question was submitted to the publicly available ChatGPT-5 web interface in 3 independent sessions, generating 75 unedited responses. Three dermatologists independently and blindly classified each response as appropriate or inappropriate, with the majority decision determining the final classification.
The investigators assessed the overall proportion of responses rated appropriate, appropriateness by category, agreement among the 3 dermatologists, reasons for inappropriate ratings, response-order effects, and question-level success, defined as at least 2 of 3 responses to a question being rated appropriate. Overall, 65% (n = 49) of the responses were classified as appropriate. Appropriateness was highest for skin care at 83% followed by nutrition and lifestyle at 78%, treatment and medications at 50%, and diagnosis and general information at 33%.
Among the 26 responses classified as inappropriate, misleading information was the most frequently cited reason, accounting for 10 responses. Five responses were considered too technical or difficult for patients to understand, and 3 were considered misleading because risks were not specified. Less frequent concerns included unnecessary or off-topic detail, incomplete or contradictory information, and nonscientific recommendations or data.
At the question level, 68% (n = 17/25) questions met the investigators' definition of success. The pattern was consistent with the primary findings, with higher success for skin care and nutrition questions and lower success for diagnosis and treatment questions. Response order was not associated with differences in appropriateness.
Agreement among the dermatologists was low. The binary classification could force clinically ambiguous responses into either the appropriate or inappropriate category. Questions near the diagnosis-treatment boundary could produce differences in clinical interpretation. The absence of a calibration session may also have led panel members to apply different thresholds for safety and adequacy. The investigators cautioned that the negative agreement value did not mean ChatGPT-5 was consistently incorrect but instead reflected differing professional thresholds within the binary assessment framework.
The study was limited to acne, potentially restricting generalizability to other dermatologic conditions, and its binary assessment did not quantify intermediate degrees of appropriateness. The responses also came from the ChatGPT-5 web interface during a defined period in August 2025, and subsequent model updates could change performance. The low inter-rater agreement meant the results should be interpreted according to the majority decision of the dermatologist panel.
The investigators cautioned against using chat-based interfaces for diagnosis or treatment guidance without appropriate risk disclosures, current guideline references, and specialist oversight.
“The final content must always be approved by a dermatologist,” wrote lead study author Nazli Caf, MD, of Istanbul Atlas University, and colleagues.
The study authors reported no funding or conflicts of interest.
Source: Medicine
