Misleading information in an artificial intelligence–generated summary may alter subsequent recall of an event.
Investigators combined an analysis of artificial intelligence (AI)-generated video summaries with an online experiment involving US adults. For the summary analysis, ChatGPT-5.5 and Gemini 2.5 Flash-Lite each summarized 2 25-second animated videos depicting a car-pedestrian accident 5 times per video, yielding 20 summaries. Among 360 participants, 328 completed both sessions and passed attention checks that were included in the final analysis. They watched 1 of the 2 videos and, 24 to 48 hours later, read a 21-sentence AI-generated summary. They were randomly assigned to a summary that either accurately described the traffic sign shown in the video or misleadingly substituted the other sign and were told that the summary had been produced either by AI or a professional human transcriber.
The primary experimental outcome was whether participants correctly identified the stop or yield sign they had originally seen. The investigators assessed whether the purported source of the summary influenced recall and whether participants’ trust in or use of AI modified the misinformation effect. A separate analysis evaluated omissions, inaccuracies, and erroneous additions in the AI summaries.
Correct traffic-sign recall was 84% among participants who received a consistent summary compared with 45% among those who received a misleading summary. The misinformation effect was observed for both original videos, regardless of whether participants initially saw a stop sign or a yield sign.
The source label did not substantially change the effect. Among participants who received misleading summaries, 46% of those told that the summary was AI-generated correctly recalled the traffic sign vs. 43% of those told that it was human written. The misinformation effect remained statistically significant within both source-label groups. Participants’ reported trust in AI and frequency of AI chatbot use did not significantly modify the association between misleading information and recall.
Performance on other questions suggested that participants retained information from the original videos rather than relying solely on the summaries. Ninety-seven percent correctly remembered that the accident occurred during the day and participants answered an average of 91% of nonprimary memory questions correctly.
Errors were present in all 20 AI-generated summaries, with each containing 7 to 21 errors. On average, 52% of predefined central details were omitted, 95% omitted the collision between the car and pedestrian, 65% contained at least 1 inaccurate description of a central detail, and 60% included at least 1 erroneous addition.
The AI-output analysis involved just 2 short animated videos, 2 AI models, 1 standardized prompt, and 20 summaries. The test assessed 1 type of misleading information, limiting generalizability to other errors, media, and more naturalistic or high-stakes settings. All participants received ChatGPT-generated text, regardless of whether the summary was labeled as AI- or human-generated; therefore, the source-label comparison reflected participants’ beliefs about authorship rather than a comparison of independently produced AI and human summaries.
“Taken together, we found that errors in AI-generated summaries can distort human memory,” wrote lead study author Mattea Sim, of Georgetown University, and colleagues.
The study authors did not provide a conflict-of-interest statement.
Source: Proceedings of the Ninth AAAI/ACM Conference on AI, Ethics, and Society
