In September 2024, the journal PNAS Nexus published a study with a slightly awkward message for scientists. David Markowitz, a communication researcher at Michigan State University, compared the plain-language summaries that researchers write for their own papers with summaries written by GPT-4 [1]. Readers found the AI versions clearer, rated the scientists behind them as more credible and trustworthy in the first experiment, and wrote better summaries of the science afterwards in the second [1].
I think the study is good news, with one catch. AI can be a real help in explaining research to the public. But a summary that reads well and makes the author look good is not automatically a summary that gets the science right. Checking that part is still the scientist’s job.
Scientists are not great at writing simply
The journal PNAS asks authors for two summaries: a technical abstract and a short significance statement meant for a wider audience. Markowitz analyzed 34,584 papers that had both [1]. The lay versions were indeed simpler, but only by a small margin, small enough that he says it is unclear whether individual readers would notice [1].
This fits something most of us already suspect. As the paper puts it, “it may be difficult for experts to write for nonexperts” [1]. After years inside a field, we stop hearing our own jargon.
That jargon has a cost. In a 2020 experiment with 650 people, Shulman and colleagues found that jargon disrupted people’s ability to process scientific information smoothly, even when the jargon terms came with definitions [2]. It also affected how much readers felt part of the science community, and through that, their interest and sense of understanding [2]. A lay summary full of technical words is not neutral. It quietly tells readers the science is not for them.
What happened when GPT-4 wrote the summary
Next, Markowitz gave 800 of those abstracts to GPT-4, along with the same instructions PNAS gives its authors [1]. The AI summaries used more common words and were more readable than the ones the scientists had written [1].
Then came the experiments. In the first, 274 people in the United States each read three summaries, randomly shown either the scientists’ version or the GPT-4 version, and rated the writer [1]. They judged the writer of the simpler AI version as more credible and more trustworthy, but also as less intelligent [1]. And in a twist I enjoyed, readers thought the simpler AI texts were more likely to have been written by a human [1].
In the second experiment, 250 new people read the summaries and then answered a question and described the study in their own words [1]. Those who read the AI versions scored higher on comprehension overall, and the clearest difference was in their written descriptions, which were rated as more accurate [1].
Reading the fine print
These are real results from preregistered experiments, and I take them seriously. But a few details shape what they can tell us.
The summaries used in the experiments were not typical ones. They were the pairs with the biggest difference in everyday words between the AI and human versions [1]. That makes sense for a first test, but it shows the effect at its strongest.
The trust results also weakened in the second experiment. With four times as many texts, the credibility effect was only marginal and the trust effect was not statistically significant [1]. The author suggests the effect may depend on the field of science [1].
Most importantly, understanding here meant understanding the summary. The multiple-choice question was written by another AI model, Gemini, and the readers’ written descriptions were scored by two AI models, GPT-4o and GPT-4 itself [1]. The multiple-choice scores alone did not differ significantly between groups [1]. None of this tests whether the AI summary faithfully captured the paper.
Imagine a researcher whose study found a small, early signal that a treatment might help a few patients. An AI summary turns that into a crisp sentence saying the treatment works. Readers understand it easily, rate the author as trustworthy, and walk away confidently believing something the study never showed.
That scenario is exactly the gap the experiments were not designed to close. Markowitz acknowledges it too. He warns that “One unintended consequence of simplifying science could be the loss of nuance or depth in the public’s understanding of complex issues” and suggests that future work have experts rate AI summaries to make sure the critical parts of the work are covered [1].

Clear is not the same as correct
There is earlier evidence on what can go wrong. In 2023, Tang and colleagues asked GPT-3.5 and ChatGPT to summarize the abstracts of Cochrane reviews, the careful syntheses of medical evidence, across six clinical areas [3]. They found the models could produce summaries that were factually inconsistent with the source and statements that were “overly convincing or uncertain” [3].
One example stays with me. The authors describe how a summary might claim atypical antipsychotics are effective for psychosis in dementia, while the review says the effect is negligible [3]. The authors also describe a certainty illusion, where a summary sounds more or less sure than the evidence it came from [3].
To be fair, the best setup they tested had factual errors in fewer than 10% of summaries [3]. That is not bad for a machine. But as the authors write, “Medical evidence summaries should be perfectly accurate” [3]. They also found that automatic scores did not reliably catch these problems, and concluded that human evaluation was still needed [3].
Put these studies side by side and the risk becomes clear. Fluent, confident writing makes readers trust the author more. If that fluent writing has quietly changed the finding, the trust goes to the wrong message.
How I would use AI for a lay summary
I write this blog to explain science to people outside the lab, and in my research I build large language model tools that pull structured information out of medical records. Both jobs have taught me the same thing: these models are very good at sounding right, and that is precisely why their output needs checking.
So here is how I would use AI for a lay summary. Give it your abstract, or better, your key results, and ask for a plain-language version. Then read it line by line against your paper. Check that every claim is one you made, that the size of the effect has not grown, and that words like “may” and “suggests” survived. If the study was in mice, make sure the summary still says mice.
AI can write a clearer summary of my research than I usually do on the first try, and I think researchers should use it. But my name goes on that summary, so checking that it still says what the science says is my job, not the model’s.
Markowitz’s study shows that clarity helps readers trust scientists and follow their work. The best reason to use AI for your lay summary is to earn that trust. The best reason to check it carefully is to deserve it.
References
[1] D. M. Markowitz, “From complexity to clarity: How AI enhances perceptions of scientists and the public’s understanding of science,” PNAS Nexus, vol. 3, no. 9, Art. no. pgae387, Sep. 2024, doi: 10.1093/pnasnexus/pgae387.
[2] H. C. Shulman, G. N. Dixon, O. M. Bullock, and D. Colón Amill, “The effects of jargon on processing fluency, self-perceptions, and scientific engagement,” Journal of Language and Social Psychology, vol. 39, no. 5-6, pp. 579-597, Oct. 2020, doi: 10.1177/0261927X20902177.
[3] L. Tang, Z. Sun, B. Idnay, J. G. Nestor, A. Soroush, P. A. Elias, et al., “Evaluating large language models on medical evidence summarization,” npj Digital Medicine, vol. 6, no. 1, Art. no. 158, Aug. 2023, doi: 10.1038/s41746-023-00896-7.
Related Articles




