In March 2024, Weixin Liang, James Zou and colleagues posted a study that tried to measure something most people in science only whispered about: how much peer review text was now being written with help from ChatGPT [1]. They looked at reviews for four large artificial intelligence (AI) conferences held after ChatGPT came out: ICLR 2024, NeurIPS 2023, CoRL 2023 and EMNLP 2023 [1]. Their estimate was that between 6.5% and 16.9% of the review text could have been substantially modified by a large language model (LLM), meaning more than spell-checking or small edits [1]. The paper was later presented at the ICML 2024 machine learning conference [1].

The study cannot say who did the thinking in any single review, and I will be careful about what it does show. My view is simple. Using AI to polish your wording is fine. Using it to form your verdict is not. And pasting someone’s unpublished manuscript into a chatbot is a separate problem.

  • An estimated 6.5% to 16.9% of review text at four AI conferences was substantially modified by an LLM [1].
  • The estimate covers a whole body of reviews; it does not identify any reviewer or prove anyone outsourced a whole review [1].
  • AI-shaped text was more common near deadlines, in low-confidence reviews, in reviews citing no one (no et al.), and among reviewers who skipped author rebuttals [1].
  • Journal reviews in the Nature Portfolio showed no significant rise [1].
  • NIH banned generative AI in its grant peer review in June 2023, citing confidentiality (NIH notice).

What the study actually measured

The method does not read a review and decide whether a machine wrote it. It estimates what share of a whole collection of text looks like AI output, without making a call on any individual document [1]. Think of it like estimating how much of a city’s water comes from one river by sampling the reservoir, not by testing each tap.

Before ChatGPT, the estimates were low: 1.6% at ICLR, 1.9% at NeurIPS and 2.4% at CoRL [1]. Afterwards they rose to 10.6%, 9.1% and 6.5% [1]. EMNLP, a conference on language processing, had no earlier data but the highest figure, about 16.9% [1]. Reviews for Nature Portfolio journals showed no significant increase [1].

When the team ran reviews through ChatGPT only to fix typos and grammar, the estimate barely moved, so the method is picking up more than proofreading [1]. And the authors are explicit that they “do not claim (nor do we believe) that many reviewers are using ChatGPT to write entire reviews outright” [1]. A reviewer could jot down bullet points and ask the tool to turn them into paragraphs, and that would still count [1].

So a real, measurable share of review text was reshaped by AI. But it is a corpus-level estimate: it does not tell you which reviewers did it, or show that anyone handed over the whole job.

Where the AI text showed up

The revealing part is where the AI-shaped text appeared. The estimated share was higher in reviews submitted within three days of the deadline; at ICLR 2024 it was 11.3% close to the deadline against 8.8% earlier [1]. It was higher in reviews where reviewers rated their own confidence as 2 or lower on a 5-point scale [1].

It was lower in reviews that contained et al., which the authors used as a rough sign that the reviewer cited other research [1]. At ICLR and NeurIPS, it was also higher among reviewers who did not respond to the authors’ rebuttals [1]. Even the vocabulary shifted. In ICLR 2024 reviews, commendable became 9.8 times more likely to appear in a sentence, and meticulous 34.7 times [1].

All of these are correlations. The authors say they “cannot make a causal claim” about the discussion effect, and they suggest one kinder explanation: because conferences are short of reviewers, scholars may agree to take on more reviews and lean on the tool to cope [1]. For citations, they note they lack a counterfactual [1]. People who reach for ChatGPT might have cited less even without it.

Imagine a reviewer with six papers due by Friday. On Thursday night, for the one they skimmed, they paste in a few rough notes and ask a chatbot to turn them into a proper review. The result is fluent, polite and generic, and nobody reading it can tell how little thinking sits underneath.

That is the pattern that worries me. The AI text clusters where the signs of human attention were thinnest.

Brown folder labeled confidential lying on a dark desk
A manuscript under review is confidential. Pasting it into a chatbot changes who can see it.

Polishing words versus handing off judgment

I build large language model tools that pull structured information out of medical records, so I know how good these systems are at turning rough input into clean prose. That is exactly why I think the line has to be drawn at judgment, not at the keyboard.

A peer review has two parts. One is the verdict: is the method sound, does the evidence support the claim, what is missing? The other is the writing that carries that verdict to the authors and editors. Asking a tool to tidy the second part is like asking a colleague to check your grammar. Asking it to supply the first part means the paper was never really reviewed by a peer.

The study hints at what is lost when that line blurs. Reviews most similar to the other reviews of the same paper had more AI-shaped text, and the authors warn that authors then “lose an opportunity to receive feedback from multiple, independent, diverse experts in their field” [1]. Peer review works because different people notice different things. If every reviewer runs their notes through the same model, the reviews start to sound alike, and some of that independence goes with them.

To be fair, the authors themselves decline to call AI use in reviewing good or bad [1]. For a reviewer writing in a second language, a tool that helps them say what they mean is a real benefit. If the thinking is theirs and the tool only shapes the sentences, I have no complaint.

The manuscript is not yours to share

There is a second issue the study flags but leaves aside: the “privacy and anonymity risks of providing unpublished work to a privately owned language model” [1]. It happens the moment you paste, whatever the tool writes back.

The US National Institutes of Health (NIH) acted on this early. In a June 2023 notice, it prohibited its peer reviewers from using large language models or other generative AI to analyze and write critiques of grant applications. Its reasoning was about confidentiality: uploading content from an application to online AI tools breaks its review rules, because “AI tools have no guarantee of where data are being sent, saved, viewed, or used in the future.”

NIH’s rule covers grant review, not journal or conference manuscripts, but the logic travels. A manuscript under review is someone else’s unpublished work, shared with you in trust. It is not your material to hand to a third party, however useful that party may be.

Let AI help you say what you think, never decide what you think. And keep other people’s unpublished work out of any tool you would not trust with your own.

For reviewers, that means a few plain habits: read the paper yourself, write your own points first, check any citation you add, and do not paste confidential text into a public chatbot. For editors and conference organizers, it means clear rules on disclosure and on what may be uploaded, plus review loads that leave time to think. The study cannot tell us who did the thinking in any given review. It does tell us that the machine’s share was higher in late, low-confidence reviews, and that is the part we can actually fix.

References

[1] W. Liang, Z. Izzo, Y. Zhang, H. Lepp, H. Cao, X. Zhao, et al., “Monitoring AI-modified content at scale: A case study on the impact of ChatGPT on AI conference peer reviews,” arXiv:2403.07183, Mar. 2024, doi: 10.48550/arXiv.2403.07183.

Saleh Ramezani

Saleh Ramezani is the founder of Better Science. Saleh believes that science literacy is crucial for navigating today’s science-driven world. Saleh is currently a post-doctoral researcher at MD Anderson Cancer Center in Houston, Texas.

Get involved

Have something to say about science? Write with us.