
In September 2023, two researchers published a simple and useful test in Scientific Reports. William Walters and Esther Isabelle Wilder asked two versions of ChatGPT, GPT-3.5 and GPT-4, to write short literature reviews on 42 topics, which gave them 84 papers and 636 references to check by hand [1]. The texts were generated in the first week of April 2023 [1].
Within that set, 55% of the GPT-3.5 references and 18% of the GPT-4 references pointed to works that do not exist [1]. Every one of them was set out in standard APA citation style.
My view is that this study says less about one chatbot than about a habit many of us have. We treat a well-formatted reference as a sign that someone has done the checking. It never was one. A plausible reference is exactly the kind of text a language model is built to produce, so the job of confirming that a source exists and says what we claim still belongs to whoever puts their name on the work.
What the study actually found
For each reference, Walters and Wilder searched Google Scholar, PubMed, Scopus, library catalogues and publisher websites, and they counted a work as real if they found a match or near match for both its title and its authors [1]. So a slightly garbled citation to a real paper counted as real, with an error, rather than as fake.
Those errors were common too. Among the references that did point to real works, 43% of GPT-3.5’s and 24% of GPT-4’s had a substantive mistake, such as a wrong year, wrong volume or wrong page numbers [1]. GPT-4 was clearly better. The authors call GPT-4 a major improvement, and add in the same breath that problems remain [1].
Two details stayed with me. Even with GPT-4, 70% of the cited book chapters were fabricated [1]. And most of the fake article, book and website references used the names of real journals, publishers and organizations [1]. An invented article in a journal you have heard of is much harder to catch than one in a journal that does not exist.
It matters which tools these numbers describe. They come from GPT-3.5 and GPT-4 as they were in spring 2023, on broad first-year college essay topics rather than specialized science [1]. By August 2023, GPT-3.5 was the free version and GPT-4 was for paid subscribers [1]. I would not quote these percentages as the error rate of any chatbot you can use today. The lesson I take from them is about how the mistakes looked.
Polished formatting is not evidence
Every reference in the study was in APA style, the format many students are taught [1]. More than 40% had small formatting slips, mostly capital letters in the wrong places, and the fake references showed the same kinds of slips as the real ones [1].
Links did not help either. Few references came with a web link at all, and links were more likely to turn up in the fabricated references than in the real ones [1]. A link that looks official is still just text until you click it.
Imagine a student handing in an essay with eight references, each with authors, a year, a journal, a volume and page numbers. The grader skims the list, sees familiar journal names and moves on. Two of those papers were never written, and nothing on the page gives that away.
This is why I think we need to drop the idea that a complete-looking citation is a checked one. Formatting is the easy part for a language model. The authors of the study put it plainly: ChatGPT is a language-processing tool, not an information-processing one [1]. It produces text that looks like a reference because it has seen so many of them. Whether the paper exists is a separate question the formatting cannot answer.

When the fake cases reached a courtroom
The most famous example came from law, not science. In June 2023, a federal judge in New York sanctioned two lawyers and their firm in Mata v. Avianca for filing court decisions that did not exist. The order says they “submitted non-existent judicial opinions with fake quotes and citations created by the artificial intelligence tool ChatGPT,” and then kept standing by them after the court questioned them.
The lawyers later admitted that six cited decisions had been generated by ChatGPT and did not exist. When one of them asked ChatGPT whether the cases were real, it told him they were and could be found in the standard legal databases. He had even tried to look up one of the cases, could not find it, and cited it anyway. At the hearing he said: “I just never thought it could be made up.” The court imposed a $5,000 penalty.
What I find most useful is how the judge framed it. The order says there is “nothing inherently improper about using a reliable artificial intelligence tool for assistance,” but that lawyers keep a gatekeeping role in making sure their filings are accurate. That is the right line for science too.
Checking stays with the author
Nature drew that same line early. In January 2023, the journal said that, at Nature and all Springer Nature journals, no large language model would be accepted as an author on a research paper, because authorship “carries with it accountability for the work, and AI tools cannot take such responsibility” [2]. Walters and Wilder end in the same place: users of ChatGPT “are cautioned to check the citations it generates” [1].
The best argument against worrying too much is that the tools keep improving. The drop from 55% to 18% in one model generation is real, and the authors say so [1]. I accept that. But 18% of references being invented is still far too many for anything you sign, and even the real references often had wrong details [1]. Better tools lower the error rate. They do not move the responsibility.
I build large language model tools that pull structured clinical information out of medical records, so I see this from the builder’s side. In that kind of work, the output only becomes useful once someone has checked it against the record it came from, and a reference list deserves the same treatment.
How to check a reference in a minute
You do not need special access to do this. For each reference you plan to rely on:
- If it has a DOI (the code that starts with 10.), paste it after https://doi.org/ and see whether it opens the paper it claims to be.
- Paste the exact title, in quotation marks, into Google Scholar or PubMed and look for a match with the same authors.
- Open the paper and confirm it says what the sentence citing it says.
If none of these finds the paper, treat it as missing until you can prove otherwise. The last step, reading the source, is the one I suspect people skip most.
A reference that looks right is only a claim that a paper exists. Whoever signs the work has to check that claim, whether a person or a chatbot wrote the list.
None of this means students and researchers should avoid AI tools. It means a reference list is a set of promises, and a chatbot can write promises it cannot keep. A minute per reference is a small price for making sure every paper you cite is one someone actually wrote.
References
[1] W. H. Walters and E. I. Wilder, “Fabrication and errors in the bibliographic citations generated by ChatGPT,” Scientific Reports, vol. 13, no. 1, Art. no. 14045, Sep. 2023, doi: 10.1038/s41598-023-41032-5.
[2] Nature, “Tools such as ChatGPT threaten transparent science; here are our ground rules for their use,” Nature, vol. 613, no. 7945, p. 612, Jan. 2023, doi: 10.1038/d41586-023-00191-1.
Related Articles



