In July 2023, the journal Science published an experiment on AI and work. Shakked Noy and Whitney Zhang, two economics PhD students at MIT, gave 453 college-educated professionals writing tasks tied to their jobs and randomly gave half of them access to ChatGPT [1]. The people with the chatbot took about 40% less time, and their work was rated 18% higher in quality [1].

These were professional writing tasks, not scientific papers. Nobody in the study wrote a methods section or analyzed a dataset. Still, it is hard not to picture the same speed-up in research.

My view is that faster writing is a real gain, and it is still a different thing from better science. A paper finished in half the time can be just as wrong, or just as unimportant, as a slow one. If we keep measuring academic productivity by how many papers we produce, AI will make us look very productive while doing little for the things that matter. That extension to science is my opinion, not something the study tested.

  • In a 2023 randomized experiment, professionals using ChatGPT took about 40% less time on writing tasks, with quality rated 18% higher [1].
  • The tasks were short, job-style writing assignments, and nobody checked the outputs for factual accuracy.
  • A 2016 study argued that rewarding high output selects for poorer research methods [2].
  • Nature‘s 2023 ground rules said no LLM tool will be accepted as a credited author on a research paper, because authorship means accountability [3].
  • I think academic productivity should count what a paper gets right, not how quickly it appears.

What the experiment actually tested

The participants were marketers, grant writers, consultants, data analysts, human resource professionals and managers, according to MIT’s news story on the study. Each did two tasks of 20 to 30 minutes, such as a cover letter for a grant application or an email about restructuring an organization. Experienced people from the same occupations graded the work without knowing who had used ChatGPT.

The design is clean. The experiment was preregistered, meaning the plan was fixed before the data came in, and access to the chatbot was assigned at random [1]. It also found that inequality between workers shrank, and that people who tried ChatGPT in the study were twice as likely to report using it in their real jobs two weeks later [1].

So I take the result seriously. For this kind of writing, the chatbot helped.

Faster is not the same as right

The researchers were open about what their setup left out. Because participants were anonymous, the tasks could not require knowledge of a specific company or customer, and the instructions were more explicit than real work usually is, as MIT News reported. They also did not think it was feasible to hire fact-checkers, so the accuracy of the outputs was never evaluated.

That last point matters most for science. In a research paper, accuracy is the whole job. A grader reading a cover letter can judge whether it reads well. A reviewer reading a methods section often cannot tell whether the analysis was run correctly or whether the result will hold up. Noy himself said the speed benefits may be smaller in real work “because you need to spend time fact-checking and writing the prompts,” in the same MIT News story.

Imagine two students writing up the same lab experiment. One uses a chatbot and hands in a polished report by lunch; the other writes slowly and hands it in the next morning. Both reports read well, but only the slow one noticed that a sample was mislabeled, and a grade based on how the writing reads would not tell them apart.

Speed and polish are what the experiment measured, and in science those were never the hard part. The hard part is being right, and then being right about something that matters.

Tall stacks of journals and papers on a wooden desk
When careers reward volume, a tool that makes papers faster can mean more papers, not better ones.

Science already rewards volume

This worry did not start with AI. In 2016, Paul Smaldino and Richard McElreath argued that when publication is a principal factor for career advancement, methods that produce more papers tend to spread, with no deliberate cheating required [2]. In their computer model of competing labs, “selection for high output leads to poorer methods and increasingly high false discovery rates” [2]. They also looked back over 60 years of studies in the behavioural sciences and found that statistical power, a basic measure of whether a study is big enough to detect what it looks for, had not improved despite repeated warnings [2].

A tool that makes writing cheaper fits neatly into that system. If a lab is rewarded for paper count, I expect the time saved by a chatbot to go mostly into more papers, rather than into bigger samples or replications. That is my prediction, not a finding.

Journals saw part of this early. In January 2023, Nature noted that ChatGPT had already produced research abstracts good enough that scientists found it hard to spot a computer had written them [3]. The editorial’s big worry was that researchers could pass off chatbot text as their own, or use the tools in a simplistic way and “produce work that is unreliable” [3].

The case for faster writing

There is a fair argument on the other side. A lot of academic writing is not science at all: grant boilerplate, cover letters, progress reports, reformatting a manuscript for yet another journal. If a chatbot trims that, researchers get hours back for thinking, experiments and teaching. The finding that workers who scored lower on their first task benefited more from ChatGPT [1] could matter too.

I agree with most of this. I build large language model tools that extract structured clinical information from medical records, so I would be the last person to say these models have no value. Writing up research takes real time too, so a tool that speeds up the writing has obvious appeal to me.

Where I part ways is on what happens to the saved time. The tool does not decide that. The incentive system does, and right now it mostly counts papers.

What productivity should mean

Nature‘s rules point in a useful direction. Under those rules, no LLM tool will be accepted as a credited author on a research paper, because “any attribution of authorship carries with it accountability for the work, and AI tools cannot take such responsibility” [3]. Researchers who use these tools should say so in the methods or acknowledgements [3]. The principle underneath is that a paper is a promise: a person stands behind every claim in it.

The San Francisco Declaration on Research Assessment, drafted in December 2012, already asked institutions to make clear “that the scientific content of a paper is much more important than publication metrics or the identity of the journal in which it was published.” It also asked them to count datasets and software as research outputs. Those ideas carry more weight now that text is cheap.

What I would count instead is whether the work answered a question someone needed answered, whether another group can reproduce it, and whether it shared the data and code that let others build on it. None of these get faster because the prose does.

AI can help us write faster, and I am glad of it. But if we keep counting papers as the measure of a scientist, we will get more papers rather than better science, and the tools will take the blame for a problem we built ourselves.

For hiring and grant committees, that means reading a few papers closely instead of counting many. For researchers, including me, it means spending the time a chatbot saves on the parts no chatbot checks: the data, the analysis, and the question of whether the result matters. Productivity in science should mean how much reliable knowledge we add, and on that measure a faster first draft is only the start.

References

[1] S. Noy and W. Zhang, “Experimental evidence on the productivity effects of generative artificial intelligence,” Science, vol. 381, no. 6654, pp. 187-192, Jul. 2023, doi: 10.1126/science.adh2586.

[2] P. E. Smaldino and R. McElreath, “The natural selection of bad science,” Royal Society Open Science, vol. 3, no. 9, Art. no. 160384, Sep. 2016, doi: 10.1098/rsos.160384.

[3] “Tools such as ChatGPT threaten transparent science; here are our ground rules for their use,” Nature, vol. 613, no. 7945, p. 612, Jan. 2023, doi: 10.1038/d41586-023-00191-1.

Saleh Ramezani

Saleh Ramezani is the founder of Better Science. Saleh believes that science literacy is crucial for navigating today’s science-driven world. Saleh is currently a post-doctoral researcher at MD Anderson Cancer Center in Houston, Texas.

Get involved

Have something to say about science? Write with us.