In March 2026, Nature published a paper from a team at Sakana AI, the University of Oxford, the University of British Columbia and the Vector Institute describing The AI Scientist, a system that comes up with research ideas, writes code, runs experiments, analyses the data, writes the whole manuscript and then reviews its own work [1]. The headline result: one of its papers scored well enough with human reviewers to clear the bar at a workshop held alongside a major machine learning conference [1].
The wording matters, so let me be precise. Nature did not publish an AI-written paper. It published a human-written paper about an AI system. The AI-written paper went to a workshop, scored above the acceptance threshold, and was then withdrawn, exactly as planned [1].
I think this is a genuine milestone. I also think it tells us at least as much about the limits of a single round of workshop review as it does about AI. Both can be true at once, and the details are what separate them.
What the AI did, and what people did
The system works on machine learning research, because those experiments run entirely on a computer [1]. For this test, the team gave it the broad theme of one workshop at the International Conference on Learning Representations (ICLR) 2025, called I Can’t Believe It’s Not Better (ICBINB), which is about where deep learning ideas fall short [1]. From there, the authors say, the whole workflow, from ideas and code to the finished manuscript, ran without any human modification [1].
People were still involved, though. The authors write that “we manually filtered the most promising outputs at each stage” [1]. They chose three papers based on whether the idea fit the workshop, whether the code ran and did what the paper said, and whether the formatting was right [1]. In Sakana AI’s March 2025 announcement, the team put it plainly: they gave the system a broad topic and “picked 3 AI-generated papers to submit.”
The filtering is disclosed, and the authors argue that without it the same papers would still exist, just alongside more weak ones and at greater cost [1]. Fair enough. Still, a person choosing the best few out of many attempts is doing some of the judging a scientist does before pressing submit.
Imagine a student who has an AI tool write twenty essays overnight, reads through them, and hands in the best three. If one of them earns a good grade, the tool did the writing. The student’s eye for which three to hand in did part of the work too.
The Nature paper reports all of this openly. The risk is in how the story gets retold as a one-line headline.
What passing actually meant
Three AI papers went into a pool of 43 submissions, and reviewers were told that some papers were AI-generated, but not which ones [1]. One paper, reporting a negative result about a way to help neural networks generalize (Sakana AI), received scores of 6, 7 and 6, an average of 6.33 [1]. The other two did not reach the bar [1].
That paper never received a final acceptance decision. The team had decided in advance to withdraw every AI submission after review, whatever the outcome, to avoid setting a precedent before the research community agrees on standards for disclosure and evaluation [1]. The organizers said it would in all likelihood have been accepted [1]. The announcement adds a detail worth knowing: because of the planned withdrawal, the organizers did no meta-review, the final step where someone weighs the reviews, so in theory the paper could still have been rejected.
The venue matters too. The workshop accepted about 70% of submissions, while the main ICLR 2025 conference accepted 32% [1]. The authors are candid about this. Their system, they write, “cannot yet meet the standards of top-tier publications nor even do so consistently for workshops” [1]. Their own researchers reviewed all three papers and concluded that none met the bar for the main conference [1].

What a workshop review can and cannot catch
This is the part I find most interesting. The developers list the system’s common failure modes themselves: naive ideas, incorrect implementations of the main idea, a lack of methodological rigor and made-up details such as inaccurate citations [1]. In the announcement, they give an example where the system credited a well-known type of neural network to the wrong authors. Most of these are exactly the problems a reviewer reading once, without the code, is poorly placed to catch.
I do not see that as a scandal. Workshops are meant to be a lighter venue, a place to share early, rough and even negative results. With seven in ten papers accepted, reviewers are mostly asking whether a paper is worth discussing, not whether every claim in it is correct. The Sakana team also points out that the back-and-forth between reviewers and authors, which improves papers at top conferences and journals, is not part of the workshop process.
So what did the experiment test? Mostly, whether a machine-made paper can read as competent machine learning research to three reviewers in a single pass. That is a real ability. But it tests the surface of a paper: the framing, the writing, the plots, experiments that look sensible. In my view, peer review is rarely set up to check whether the code does what the text claims, and human authors can slip through the same gap.
Why it is still a milestone
The strongest case for the milestone is easy to state. Starting from a broad topic, a machine produced a complete study that scored, in the team’s words, “higher than many other accepted human-written papers at the workshop,” according to Sakana AI. The authors also report that paper quality, as judged by their own automated reviewer, rose steadily as the underlying language models improved [1]. If that trend holds, the gap to main conferences may not last long.
The authors name the downside themselves: systems like this could overload review systems and add noise to the scientific literature [1]. To me, that is the real headline. A system that clears a 70% bar one time in three, and can be run again and again, could put real pressure on reviewers who are already stretched thin.
In my own research I build deep-learning models for breast MRI, including the pipelines that process the images and train and test the models. That work taught me how much sits behind a single results table, and how little of it a reviewer ever sees on the page.
The AI Scientist showed that a machine can write a paper that clears a light, single-round review. It did not show that the paper was right, and it reminds us that peer review was never built to prove that either.
My wish list is short. Venues should ask for code with every submission, human or AI, and some reviewers should actually run it. Clear rules on disclosing AI-generated work should be settled before the next experiment, which is the gap the authors say their withdrawal rule was meant to respect [1]. And when you see a headline saying an AI paper passed peer review, check three things: which venue, who chose what got submitted, and whether it was actually published. Here the answers are a workshop, humans, and no.
References
[1] C. Lu, C. Lu, R. T. Lange, Y. Yamada, S. Hu, J. Foerster, et al., “Towards end-to-end automation of AI research,” Nature, vol. 651, no. 8107, pp. 914-919, Mar. 2026, doi: 10.1038/s41586-026-10265-5.
Related Articles




