
In February 2025, The Lancet Digital Health published screening performance results from the full Mammography Screening with Artificial Intelligence (MASAI) trial. Across four screening sites in southwest Sweden, 105,934 women were randomly assigned either to screening supported by artificial intelligence (AI) or to the usual standard of two radiologists reading every mammogram [1]. The AI arm found 338 cancers against 262, a detection rate of 6.4 per 1,000 women compared with 5.0 [1].
Those numbers will get the headlines. What interests me more is how they were produced: not a model scoring well on a stored set of old images, but a randomized trial inside a working national screening program, with real women, real radiologists and real follow-up. My view is simple: that kind of evidence is what moves medical AI from promising to trustworthy, and we should expect it far more often.
What a benchmark can and cannot tell you
Most medical AI is judged the same way. Researchers collect a batch of past images where the answer is already known, hold some back, and see how often the model gets them right. That is a benchmark, and it is a good first step: it tells you whether the model learned something real.
It does not tell you what happens when the model is switched on in a clinic. The images were taken in the past, and nobody’s care changed because of the answer. So a benchmark cannot show whether doctors trust the tool too much or too little, whether it slows the workflow, or whether patients end up better off.
A 2020 review in the BMJ by Nagendran and colleagues looked at deep learning studies that compared AI with clinicians in medical imaging. They found only 10 records of randomized trials, two of them published [2]. Of 81 non-randomized studies, only nine were prospective and just six were tested in a real world clinical setting [2]. Yet 61 of those 81 said in their abstract that the AI was at least comparable to, or better than, clinicians [2].
Imagine a new car that aces every test on a closed track. It brakes perfectly, corners perfectly and never runs a red light. You would still want to know how it does on a wet Monday morning in real traffic, with real drivers around it, before you put your children in the back seat.
That gap between the track and the road is what trials are for. The review’s own conclusion was blunt: most non-randomized studies “are not prospective, are at high risk of bias, and deviate from existing reporting standards” [2].
How MASAI was built differently
The MASAI team said plainly why they ran a trial. Retrospective studies had shown promising results for AI in mammography, “however, to our knowledge, a randomised trial has not yet been conducted” [3].
The design was careful in ways a benchmark never has to be. Women were randomly allocated, so the two groups were comparable from the start. The AI (one commercial product, one version) sorted exams into single or double reading by radiologists and highlighted suspicious areas [1]. And the trial had a safety brake built in.
That brake was the step published in The Lancet Oncology in August 2023. After 80,033 women had been enrolled, the team ran a prespecified clinical safety analysis, with a minimum acceptable cancer detection rate set in advance [3]. AI-supported screening cleared it, so the trial was not halted and carried on toward its main question [3]. Our earlier piece on the human in the loop looked at what those interim numbers mean for how radiologists react to AI. Here I care about something else: the order of steps. Check safety early, keep going only if it is safe, then answer the main question with the full sample.

What the trial measured, and what it did not
Cancer detection went up: a ratio of 1.29 in favor of AI [1]. The extra invasive cancers were mainly small and had not spread to the lymph nodes [1]. Workload went down, with 61,248 screen readings in the AI arm against 109,692 in the control arm [1]. Recalls and false positives were not significantly higher with AI [1]. The authors wrote that the findings suggest AI “contributes to the early detection of clinically relevant breast cancer and reduces screen-reading workload without increasing false positives” [1].
What the February 2025 paper did not report is the trial’s primary endpoint: the interval cancer rate, meaning cancers that show up between screening rounds, which is a key sign of what screening missed. That needed two years of follow-up after enrollment [3]. Later update: the interval cancer results have since been published in The Lancet in January 2026, reporting 1.55 versus 1.76 per 1,000 women, which met the trial’s non-inferiority target.
And one outcome is missing entirely. Deaths from breast cancer were not among the outcomes listed [1, 3]. Finding more small cancers is good news only if it leads to fewer women dying, and some extra detection in any screening program may be cancers that would never have caused harm. MASAI cannot settle that.
The fair case against waiting for trials
Trials have real costs. MASAI randomized more than 100,000 women over about 20 months, and its main result needed two more years on top [1, 3]. AI products are updated faster than that, so the version tested may be old by the time a trial reports.
There is also the question of where the trial ran. In this trial, standard screening meant double reading, with two radiologists reading each mammogram, and that was the comparison [1]. A health system that uses a single reader, a different AI product or a different population might see different results. One well run trial in one region of one country is strong evidence for that setting, and weaker evidence for everywhere else.
I accept both points. They are reasons to run trials smarter, with built-in safety checks like MASAI’s and more trials in more places. They are not reasons to accept benchmark scores as the finish line.
What I would ask of the next AI tool
I build deep-learning models that use breast MRI to help tell inflammatory breast cancer apart from other locally advanced breast cancers, including the pipelines that train and test them. Building a testing pipeline makes it clear what a test set can tell you about a model, and what it cannot tell you about the patients.
So when I read about a new medical AI tool, I ask what evidence stands behind it. A high score on a retrospective dataset means the idea is worth testing. A prospective study in a real clinic means it survived contact with practice. A randomized trial with outcomes chosen in advance means we can finally compare it fairly with what we already do.
Benchmarks tell us which AI tools deserve a trial. Randomized trials in real clinics, like MASAI, tell us which ones deserve our trust, and we should stop confusing the two.
If you are a patient, a clinician or a hospital buyer, the question to ask is the same: was this tested on people, in a setting like ours, against the care we give now? If the honest answer is “only on old images,” the tool may still be promising. It just has not earned trust yet.
References
[1] V. Hernström, V. Josefsson, H. Sartor, D. Schmidt, A.-M. Larsson, S. Hofvind, et al., “Screening performance and characteristics of breast cancer detected in the Mammography Screening with Artificial Intelligence trial (MASAI): A randomised, controlled, parallel-group, non-inferiority, single-blinded, screening accuracy study,” The Lancet Digital Health, vol. 7, no. 3, pp. e175-e183, Mar. 2025, doi: 10.1016/S2589-7500(24)00267-X.
[2] M. Nagendran, Y. Chen, C. A. Lovejoy, A. C. Gordon, M. Komorowski, H. Harvey, et al., “Artificial intelligence versus clinicians: Systematic review of design, reporting standards, and claims of deep learning studies,” BMJ, vol. 368, Art. no. m689, Mar. 2020, doi: 10.1136/bmj.m689.
[3] K. Lång, V. Josefsson, A.-M. Larsson, S. Larsson, C. Högberg, H. Sartor, et al., “Artificial intelligence-supported screen reading versus standard double reading in the Mammography Screening with Artificial Intelligence trial (MASAI): A clinical safety analysis of a randomised, controlled, non-inferiority, single-blinded, screening accuracy study,” The Lancet Oncology, vol. 24, no. 8, pp. 936-944, Aug. 2023, doi: 10.1016/S1470-2045(23)00298-X.
Related Articles



