In May 2023, the journal Radiology published a clever and slightly uncomfortable experiment. Dratsch and colleagues asked 27 radiologists from three German university hospitals to read 50 mammograms each, with help from what they were told was a new artificial intelligence (AI) system [1]. There was no real AI. The researchers had built a simple display program that showed a suggested rating for each breast, prepared in advance, and on 12 of the 50 cases that suggestion was deliberately wrong [1].
On those 12 cases, most of their ratings were wrong too [1].
It would be easy to read this as proof that AI makes doctors worse. I do not think that is the lesson. My reading is simpler: a human in the loop is a design choice, not a safety guarantee. Whether the person catches the machine’s mistakes depends a lot on when and how the machine’s answer is put in front of them.
A fake AI and 27 real radiologists
For the first 10 cases the pretend AI was always right, which the authors say was meant “to establish the credibility of the AI” [1]. Then came 40 more cases, and hidden among them were the 12 wrong suggestions: half rated a breast as more suspicious than it really was, half as less suspicious [1].
When the suggestion was right, readers at every experience level rated about 80% of mammograms correctly [1]. When it was wrong, the least experienced readers got about 20% right, the middle group about 25%, and the most experienced about 45% [1].
Experience helped, but only partly. The most experienced group did significantly better than the others on the wrong-suggestion cases and still got fewer than half of them right [1]. That group was also tiny, just five people, so I would not lean hard on its exact number [1].
What the headline got wrong
The publisher’s announcement was titled “AI Bias May Impair Radiologist Accuracy on Mammogram”. That wording mixes up two different things. Bias inside an algorithm is a flaw in the software. Automation bias, which is what this study was about, is a human habit: the tendency to go along with what a machine suggests.
The same release said that when the AI suggested the wrong category, readers’ “accuracy fell to less than 20%.” Fell implies a before and after. The study never measured one. The comparison was between different cases read by the same people, some with a right suggestion and some with a wrong one. The authors say so themselves: “we did not investigate the performance of radiologists without the use of the purported AI-based system” [1].
Imagine a teacher who hands out an answer key with a few wrong answers mixed in, then finds that many students copied those wrong answers. That tells you the key pulled them along. It does not tell you how those students would have done with no key at all.
That is the distinction the headline missed, and it matters for how we read every study like this one.
An old problem with a new face
None of this is new. In a 2010 review, Parasuraman and Manzey concluded that automation bias “occurs in both naive and expert participants, cannot be prevented by training or instructions, and can affect decision making in individuals as well as in teams” [2].
In a 2021 study, Gaube and colleagues gave 138 radiologists and 127 internal and emergency medicine physicians chest X-ray cases with written advice [3]. All of the advice came from human experts, but some of it was labelled as coming from AI. Accuracy was significantly worse with inaccurate advice, whatever the claimed source [3]. Radiologists rated the advice labelled as AI lower, yet, in the authors’ words, “their expressed aversion against algorithmic advice did not affect their reliance on it” [3].
I find that last result the most telling. Being sceptical of AI, and saying so, did not protect anyone. That is why I think simply telling users to stay aware of automation bias is a weak safeguard on its own.

The case for AI still stands
The best evidence from real patients points the other way. In the Mammography Screening with Artificial Intelligence (MASAI) trial in Sweden, 80,033 women were randomly assigned to AI-supported screening or to standard reading by two radiologists without AI [4]. In an early safety analysis, the AI arm found 244 cancers against 203, a similar detection rate, and cut the screen-reading workload by 44.3% [4].
MASAI’s radiologists were not shielded from the AI. They could see its risk score for every exam, and marks on the image for the higher-risk ones, while they read [4]. If automation bias were a fatal flaw, that setup should have struggled across a whole screening programme. It did not.
Dratsch and colleagues make a similar point themselves: an AI system “may help novices perform on a level similar to that of more experienced radiologists” [1]. Their test was a deliberate stress test, and I accept that. Automation bias is a cost to weigh against a real benefit, and no reason to throw the tools away. What a large trial like MASAI does not tell us is how often an individual reader followed a wrong cue, and for the one patient behind that cue, that is the number that matters.
Read first, then look at the AI
In Dratsch’s study, the AI’s answer was on screen from the very start of every case [1]. Gaube and colleagues note that when doctors ask a colleague for advice, “they typically ask for advice after their initial review of the case,” while a decision aid shown up front may push them to look for evidence that confirms it [3]. They suggest that giving AI advice only on request may help [3].
My own view is that for many reading tasks, the reader should form an opinion first and only then see the AI output. A second check is only as useful as it is independent. A reader who sees the machine’s answer before forming their own is partly repeating it rather than testing it. The cost is time, so this needs testing rather than assuming.
Imagine a radiologist reading her fortieth mammogram of the afternoon, and the software has already flagged it as low risk. A faint shadow in one corner is easy to wave off when the screen says all is well. Had she looked first and seen the score afterwards, that shadow would have had her full attention before anything told her to relax.
The question is a practical one for me. I lead BRACE, a deep-learning framework that uses breast MRI to help distinguish inflammatory breast cancer from other locally advanced breast cancers, and I have built web apps for image review and reader studies. As we extend BRACE through multi-reader evaluation, deciding when and how readers see the model’s output is part of the work.
A human in the loop is only as good as the loop we design around them. I would rather readers form their own view first and see the AI second, and I would want every study of AI-assisted reading to show how often readers followed a wrong answer.
If you read news about AI in medicine, ask one question: was there a fair comparison without the AI? If you build or buy these tools, ask another: when does the reader see the answer? Being the human in the loop does not make you an independent check. The loop decides that, and we are the ones who design it.
References
[1] T. Dratsch, X. Chen, M. Rezazade Mehrizi, R. Kloeckner, A. Mähringer-Kunz, M. Püsken, et al., “Automation bias in mammography: The impact of artificial intelligence BI-RADS suggestions on reader performance,” Radiology, vol. 307, no. 4, Art. no. e222176, May 2023, doi: 10.1148/radiol.222176.
[2] R. Parasuraman and D. H. Manzey, “Complacency and bias in human use of automation: An attentional integration,” Human Factors, vol. 52, no. 3, pp. 381-410, Jun. 2010, doi: 10.1177/0018720810376055.
[3] S. Gaube, H. Suresh, M. Raue, A. Merritt, S. J. Berkowitz, E. Lermer, et al., “Do as AI say: Susceptibility in deployment of clinical decision-aids,” npj Digital Medicine, vol. 4, no. 1, Art. no. 31, Feb. 2021, doi: 10.1038/s41746-021-00385-9.
[4] K. Lång, V. Josefsson, A.-M. Larsson, S. Larsson, C. Högberg, H. Sartor, et al., “Artificial intelligence-supported screen reading versus standard double reading in the Mammography Screening with Artificial Intelligence trial (MASAI): A clinical safety analysis of a randomised, controlled, non-inferiority, single-blinded, screening accuracy study,” The Lancet Oncology, vol. 24, no. 8, pp. 936-944, Aug. 2023, doi: 10.1016/S1470-2045(23)00298-X.
Related Articles




