In May 2024, the journal Radiology: Artificial Intelligence published an update to CLAIM, the Checklist for Artificial Intelligence in Medical Imaging [1]. It is not an experiment and it does not test any model. It is a list of things a research paper about imaging AI should tell its readers, written with input from a panel of 72 physicians, AI scientists, journal editors and statisticians who completed the process [1]. The authors call it “an educational tool for both authors and reviewers” [1].

That sounds dry, but the checklist carries a message the public rarely hears. Most news about medical AI leads with a single number, usually the AUC. A high AUC tells you a model is good at one specific job. It does not tell you the model will help a patient, and reading the CLAIM items one by one is a good way to see why.

  • AUC measures how well a model ranks sick patients above healthy ones; it does not tell you whether the model’s risk numbers are right [2].
  • CLAIM asks authors to say who was tested, and to explain it if they never tested the model on outside data [1].
  • A model’s risk estimates can be off even when its ranking is good, and a lower AUC model can be the more useful one [2].
  • Higher accuracy has not always meant better outcomes for patients [4].
  • Of 81 non-randomised studies comparing deep learning with doctors, only six were tested in a real clinical setting [3].

What an AUC actually measures

AUC stands for area under the curve, the curve being a receiver operating characteristic (ROC) curve. The name is intimidating, but the idea is simple. A good model “should give higher risk estimates for patients with the event than for patients without the event,” which statisticians call discrimination, and the AUC is the usual way to put a number on it [2].

In plain words: pick at random one patient who has the disease and one who does not. The AUC is the chance that the model gives the sick patient the higher score, with ties counted as half. An AUC of 0.5 is no better than a coin flip. An AUC of 1.0 means the model ranks every sick patient in the test group above every healthy one.

Imagine a teacher who can always tell which of two essays is better, but who gives every essay a grade two letters too high. Her ranking is perfect. Her grades would still mislead any student who took them at face value.

That is a real skill, and a model with a poor AUC is not worth much. But ranking is only the first question, and the rest of this post is about what the AUC leaves out.

Who was in the test

An AUC is always measured on some particular group of patients, and it only describes that group. CLAIM asks authors to “describe how well the data align with the intended use and target population of the model” [1]. It also asks them to spell out who was included and excluded, down to the hospital setting and patient age, sex and race [1].

In a 2020 review in The BMJ, Nagendran and colleagues looked at studies comparing deep learning with expert doctors on medical images [3]. Of 81 non-randomised studies, “only nine were prospective and just six were tested in a real world clinical setting” [3]. Prospective means the patients were enrolled and the data collected going forward, after the study was planned, rather than pulled from old records, and CLAIM asks authors to “evaluate AI models in a prospective setting, if possible” [1].

A new hospital is a new exam

A model learns the scanners, patients and habits of the places that trained it, and a new hospital changes all of those. CLAIM prefers the plain term external testing for checking a model on data from another site, and asks authors who skip it to “note and justify this limitation” [1].

Kelly and colleagues, writing in BMC Medicine in 2019, argued that judging real-world performance needs testing on “adequately sized datasets collected from institutions other than those that provided the data for model training” [4]. In one chest X-ray study, the model’s specificity at one fixed cutoff ranged from 0.566 to 1.000 across five independent datasets [4]. Specificity is how often healthy people are correctly told they are healthy. Same model, same threshold, very different results depending on where it was used.

That threshold point is easy to miss. The AUC summarises every possible cutoff at once, but in the clinic someone has to pick one. The cutoff that looked sensible in the original hospital can flag far too many healthy people somewhere else.

Doctor showing results on a tablet to an older couple
The real test for a medical AI tool is whether care gets better for the people in front of the doctor.

Ranking well is not the same as being right about risk

Many models give a probability, such as a 35% chance of cancer. Whether those numbers can be trusted is called calibration. Van Calster and colleagues, in a 2019 paper that called calibration the Achilles heel of predictive analytics, put it bluntly: “estimated risks can be unreliable even when the algorithms have good discrimination” [2].

One common cause is a change of setting. A model built where a disease is common “may systematically give overestimated risk estimates when used in a setting where the incidence is lower” [2]. The line I would most like readers to remember: poor calibration “may make an algorithm less clinically useful than a competitor algorithm that has a lower AUC but is well calibrated” [2]. CLAIM reflects this too, asking authors to use ROC analysis and, where it fits, calibration curves [1].

Whether patients end up better off

Kelly and colleagues noted that none of the usual performance measures “ultimately reflect what is most important to patients, namely whether the use of the model results in a beneficial change in patient care” [4]. They described a randomised trial of automated reading of fetal heart monitoring during labour that found no improvement in outcomes for mothers or babies, calling it “a cautionary example of how higher accuracy enabled by AI systems does not necessarily result in better patient outcomes” [4].

Meanwhile, the claims in study abstracts ran ahead of the evidence. In the Nagendran review, 61 of 81 studies said in their abstract that the AI was at least comparable to, or better than, clinicians [3]. Only 31 said further prospective studies or trials were needed [3]. CLAIM asks authors to describe the intended use and clinical role of their model, and to “discuss any issues that would impede successful translation of the model into practice” [1].

The fair counterpoint is that a checklist cannot make a study good. CLAIM itself says it was “not designed as a scoring system” [1]. A paper can tick every box and still describe a model nobody needs, and I agree with that. But a checklist can make gaps visible, and visible gaps are much harder to hide behind a single impressive number.

I think about this in my own work. I lead BRACE, a deep-learning framework that uses breast MRI to help tell inflammatory breast cancer apart from other locally advanced breast cancers, and we are extending it through multi-reader evaluation and external validation because a number measured on your own data answers only the first question.

A high AUC means a model has learned to rank. It earns the word success only after it holds up on new patients, gives honest risk estimates at a sensible cutoff, and makes someone’s care better.

So the next time you read that an AI model achieved an AUC of 0.95, look for four more things in the story. Check who the patients were and whether they resemble the people the tool is meant for. Look for a test at a hospital that had nothing to do with building it. Look for any word on whether its risk numbers can be trusted and what cutoff doctors will use. And look for evidence that patients actually did better. If those pieces are missing, the AUC is only the first chapter, and the ending has not been written yet.

References

[1] A. S. Tejani, M. E. Klontzas, A. A. Gatti, J. T. Mongan, L. Moy, S. H. Park, et al., “Checklist for Artificial Intelligence in Medical Imaging (CLAIM): 2024 Update,” Radiology: Artificial Intelligence, vol. 6, no. 4, Art. no. e240300, Jul. 2024, doi: 10.1148/ryai.240300.

[2] B. Van Calster, D. J. McLernon, M. van Smeden, L. Wynants, and E. W. Steyerberg, “Calibration: the Achilles heel of predictive analytics,” BMC Medicine, vol. 17, Art. no. 230, Dec. 2019, doi: 10.1186/s12916-019-1466-7.

[3] M. Nagendran, Y. Chen, C. A. Lovejoy, A. C. Gordon, M. Komorowski, H. Harvey, et al., “Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies,” BMJ, Art. no. m689, Mar. 2020, doi: 10.1136/bmj.m689.

[4] C. J. Kelly, A. Karthikesalingam, M. Suleyman, G. Corrado, and D. King, “Key challenges for delivering clinical impact with artificial intelligence,” BMC Medicine, vol. 17, Art. no. 195, Oct. 2019, doi: 10.1186/s12916-019-1426-2.

Saleh Ramezani

Saleh Ramezani is the founder of Better Science. Saleh believes that science literacy is crucial for navigating today’s science-driven world. Saleh is currently a post-doctoral researcher at MD Anderson Cancer Center in Houston, Texas.

Get involved

Have something to say about science? Write with us.