Saleh Ramezani
Table of contents
Saleh Ramezani
Table of contents

In February 2025, The BMJ published FUTURE-AI, a guideline for building medical artificial intelligence (AI) that is trustworthy and ready for real clinics [1]. It came from a consortium of 117 experts in 50 countries [1], who worked toward agreement over 24 months [1]. The result is 30 best practices that run across the entire lifecycle of a tool, from design and validation to deployment and monitoring [1].

This is a consensus document. It does not test any AI tool, and nobody should read it as a trial. But one line in it stopped me: “Unlike medical equipment, AI currently lacks universally accepted measures for quality assurance” [1].

I think that sentence names the real gap. Too often, a medical AI tool is treated like a product launch: validate once, switch on, move on. My argument is that clinical AI should be treated the way radiation oncology treats its treatment machines, as something that is checked on a schedule for as long as it touches patients.

  • FUTURE-AI, a 2025 consensus of 117 experts, sets out 30 best practices that cover AI from design through deployment and monitoring [1].
  • The guideline says AI, unlike medical equipment, still lacks accepted quality assurance measures [1].
  • Medical linear accelerators are checked against a baseline set at acceptance and commissioning, with daily, monthly and annual tests [2].
  • A hospital switched off a widely used sepsis alert model in April 2020 after the mix of patients changed during the pandemic [3].
  • FUTURE-AI calls for periodic audits, such as once a year, to catch drift and falling performance [1].

How a treatment machine earns trust

A linear accelerator is a machine that delivers radiation therapy. Before it treats anyone, physicists run acceptance tests and commission it, measuring how it behaves in detail. That becomes the baseline. The American Association of Physicists in Medicine (AAPM) put the goal of machine quality assurance (QA) plainly in its Task Group 142 report: to make sure the machine’s characteristics “do not deviate significantly from their baseline values acquired at the time of acceptance and commissioning” [2].

Why keep checking a machine that passed its tests? Because machines change. The report lists sudden problems, such as malfunction or component failure, and also “gradual changes as a result of aging of the machine components” [2]. Many of those baseline numbers also go into the treatment planning software, so they can affect the plan for every patient treated on that machine [2]. A small, quiet drift is not a small problem.

So the report lays out tables of daily, monthly and annual checks [2]. It recommends action levels, so a result can call for an inspection, a scheduled action, or immediate corrective action [2]. And while QA is a team effort, the report recommends that overall responsibility sit with one person, the qualified medical physicist [2].

Imagine a bathroom scale that read correctly the day you bought it. Over a few years its spring weakens, and it shows you two pounds lighter than you are. Nothing looks broken, and you would only find out by checking it against a known weight now and then.

Passing the first test tells you a machine was right on day one. It says nothing about month eighteen.

What FUTURE-AI asks for after launch

Much of FUTURE-AI reads like a QA program for software. It asks for continuous monitoring of what goes into a model and what comes out, to catch things like missing or out-of-range inputs and “erroneous or implausible AI outputs” [1]. That is the AI version of a daily check.

It also asks for periodic audits on a defined timeline, with every year given as an example [1]. The purpose is to spot “data or concept drifts, newly occurring biases, performance degradation” and changes in how clinicians use the tool [1]. That is closer to the annual check.

Two more recommendations stood out to me. A tool should be tested for local clinical validity at each site, and recalibrated if it performs worse there [1]. That sounds a lot like commissioning. And after deployment, the roles of risk management, auditing, maintenance and supervision should be assigned, for example to IT teams or hospital administrators, with responsibility for AI errors clearly specified [1]. That is the ownership question physicists answered long ago, and TG-142 goes a step further by giving overall responsibility to one person.

Clinical team reviewing medical scans across several monitors in an imaging room
A model that worked at launch can drift as patients, scanners and practice change.

When the data moved under a model

Finlayson and colleagues describe dataset shift as what happens when a model underperforms because the data it sees in use no longer match the data it was built on [3]. Their main example comes from a hospital sepsis alert. The University of Michigan Hospital had implemented a widely used sepsis alert model from Epic Systems, and in April 2020 it “had to be deactivated because of spurious alerting” owing to changes in patients’ demographic characteristics during the COVID-19 pandemic [3].

The hospital’s clinical AI governing committee decided to take it out of use [3]. The authors call it an extreme example and say many causes of dataset shift are more subtle [3]. That is the part that worries me more. A flood of false alarms gets noticed. A slow slide in accuracy might not.

Imagine a model trained to flag patients at risk of a serious infection. Then the hospital switches to a new lab test that reports results slightly differently. The model keeps running and keeps producing scores, and nobody notices that the scores quietly mean something different now.

This is why monitoring has to be planned from the start. Finlayson and colleagues recommend a governance committee with the right mix of expertise, an ongoing monitoring process for dataset shift, and a way for frontline staff to flag concerns [3]. In physics terms, that is a QA team, a schedule, and a clear route to report a problem.

The fair objection

There is a reasonable case against pushing this analogy too far. A linear accelerator has physical quantities you can measure with an instrument against a known standard. Many AI tools have no clean equivalent of a calibrated dose meter, and the right answer for a patient may not be known right away. FUTURE-AI itself notes a related gap: “For some of the recommendations, no clear standard on how these should be addressed yet exists” [1].

Consensus is also not proof. The guideline rests on expert agreement, not on a trial showing that yearly AI audits improve outcomes for patients. And QA costs staff time that small hospitals may not have.

I accept all of that. But radiation QA was not born finished either. TG-142 was written partly to update the recommendations of an earlier AAPM report, and to add checks for newer technology [2]. You do not wait for a perfect standard before you start checking. You start with a baseline, a schedule and an owner, and you improve the tests as you learn.

Ownership is the missing piece

I trained in medical physics, and as a graduate student I worked on measuring proton beam dose with scintillation detectors. Now, as a researcher in radiation oncology, I build imaging AI and automated data pipelines, and quality assurance and workflow automation are the parts of clinical physics that interest me most.

From that seat, the question I would ask of any clinical AI tool is simple: who checks it next month? If nobody can answer with a name, a schedule and a baseline, the tool has a launch date but no QA program.

Validation is where a medical AI tool’s safety record begins, and it cannot be where it ends. I would like hospitals to commission AI the way physicists commission a treatment machine, then keep checking it for as long as it runs.

If you work in a hospital that is buying or building an AI tool, ask what the baseline is, how often it will be re-checked, and who owns the result. If you build these tools, plan the monitoring before the launch party. A treatment machine is never finished being tested, and a model deserves the same attention, because the patients it serves keep changing.

References

[1] K. Lekadir, A. F. Frangi, A. R. Porras, B. Glocker, C. Cintas, C. P. Langlotz, et al., “FUTURE-AI: international consensus guideline for trustworthy and deployable artificial intelligence in healthcare,” BMJ, vol. 388, Art. no. e081554, Feb. 2025, doi: 10.1136/bmj-2024-081554.

[2] E. E. Klein, J. Hanley, J. Bayouth, F.-F. Yin, W. Simon, S. Dresser, et al., “Task Group 142 report: Quality assurance of medical accelerators,” Medical Physics, vol. 36, no. 9, pp. 4197-4212, Sep. 2009, doi: 10.1118/1.3190392.

[3] S. G. Finlayson, A. Subbaswamy, K. Singh, J. Bowers, A. Kupke, J. Zittrain, et al., “The clinician and dataset shift in artificial intelligence,” New England Journal of Medicine, vol. 385, no. 3, pp. 283-286, Jul. 2021, doi: 10.1056/NEJMc2104626.

Saleh Ramezani

Saleh Ramezani is the founder of Better Science. Saleh believes that science literacy is crucial for navigating today’s science-driven world. Saleh is currently a post-doctoral researcher at MD Anderson Cancer Center in Houston, Texas.

Get involved

Have something to say about science? Write with us.