Table of contents
Table of contents

On 16 April 2024, The BMJ published TRIPOD+AI, an updated guideline for how researchers should report clinical prediction models [1]. Most of these tools estimate how likely it is that a patient has a condition now, or will have a particular outcome later [1]. The new version covers models built with classic regression and with machine learning alike, and it replaces the 2015 checklist, which the authors say “should no longer be used” [1].

A reporting guideline sounds like the dullest possible news. My view is simple: a performance number without a clear account of what was done, to whom, with which data, how the data were split and how the gaps were handled is a claim nobody can check. Readers should not have to take “the model performed well” on faith.

  • TRIPOD+AI, published in April 2024, sets out what prediction model studies should report, whether they use regression or machine learning [1].
  • It asks authors to describe their data sources, how data were split, how missing data were handled, and whether data and code are available [1].
  • A review of 152 machine learning prediction model studies from 2018 and 2019 found the median paper reported just 38.7% of the items the earlier guideline asked for [2].
  • The guideline’s authors say plainly that it is not a quality appraisal tool, so a complete report can still describe a poor model [1].

One number, many hidden choices

I build deep-learning models that use breast MRI to help tell inflammatory breast cancer apart from other locally advanced breast cancers, and that work includes the image processing, training and testing pipelines. I know how many decisions sit under a single performance figure: who was included, how images were prepared, which cases were held back for testing, how settings were tuned, and what happened to incomplete records.

Each of those choices can move the final number, sometimes a lot. None of them shows up in the number itself.

The TRIPOD+AI authors write that incomplete or inaccurate reporting keeps readers, a group they say includes patients and the general public, from judging the study’s methods and having confidence in its findings [1]. They go further: poor reporting “might also mask flaws in the design, data collection, or conduct of a study” that could cause harm if the model were used in care [1].

Imagine a friend tells you she ran a mile in five minutes. Impressive, until you learn it was downhill, with a strong wind behind her, and she measured the distance on her phone. The time was real. You just could not judge it without the details.

A model’s accuracy is the same: it may be honest, but it means little until you know the conditions.

What the guideline actually asks for

The TRIPOD+AI checklist has 27 main items [1]. A few stand out to me as the ones a reader most needs.

Where the data came from. Authors should describe the data sources separately for building the model and for testing it, explain why they used those data, and say how representative they are [1]. They should also give the dates the data were collected and the rules for who was eligible [1].

How the data were split. Authors should explain how the data were used, including whether they were divided into parts for building and for testing [1]. The guideline’s glossary adds that test data should not overlap with the data used to train the model, tune its settings or choose between versions [1]. If the same patients sit on both sides of that line, the test is partly a memory check.

What happened to missing data. Authors should describe how missing values were handled and give reasons for leaving any data out [1]. Real clinical records are full of gaps, and quietly dropping incomplete cases can change who the model was really built for.

How the model was built. Authors should spell out every model building step, including how settings were tuned [1]. They should also say whether the data and the analysis code are available [1], and report performance with confidence intervals, including for key subgroups of patients [1].

I also like one small rule. When a detail is unknown or does not apply, authors should say so clearly rather than leave it out [1]. Silence and “we don’t know” are different answers, and readers deserve to know which one they are getting.

Close-up of a clinician's hands filling in a form on a clipboard
Reporting guidance helps readers judge a model, but it is not a quality score.

The gap this guideline is trying to close

The problem was well documented before the update. In a 2022 systematic review, Andaur Navarro and colleagues examined 152 studies that used supervised machine learning to build clinical prediction models, found through a PubMed search covering 2018 and 2019 [2]. They scored each paper against the original TRIPOD checklist.

The median paper reported 38.7% of the applicable items [2]. Only 44 of the 152 studies described how missing data were handled [2]. The full model was presented in just 6 of the 116 studies where that applied, and predictive performance was reported completely in only 9 of 152 [2].

Not everything was missing. Nearly every study said where its data came from [2]. But the details a reader would need to rebuild, check or use the model were usually absent, and the review’s authors noted that most of the models “were unavailable for replication, assessment, or clinical application” [2].

One caveat: that review excluded studies that used machine learning to read images [2], so its numbers do not directly describe imaging AI papers. I would not stretch them that far, though I see no reason to think imaging papers are immune.

Why a checklist is not a quality score

There is a fair case against leaning too hard on reporting guidelines, and it has three parts.

The first point is one the guideline’s own authors make: the recommendations are about “transparently reporting how prediction model research was conducted,” not about how to build a model, and “the checklist is not a quality appraisal tool” [1]. A perfectly filled checklist can describe a model trained on too few patients, tested on the wrong ones, and useless in practice. Complete reporting makes those flaws visible. It does not remove them.

The second point is box-ticking. Some authors will treat a required checklist as paperwork and write one vague sentence per item. The TRIPOD+AI authors seem aware of the risk, describing one item, on patient and public involvement, as meant to go “beyond a mere tick box exercise” [1].

The third point is fit. The 2022 review’s authors acknowledged that some items in the older checklist may suit machine learning studies less well [2], and that some methods can handle missing values by design, which may partly explain why that item was so rarely reported [2].

A checklist cannot make a model good, and I would never read a completed one as a stamp of quality. What it can do is make a weak model easy to spot, and that alone is reason enough to ask for it.

I accept all three points. None of them argues for less disclosure. They argue for reading a completed checklist as the start of an appraisal, never the end of it.

What I would like readers and authors to do

If you write these papers, the guideline’s authors recommend using it early in the writing process [1]. I would open it even sooner, while the study is still being planned, and treat each item as a question a skeptical colleague will ask. If your data or code cannot be shared, say why.

If you read these papers, including news stories about them, look past the headline number. Ask who the patients were, where and when the data were collected, how the test set was kept apart, and what happened to incomplete records. A paper that only tells you the model performed well has told you almost nothing yet.

References

[1] G. S. Collins, K. G. M. Moons, P. Dhiman, R. D. Riley, A. L. Beam, B. Van Calster, et al., “TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods,” BMJ, vol. 385, Art. no. e078378, Apr. 2024, doi: 10.1136/bmj-2023-078378. (Correction published 18 Apr. 2024 updating author affiliations only.)

[2] C. L. Andaur Navarro, J. A. A. Damen, T. Takada, S. W. J. Nijman, P. Dhiman, J. Ma, et al., “Completeness of reporting of clinical prediction models developed using supervised machine learning: a systematic review,” BMC Medical Research Methodology, vol. 22, no. 1, Art. no. 12, Jan. 2022, doi: 10.1186/s12874-021-01469-6.

Saleh Ramezani

Saleh Ramezani is the founder of Better Science. Saleh believes that science literacy is crucial for navigating today’s science-driven world. Saleh is currently a post-doctoral researcher at MD Anderson Cancer Center in Houston, Texas.

Get involved

Have something to say about science? Write with us.