On 18 December 2024, The Lancet Digital Health published online a set of recommendations with a long name and a fairly simple request [1]. The STANDING Together recommendations ask the people who build health datasets, and the people who train artificial intelligence (AI) on them, to be open about who is in the data and who is not. There are 29 of them, and they appeared in the journal’s January 2025 issue [1].

Reading them, I kept coming back to one idea. In medical AI papers, the patients who are missing from a dataset, or hidden inside it, too often get a single line in the limitations section. I think they belong much closer to the center of how we judge a model. A single average performance number can look excellent while one group of patients is quietly being failed, and a large, diverse dataset does not fix that on its own.

  • STANDING Together is a set of 29 recommendations, shaped with input from more than 350 people in 58 countries [1].
  • They come in two parts: one for documenting health datasets, one for using them [1].
  • They ask teams to name the groups missing from a dataset and to report how a model performs for each group, next to the overall number [1].
  • Chest X-ray models trained on over 700,000 images underdiagnosed several underserved groups, including female patients [2].
  • A widely used care algorithm underestimated how sick Black patients were because it predicted costs rather than illness [3].

How the recommendations were built

The draft items came from a systematic review of existing approaches and a survey of stakeholders [1]. They were then refined through a Delphi study, in which a panel votes on each item, sees how others voted, and votes again. In this case, 194 participants from 25 countries voted on 32 candidate items over three online rounds and one in-person meeting [1]. That final meeting took place over two days in Birmingham, UK, in June 2023 [1].

A public consultation and an interview study widened the circle: in total, more than 350 representatives from 58 countries contributed [1].

The authors also turned the lens on themselves. They point out that 58 countries should be seen against the 193 member states of the United Nations, and that relying on English and on electronic communication will have left some people out [1]. That is exactly the habit they ask of everyone else.

Two jobs: describe the data, then use it honestly

The first part holds 18 recommendations for documenting datasets, written for the people who collect and curate them [1]. The second part holds 11 recommendations for using datasets, written for the people who build and test AI tools with them [1].

On the documentation side, one item asks for a summary of the groups present in a dataset and, just as important, for curators to “Highlight any known missing groups within the dataset and any reason(s) for their missingness” [1]. The authors also explain how people can vanish from data even when they are technically in it. A group might be too small to show up, lumped into a catch-all category such as other or mixed, or hidden because nobody looked at combinations such as age and gender together [1].

On the use side, teams are asked to decide in advance which groups might be at risk of worse performance or harm, and then to report how the tool performs for each of those groups compared with its overall performance [1]. They are also asked to look for gaps in groups they did not think of beforehand [1]. The authors say the recommendations are meant to prompt questions rather than serve as a checklist, and that missing information about a dataset’s limitations should itself count as a limitation [1].

Imagine a skin-check app that reports 94% accuracy across ten thousand photos. Now imagine that only a few hundred of those photos show darker skin, and nobody ever checked the app’s accuracy on them separately. The headline number is true, and it still tells a person with darker skin almost nothing about whether the app will work for them.

That is the gap these recommendations try to close. The headline figure is accurate, which is exactly what makes it easy to trust too far.

What an average can hide

A 2021 study in Nature Medicine by Seyyed-Kalantari and colleagues shows why [2]. They note that standard practice for medical image classifiers is to report performance on the overall population, whatever subgroup each patient belongs to [2]. They trained chest X-ray models on three large public datasets and on a combination of all three, 707,626 images from 129,819 patients [2].

The models “consistently and selectively underdiagnosed under-served patient populations” [2]. Underdiagnosis here means the model labels someone who has disease as healthy, which can delay their care [2]. In MIMIC-CXR, the only dataset that recorded race and insurance, female patients, patients under 20, Black patients, Hispanic patients and patients with Medicaid insurance all had higher rates of it [2]. Some combinations fared worse still, such as Hispanic female patients [2].

The authors are careful about limits. They say their results may not hold in health systems where sex, race or insurance work differently, and that the labels, generated automatically from radiology reports, could themselves be a large source of bias [2]. Those caveats matter, but they do not change the lesson: the overall score hid a pattern.

Doctor holding up and studying a chest X-ray
In one large study, chest X-ray models missed disease more often in some patient groups than others.

Being in the data is not the same as being treated fairly

Female patients were not a missing group in those chest X-ray data. They made up about 45% of the combined dataset [2], and the models still underdiagnosed them more often [2]. Showing up in the numbers did not protect them.

In 2019, Obermeyer and colleagues examined a widely used commercial algorithm that helped health systems pick out patients with complex needs for extra help [3]. At the same risk score, Black patients were considerably sicker than White patients [3]. The algorithm predicted health care costs rather than illness, and because unequal access to care means less money is spent on Black patients, cost was a biased stand-in for need [3]. By some measures of predictive accuracy, cost still looked like an effective stand-in [3]. The authors estimated that fixing the disparity would raise the share of Black patients getting additional help from 17.7% to 46.5% [3].

Black patients were in that data. The problem was what the algorithm was told to predict. STANDING Together uses this same case as an example and notes that once the bias was found, it could be reduced by reformulating the algorithm [1]. The authors go further: even perfectly measured data would not stop social and structural inequalities from being written into a dataset [1]. More diverse data is a good start. It is not a guarantee.

What to ask of every health dataset

This is close to home for me. I build deep-learning models that use breast MRI to help tell inflammatory breast cancer apart from other locally advanced breast cancers, and in that kind of work a single accuracy number is the easiest thing to report and the least informative.

Who is missing from a dataset, and how a model does for the people who are there, should sit next to the headline result, not in a footnote. A diverse dataset is a start; checking performance group by group is the part that tells you whether it worked.

If you build models, write down before you start which groups might be failed, and report results for each of them beside the overall figure. If you curate data, say who is missing and why, even when the answer is uncomfortable. And if you read about a medical AI tool that performs brilliantly, ask a simple question: brilliantly for whom? The patients who are absent from a dataset, or buried inside its average, are still real patients, and they will meet the tool in the clinic whether or not anyone measured how it works for them.

References

[1] J. E. Alderman, J. Palmer, E. Laws, M. D. McCradden, J. Ordish, M. Ghassemi, et al., “Tackling algorithmic bias and promoting transparency in health datasets: the STANDING Together consensus recommendations,” The Lancet Digital Health, vol. 7, no. 1, pp. e64-e88, Jan. 2025 (published online Dec. 2024), doi: 10.1016/S2589-7500(24)00224-3.

[2] L. Seyyed-Kalantari, H. Zhang, M. B. A. McDermott, I. Y. Chen, and M. Ghassemi, “Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations,” Nature Medicine, vol. 27, no. 12, pp. 2176-2182, Dec. 2021, doi: 10.1038/s41591-021-01595-0.

[3] Z. Obermeyer, B. Powers, C. Vogeli, and S. Mullainathan, “Dissecting racial bias in an algorithm used to manage the health of populations,” Science, vol. 366, no. 6464, pp. 447-453, Oct. 2019, doi: 10.1126/science.aax2342.

Saleh Ramezani

Saleh Ramezani is the founder of Better Science. Saleh believes that science literacy is crucial for navigating today’s science-driven world. Saleh is currently a post-doctoral researcher at MD Anderson Cancer Center in Houston, Texas.

Get involved

Have something to say about science? Write with us.