Table of contents
Saleh Ramezani
Table of contents

Picture a letter from a screening program. Your test came back positive, and the test is described as 99% accurate. Most of us would read that as a 99% chance of having cancer. It is often far from the truth.

In July 1994, the statisticians Douglas Altman and Martin Bland explained why in a one-page note in the BMJ [1]. “The whole point of a diagnostic test is to use it to make a diagnosis,” they wrote. What matters is how often the test gets it right for the person holding the result.

My view is simple. A positive screening result is a reason to take the next step, not a verdict. A label of 99% accuracy tells you less than it seems. The missing piece is how common the disease is among the people being tested.

  • A cancer test described as 99% accurate can still be wrong most of the times it gives a positive result. This happens when the cancer is rare.
  • Take a made-up example: 100,000 people screened for a cancer that 1 in 1,000 people have. A test that is right 99% of the time flags about 1,100 people, and only 99 of them have cancer.
  • In the same made-up example, if 1 in 100 people had the cancer instead, about half of the test’s positive results would be real.
  • In 1994, two statisticians wrote in the medical journal BMJ about screening the general public. They said that when a disease is rare, many positive results are bound to be false alarms, even with a very good test.
  • A positive screening result usually means you need a follow-up test, not that you have cancer. Ask your doctor how often a positive result from that test turns out to be real.

What can accuracy mean?

The word accurate can hide at least three different numbers. Altman and Bland defined the first two in an earlier note from June 1994 [2]. Sensitivity is the share of people with the disease whom the test correctly flags. Specificity is the share of people without the disease whom the test correctly clears.

Their example was a liver scan checked against a firm diagnosis in 344 patients. Of the 258 patients with liver disease, the scan flagged 231, a sensitivity of 90%. Of the 86 patients without it, the scan cleared 54, a specificity of 63%.

A third meaning is overall accuracy: the share of all results that are right. Think of a disease that affects 1 in 1,000 people. A fake test that tells everyone they are healthy would be right 99.9% of the time, and it would never find a single case.

None of these numbers answers a patient’s real question. As Altman and Bland put it, “In clinical practice, however, the test result is all that is known”. You know your result, and you want to know what it means.

The two numbers a patient needs

These answers have their own names [1]. Positive predictive value is the share of people with a positive result who really have the disease. Negative predictive value is the share of people with a negative result who really do not.

For the liver scan, the positive predictive value was 88%: 231 of the 263 patients with an abnormal scan really had liver disease. Of the 81 patients with a normal scan, 54 really were free of it. That is a negative predictive value of about 67%, or 2 in 3. The printed note gives 0.59, but 54 divided by 81 is 0.67.

The key point is that these two numbers are not fixed features of a test. They depend on how common the disease is in the group being tested, which doctors call prevalence. In the liver study, 3 out of 4 patients had disease. Altman and Bland worked out that if only 1 in 4 did, the same scan’s positive predictive value would fall to 45%. Its negative predictive value would rise to 95%.

A made-up example: screening 100,000 people

The numbers below are invented to show the arithmetic, not taken from any real test. Suppose we screen 100,000 people for a cancer that 1 in 1,000 of them have. That means 100 people have the cancer and 99,900 do not.

Now give the test 99% sensitivity and 99% specificity. Of the 100 people with cancer, the test flags 99 and misses 1. The trouble is in the healthy group. The test is wrong about 1% of them, and 1% of 99,900 is 999 people.

So the test gives 99 true positives plus 999 false positives. That is 1,098 positive results in all, and only 99 are real. The positive predictive value is about 9%, roughly 1 in 11. About 10 of every 11 people who get a positive letter are healthy.

The negative results tell a happier story. Of the 98,902 negative results, only 1 is wrong. The overall accuracy is exactly 99%, just as the label promised. The label was true, but it answered a different question.

Imagine a hall holding everyone in this example whose test came back positive, about 1,100 people. Fewer than 100 of them have cancer. The other thousand or so are healthy people waiting, worried, for their next appointment.

Nothing about that test is broken. A tiny error rate applied to a very large number of healthy people still adds up to a crowd.

Woman at home reading a printed letter with medical test results
When a disease is rare, many positive results from even a very good test are false alarms.

Why does a rare disease mean more false alarms?

Now run the same made-up test on a group where 1 in 100 people have the cancer. Out of 100,000 people, 1,000 are sick and 99,000 are not. The test flags 990 of the sick people. It also flags 1% of the healthy ones, which is another 990 people. Now half of the positive results are real.

The test did not change; only the group did. Altman and Bland wrote that a very rare disease keeps the positive predictive value well below certainty, even for a very good test [1]. They concluded that in screening the general population, “it is inevitable that many people with positive test results will be false positives.”

They also noted the flip side. The rarer a condition is, the more you can trust a negative result, and the less you can trust a positive one.

The fair counterpoint is that false alarms can be a price worth paying. A screening test may be tuned to accept more false alarms so that it misses fewer cancers. A missed cancer can cost a life, while a false alarm often costs a follow-up test and some worry. That trade is reasonable if people are told about it honestly. I work on AI for breast cancer imaging, and the same trade-off appears there.

What should you ask after a positive screening result?

A positive screening result means you need a closer look. It does not mean you have cancer. A few questions can help:

  • Ask how many people with a positive result on this test turn out to have cancer.
  • Ask how common this cancer is in people of your age and risk.
  • Ask what the next test is, and how much more certain it will make the answer.

The first question matters most, because it asks for the positive predictive value in plain words. If most positive results are false alarms, that is useful news. It means your worry rests on a small chance, not a likely one.

A test’s accuracy describes the test. What you need to know is what your own result means, and that depends on how rare the disease is. A positive screening result is a question, not an answer.

Start with how many people have the disease, then count the false alarms among everyone else. After a few tries, a headline about a 99% accurate test stops sounding like a promise. It becomes the start of a better question.

References

[1] D. G. Altman and J. M. Bland, “Statistics Notes: Diagnostic tests 2: predictive values,” BMJ, vol. 309, no. 6947, p. 102, Jul. 1994, doi: 10.1136/bmj.309.6947.102.

[2] D. G. Altman and J. M. Bland, “Statistics Notes: Diagnostic tests 1: sensitivity and specificity,” BMJ, vol. 308, no. 6943, p. 1552, Jun. 1994, doi: 10.1136/bmj.308.6943.1552.

Saleh Ramezani

Saleh Ramezani is a researcher trained in medical physics. He believes that science literacy is crucial for navigating today’s science-driven world.

Get involved

Have something to say about science? Write with us.

Related Articles