
A CT scan of the abdomen shows the liver, pancreas, kidneys, bowel and more in one set of pictures. The radiologist, the doctor who reads the scan, has to check every organ for hundreds of possible problems. A small tumor can hide among gray shapes that all look alike. Can an AI tool help a doctor catch what a busy eye might miss?
A study published in Science on September 17, 2026, tested one such tool, called RADAR [1]. The team trained it on 424,911 stored abdominal CT exams and their written reports from one large hospital in China. Then they tested it on 146 kinds of findings, from tumors to blocked bowels. In a reading test with 26 radiologists, the doctors caught about 10 more real findings out of every 100 when the AI helped them.
That is a real gain. But every test used stored scans, not new patients. And with the AI, the doctors also flagged a few more problems that were not there.
What did the researchers build?
Most medical AI tools do one job, and many need large sets of images labeled by hand. RADAR learned from reports that radiologists had already written, so no one had to mark findings on the images by hand. The training data came from one hospital, the First Affiliated Hospital of Zhejiang University School of Medicine in Hangzhou, China. It covered exams from January 2010 to December 2023.
The model splits each scan into body parts, such as the liver, and matches each part to the report sentences about it. To check for a finding, it compares the scan with two short phrases, one saying the finding is there and one saying it is not. The phrase that fits better becomes its answer.
Eight authors work at Alibaba’s DAMO Academy, a research lab of the tech company Alibaba. The paper says six authors hold Alibaba stock as part of their pay, and six have filed a patent related to the work.
How well did the AI do on its own?
To score the model, the team used a measure called AUC. It asks how often the model ranks a scan that has a finding above a scan that does not. A score of 1 is perfect, and 0.5 is a coin toss.
On 39,160 exams from the same hospital, taken in the first half of 2024, after the training period, RADAR averaged 0.913 across the 146 findings. The next best AI model, retrained on the same data, scored 0.776. The team then tested it in harder settings:
- Emergencies: 27,267 emergency cases, a kind of scan left out of training. It scored 0.904.
- Eight other hospitals: 24,239 exams from eight other hospitals in China. Scores ranged from 0.874 to 0.912.
- Cancers confirmed by tissue tests: 4,333 patients at two Chinese hospitals, with liver, pancreas, stomach or colorectal cancer, or no cancer. Scores ranged from 0.891 to 0.984.
- A US hospital: 5,137 scans from Stanford Hospital, mostly from Western patients. On the 21 findings it could check there, it scored 0.883 with no extra training.
Those are strong numbers. But a high AUC on stored scans is only a first step. I explained the reasons in a post on why a high AUC is not clinical success.
What does beating 23 of 26 radiologists mean?
Yicai Global, a Chinese business news site, wrote that the model’s average performance beat that of 23 of the 26 radiologists [2]. Here is what the paper actually tested.
The team picked 300 cases from the 2024 test set at the same hospital, covering 61 kinds of findings. The 26 radiologists came from 14 institutions. Some were senior, with more than 10 years of experience, and some were junior. Some worked at hospitals in the top grade of China’s national hospital ranking, and some at lower-graded ones. Each doctor read all 300 scans alone and wrote a full report.

Diagram by Better Science based on Zhang et al., Science (2026), doi:10.1126/science.aec6129.
The doctors rarely raised false alarms. When a finding was absent, they correctly left it out 98.8% of the time. But they missed many real findings, and how many depended on their experience and the grade of their hospital.
How was the contest judged? Each doctor gives one result: a share of real findings caught and a share of false alarms. The AI can be set to be more or less cautious, so it gives a whole curve of possible results. Most doctors’ results sat below that curve. Only three senior radiologists did slightly better than the AI. In plain terms, for 23 of the 26 doctors, the AI caught more real findings when it was set to that doctor’s own false-alarm level.
That is a fair comparison, but not a test of real work. The paper does not say the doctors were told why each scan was ordered or given the patient’s history. In a clinic, knowing that someone has sharp pain on the right side changes where a doctor looks.
Did the AI make doctors better?
After a break of at least one month, the same doctors read the same 300 scans again, this time in a random order. This time RADAR listed the findings it suspected. It also showed a heat map, a colored overlay that marks the spots behind each suggestion.
With the AI’s help, the doctors’ sensitivity rose by 10.0 points. Sensitivity is the share of real findings that get caught. Reading off the paper’s chart, the average went from roughly 46 out of every 100 real findings to roughly 56. So the gain is about 10 points, not a relative 10%.

Chart by Better Science from data in Zhang et al., Science (2026), Fig. 4D, doi:10.1126/science.aec6129.
Every group improved. Junior doctors at lower-graded hospitals gained about 10 points. With the AI, they caught 7 points more than senior doctors at the same kind of hospital did alone, though that gap could be chance. Gains were also clear for cancers (7.3 points), emergencies (8.5 points) and diseases of hollow organs such as the stomach and bowel (about 9.5 points). Reading time per case fell by 30.7%.
Imagine a radiologist near the end of a long shift. She has many abdominal scans left to read. A small spot on the pancreas is easy to skip at that hour. An AI tool that circles the spot could make her stop and look again.
That is the hope. The next question is what the help cost.
What about false alarms?
A false alarm is a flag for a finding that is not really there. The measure for this is called specificity, the share of truly clear cases that are correctly called clear.
With the AI’s help, the doctors’ specificity dipped from 98.8% to 98.2%. Here is my own arithmetic on that. Out of 1,000 checks where a finding was truly absent, the doctors flagged about 12 by mistake on their own. With the AI, they flagged about 18.
That is a small change per check. But each report covers dozens of possible findings, and each false alarm can mean another scan, a biopsy or weeks of worry. My post on why an accurate test can still give many false alarms explains how fast these add up. The paper does not say what the extra flags were, or whether they would have led to more tests.

Can a test on stored scans show that patients benefit?
No. Every part of this study used scans that had already been taken and read. That is called a retrospective study. It can show that a tool spots findings on a screen. It cannot show that patients do better.
There is also the answer key. For most tests, the team took the correct findings from the original radiology reports, pulled out by language software. Radiologists settled cases where the software disagreed and spot-checked the rest. Reports can miss things or include errors. The authors list “the absence of a pathological gold standard” as a limitation. That means most findings were not confirmed by tissue tests, although four cancers were, and the AI held up well on them.
They also say their testing in other populations “is not yet comprehensive.” Scans from only one US hospital were tested. Without extra training, the AI could check just 21 of the 30 findings there.
My own reading adds two points, which the authors do not make. First, the doctors always read without the AI first and with it second. Even with a month between sessions, practice or memory may explain some of the gain. Second, working alone, the doctors caught under half of the real findings on average. The team picked these cases to include complex ones. So I read this as a very demanding test, or a strict answer key, more than a picture of everyday care.
The code page is careful too. It says the model “is currently intended for research purposes only” and that “prospective clinical studies are still required” [3]. A prospective study follows new patients forward in time, in real clinics.
Could doctors lean on the AI too much?
People tend to trust an automatic suggestion, even when it is wrong. This is called automation bias, and I wrote about it in a post on keeping humans in the loop.
This study did not measure that risk. A doctor who leans on the tool for years may build less skill of their own. And saving nearly a third of the reading time could ease a heavy workload, or become a reason to assign more scans.
What is new here?
It would be unfair to stop at the cautions. Many imaging AI tools check for one disease, while this one covered 146 findings across 18 body parts.
The reader test also asked the right question: do doctors do better with the AI? The team also shared its code and model, so other groups can test it on their own scans.
My reading is that this is strong early evidence that one AI model can help radiologists catch more on abdominal CT scans. The cost is a few more false alarms. It is not proof that it beats doctors in real care, and it says nothing yet about patient outcomes.
As a researcher trained in medical physics who works on AI for medical imaging, I find the reader test the most useful part. The next step should be a trial in real hospitals, with and without the tool, in more than one country. It should count missed findings, false alarms, extra tests and how patients do.
Bottom line
With RADAR’s help, 26 radiologists caught about 10 more real findings out of every 100 and read faster. By my arithmetic, their false alarms rose from about 12 to about 18 per 1,000 checks. The claim that it beat 23 of 26 radiologists is real, but it came from a screen-based test with answers drawn mostly from old reports.
What we still need is a test with new patients in real clinics. If you have a CT scan, your doctor and radiologist are still the people to talk with about your results.
References
[1] Q. Zhang, J. Zhang, W. Cao, Z. Lu, W. Chang, H. Ding, et al., “An expert-level generalist AI for abdominal CT diagnosis,” Science, vol. 393, no. 6817, Art. no. eaec6129, Sep. 2026, doi: 10.1126/science.aec6129.
[2] S. Dou, “Alibaba’s DAMO Academy Pushes Beyond Single Disease-Detecting AI With New Diagnostic Model,” Yicai Global, Sep. 18, 2026. [Online]. Available: https://www.yicaiglobal.com/news/alibabas-damo-academy-debuts-generalist-ai-for-nearly-150-abdominal-conditions
[3] Alibaba DAMO Academy, “RADAR: An Expert-Level Generalist AI for Abdominal CT Diagnosis,” GitHub repository, accessed Oct. 2, 2026. [Online]. Available: https://github.com/alibaba-damo-academy/damo-radar
Related Articles


