Research

Evaluation of Google's AI for Mammogram Reading in UK Clinical Protocols

Two retrospective and live studies tested Google’s mammography AI against standard UK double-reading protocols.

Two studies examined how Google’s AI system for detecting breast cancer in mammograms would integrate into typical United Kingdom clinical workflows. The research was led by Christopher J. Kelly, Marc Wilson, and colleagues at Google, Imperial College London, University of Surrey, Royal Surrey National Health Service Foundation Trust, and several NHS Breast Screening Centres.

How the system works

Google’s system comprises three convolutional neural networks: one produces embeddings from mammography images, another identifies candidate cancerous regions, and the third assigns a probability of cancer.

Tests and key findings

The authors ran two retrospective analyses and a live integration test to evaluate the system in a standard UK diagnostic pathway.

  • Retrospective test (116,000 mammograms, 2016): using scans from five hospitals, the team selected images from the same women taken up to 39 months apart and compared the AI’s diagnoses with those of human experts. The AI achieved a sensitivity of 0.541 (proportion of true positives identified), significantly higher than the first of two human readings at 0.437. Specificity was 0.943 for the AI versus 0.952 for the human reader — lower but statistically equivalent. The AI also identified 25 percent of cases that human readers had initially missed but that became apparent within three years.

  • Simulation replacing the second reader (46,000 scans): the authors modelled outcomes if the AI replaced the second of two human evaluators. In that role the system achieved slightly better sensitivity and specificity, suggesting that using AI as the second reader could save time while maintaining or improving accuracy. Under clinics’ protocols, cases where cancer was detected or where AI and human disagreed were sent to an arbitration panel. The AI sent 1,800 more cases to arbitration (an absolute increase of 4 percentage points; 5,300 cases in total). Assuming arbitration takes five times the human effort of a single read, the authors estimated that, despite more arbitrations, the AI would reduce overall human effort by roughly 40 percent.

  • Live integration test (2023–2024, 12 clinics, ~9,250 fresh scans): the system labelled new screening scans of women aged 50 to 70 at 12 clinics over a few months. The test did not alter patient care: clinicians made diagnoses as usual and neither clinicians nor patients were informed of the AI assessments. The AI’s median processing time from screening to interpretation was 17.7 minutes, compared with more than two days for the first of two human evaluations. Three months later the researchers checked ground truth for cancer outcomes; as in the retrospective study, the AI showed higher sensitivity than the first human reader and lower but statistically equivalent specificity.

Clinical and operational implications

The studies indicate that Google’s AI identified more cancers and earlier in a typical UK diagnostic workflow, while substantially speeding up image processing. This is clinically significant given that about 2.3 million women are diagnosed with breast cancer annually worldwide and roughly 760,000 die from the disease each year; early detection is critical for improving survival.

At the same time, the research highlights practical challenges: the AI increased the number of cases sent to arbitration, and some clinicians reported distrust of the system’s outputs. The authors note that building clinician trust may require educating physicians about how AI systems work and improving the explainability of AI outputs.

Background and commercialization

Computer-aided detection (CAD) in mammography dates to the 1990s and 2000s, but progress accelerated in the mid-2010s as deep-learning models trained on large mammography datasets began outperforming older methods. In 2020 Google researchers showed that an AI system could match or exceed expert radiologists in screening mammograms while reducing both false positives and false negatives. In late 2022 Google licensed the system to iCAD, and in 2023 Google and iCAD expanded their partnership into a 20-year worldwide commercialization agreement to use Google’s AI as an independent second reader of 2D mammography. The partnership is pursuing regulatory approvals for deployment in double-reading screening workflows.

Conclusion

The retrospective and live evaluations suggest Google’s mammography AI could be a useful adjunct in breast screening programs by increasing sensitivity, accelerating reading times, and reducing human workload. However, questions of clinician trust and the operational impact of increased arbitration remain important issues to address before broad deployment.