Research

AI-generated text

AlphaGenome Atlas: precomputed predictions for 9 billion single-nucleotide variants to accelerate genomic research

Google DeepMind has released AlphaGenome Atlas, a public resource providing precomputed predictions for roughly 9 billion possible single-nucleotide variants across the human genome.

AlphaGenome Atlas: precomputed predictions for 9 billion single-nucleotide variants to accelerate genomic research

Google DeepMind has introduced AlphaGenome Atlas, a resource that provides precomputed predictions for roughly 9 billion possible single-nucleotide variants (SNVs) across the human genome. The Atlas offers molecule-level effect predictions for each possible single-letter DNA change and is available for academic use via a free, web-based portal.

What it is and why it matters

Interpreting how genetic variants alter molecular biology remains a major bottleneck for understanding disease and biology. Laboratory testing of every possible SNV (~9 billion) is infeasible. AlphaGenome, an AI model from Google DeepMind, predicts variant impacts on biological processes; AlphaGenome Atlas scales those predictions across the entire genome and packages them into an accessible dataset and tools.

Contents and main features

AlphaGenome Atlas bundles several interconnected resources:

  • Molecular effect predictions: thousands of predicted molecular effects per variant across multiple aspects of gene regulation, spanning hundreds of human and mouse cell types and tissues.
  • AVI score (AlphaGenome Variant Impact): a single numeric score combining predictions from AlphaGenome and AlphaMissense (the model for predicting effects of protein-altering variants) to prioritize variants by likely impact.
  • AVI feature attributions: links each AVI score to the biological features driving it (for example, predicted disruption to RNA splicing or gene expression).
  • DNA sequence motifs: a comprehensive collection of over 2,500 recurrent short DNA sequences (motifs) and their genomic locations.

These resources support tasks from rapid variant ranking to in-depth functional analysis across coding and non-coding regions (the genome is ~2% coding and ~98% non-coding).

Technical scale and access

  • The Atlas is approximately a 1-petabyte dataset, more than 30 times larger than the AlphaFold Database.
  • AlphaGenome Atlas is accessible via an intuitive website portal, the AlphaGenome API, and as a skill integrated with Google Antigravity.
  • The AlphaGenome base model is available for academic use on GitHub and through the AlphaGenome API; commercial availability on Google Cloud (Model Garden and Cloud) is planned. The Atlas is available for non-commercial use from launch via the website.

Early academic use cases

AlphaGenome Atlas has already been applied by external research collaborators in multiple areas:

  • Solving unsolved rare diseases: in collaboration with the GREGoR Consortium, researchers used the AVI score to prioritize candidate variants in unsolved rare disease cases. Laura Covill and Anne O’Donnell-Luria from the Broad Institute and colleagues identified and experimentally validated a variant affecting the DNM1 gene that is strongly linked to epileptic encephalopathy. AlphaGenome’s predictions indicated the variant created an incorrect splice site, causing an abnormal protein extension; experimental screens validated this mechanism and revealed nearby variants with similar effects.

  • Mapping rare non-coding variants related to protein levels and complex traits: Gareth Hawkes (Medical Research Council fellow, University of Exeter) applied Atlas predictions to whole-genome data from over 54,000 UK Biobank participants. By grouping rare variants according to predicted molecular effects, Hawkes detected 22% more non-coding genetic associations than standard approaches and identified regulatory variants that influence levels of circulating proteins such as PLA2G7 (linked to aging) and EGLN1 (a cellular oxygen sensor). Focusing on the 1% of non-coding variants Atlas predicts to be most impactful, he found 19 genetic regions for further targeted study of body mass index.

  • Identifying regulatory motifs: Julia Zeitlinger and Melanie Weilert at the Stowers Institute for Medical Research used the motif annotations to classify which transcription factors affect only DNA accessibility versus those that can activate or repress gene expression.

Performance and interpretability

According to the developers, the AVI score shows best-in-class performance on multiple variant pathogenicity and rare disease benchmarks. AVI feature attributions provide interpretable links to molecular processes (e.g., splicing, expression) that underlie the score, helping researchers form testable biological hypotheses.

Future directions and integration

The Atlas is presented as a baseline resource rather than an endpoint: as AlphaGenome and related AI models improve, the genomic maps will become more comprehensive and precise. Atlas resources are designed to integrate with broader agentic systems, such as Google Antigravity, to support end-to-end scientific workflows.

Availability and deployment

AlphaGenome Atlas is available for non-commercial academic use via the project website and programmatically through the AlphaGenome API. The AlphaGenome base model is on GitHub and accessible via the API; commercial cloud availability is planned on Google Cloud (Model Garden and Cloud offerings).

Acknowledgements

The work involved collaborations with University of Exeter, Broad Institute, Boston Children’s Hospital, Stowers Institute for Medical Research, Harvard University, Memorial Sloan Kettering Cancer Center, Center for Genomic Medicine at Massachusetts General Hospital, and the University of Kansas Medical Center.

Contributors named by the project include Jun Cheng, Kyle R. Taylor, Lauren Nicolaisen, Joshua Pan, Clare Bycroft, Matteo Perino, Tom Ward, Raina W. Thomas, Natasha Latysheva, Gareth Hawkes, Laura E. Covill, Melanie Weilert, Maile J. Hirschmann, Xi Dawn Chen, Robin N Beaumont, V Kartik Chundru, Michael N Weedon, Simon Bourdareau, Hoyin Chu, Dhavi Hariharan, Thais Kagohara, Lucas Tenório, Yosuke Ushigome, Amanda Stafford, Courtney A. Shearer, Barbara Ikica, Ada Fang, Mouad Naciri, Victoria Johnston, Richard Green, Elisa Lai Hong Wong, Vincent Dutordoir, Anne Mottram, Adam Gayoso, Eirini Arvaniti, Guido Novati, Heidi L. Rehm, Fei Chen, Caleb A. Lareau, Caroline F Wright, Anne O'Donnell-Luria, Julia Zeitlinger, Pushmeet Kohli, Žiga Avsec, among many others. The team also thanked numerous colleagues for technical support, feedback and communication expertise.

Conclusion

AlphaGenome Atlas delivers a large, precomputed map of predicted molecular effects for every possible single-nucleotide variant in the human genome. By combining variant prioritization (AVI), feature attributions, motif annotations and broad accessibility, the resource aims to accelerate discovery in rare disease research, population genetics and fundamental molecular biology, while remaining extensible as models and data improve.