AlphaGenome Atlas turns nine billion possible DNA changes into a research map
Google DeepMind’s new AlphaGenome Atlas precomputes predicted molecular effects for every possible single-letter change in the human genome. Its value is practical: it helps researchers decide what to test next, while leaving diagnosis and clinical interpretation to experiments and genetics professionals.
A genetic variant can be only one DNA letter away from its reference sequence and still be difficult to interpret. The problem is not a lack of possible variants. It is the opposite: the human genome contains so many possible changes, in so many regulatory contexts, that researchers cannot test them one by one in the laboratory. The majority are not in protein-coding regions, either. They sit in the less familiar parts of the genome that help decide when, where and how strongly genes are switched on.

On 8 September 2026, Google DeepMind introduced AlphaGenome Atlas, a searchable catalogue of predictions for roughly nine billion possible single-nucleotide variants across the human genome. The number refers to every possible one-letter substitution at each position in a reference genome: at each DNA base, the three alternative bases create three possible changes. The Atlas precomputes what the AlphaGenome model predicts those substitutions could do to molecular processes such as gene expression, RNA splicing, chromatin activity and DNA contacts.
That description matters. The Atlas is not a list of nine billion confirmed disease mutations. It is not a database of personal diagnoses, and it does not establish that a predicted molecular change will cause illness. Its immediate contribution is more modest and more useful: it gives biologists a way to rank hypotheses before spending time and money on experiments.
The bottleneck is interpretation, not sequencing
Modern sequencing can identify far more genetic differences than clinical or research teams can confidently explain. A sequence report may contain variants that are already well understood, variants with strong but incomplete evidence, and variants of uncertain significance. The last category is not a diagnosis. It is a statement that the available evidence does not yet support a reliable classification.
The difficulty is especially pronounced outside protein-coding regions. Protein-coding exons account for only about 2% of the human genome, while regulatory DNA helps control the activity of genes across different cell types and tissues. A change in a regulatory element may affect whether a gene is transcribed, how much RNA is made, or how an RNA transcript is spliced. The relevant evidence may be distributed across a long stretch of DNA and may matter only in a particular biological setting.
This makes a simple lookup table based only on the nearest gene inadequate. A variant can influence a distant gene, have different effects in different tissues, or alter a molecular step without producing an obvious clinical consequence. Researchers therefore need more than a position on a chromosome. They need a set of testable predictions about what the sequence change might alter.
AlphaGenome was built for that kind of prediction. It can process DNA sequences of up to one million base pairs and produce many types of output at high resolution. The model compares a reference sequence with a sequence carrying a variant, then estimates changes in molecular properties. Those outputs include gene-expression signals, transcription-factor binding, chromatin accessibility, three-dimensional DNA contacts and RNA-splicing patterns.
The Atlas applies that approach at a scale that individual research groups normally could not reproduce for every possible single-letter change. Google DeepMind says the resulting precomputed dataset is about one petabyte. Instead of asking every researcher to run the model for each candidate variant, the service makes the predictions available through a web interface and an API for academic research.
What the Atlas actually adds
The original AlphaGenome model already allowed researchers to score selected variants. The new Atlas changes the workflow. It shifts part of the work from on-demand computation to a prepared reference resource, so a scientist can begin by exploring a region or a candidate variant rather than first building a large computational pipeline.
That change is useful in several situations. A rare-disease team may have a list of variants found in a patient or family and need to decide which non-coding changes deserve functional testing. A population-genetics group may want to examine whether rare variants near a gene are likely to have similar regulatory effects. A laboratory studying gene regulation may need to choose a small set of substitutions for reporter assays, CRISPR perturbations or RNA measurements. In each case, the model can help narrow the search.
The Atlas also introduces an AlphaGenome Variant Impact score, or AVI score. DeepMind describes it as a combined score that brings together AlphaGenome predictions for regulatory effects and AlphaMissense predictions for protein-altering variants. A single score cannot replace the underlying evidence, but it can help researchers sort a large candidate list and inspect the variants that appear most consequential across the available signals.
This is the practical distinction between prediction and proof. A score is useful when it improves the order in which scientists investigate questions. It becomes misleading when it is treated as a final verdict. The Atlas is best understood as a prioritization layer over biology, not as an automated authority on what a variant means for a person.
Early examples show the intended use
Google DeepMind describes two early applications. In one, researchers at the Broad Institute used the AVI score to prioritize variants in an unsolved rare-disease case. The system highlighted a variant in the DNM1 gene and predicted that it could create an abnormal splice site. The prediction provided supporting evidence in a case that was then resolved through the wider process of genomic investigation.
The important detail is the word “supporting.” The result did not come from a model alone. A rare-disease diagnosis generally depends on the patient’s phenotype, family relationships, population frequency, the gene–disease connection, prior observations and, where available, functional evidence. A prediction about splicing can make a candidate more attractive for laboratory work or expert review, but it does not remove the need for those other lines of evidence.
In another example, a researcher used Atlas predictions alongside data from more than 54,000 UK Biobank participants. Grouping variants according to their predicted molecular effects reportedly revealed 22% more non-coding genetic associations than an approach that did not make the same use of predicted impact. The analysis also identified 19 genomic regions linked to body-mass index among the highest-impact variants.
That result is an example of how a model may help with statistical organization rather than offer a new medical conclusion. Researchers often face a noisy collection of rare variants, each observed too infrequently to analyse alone. If predictions can group variants that appear to perturb the same regulatory mechanism, those groups may become easier to study. The association still needs independent analysis and biological validation.
The examples also reveal what the Atlas is not designed to do. It does not explain a person’s complete phenotype from their genome. It does not calculate an individual’s future health. It does not determine how a variant behaves in every tissue, at every age, under every environmental condition. It identifies molecular hypotheses that can be combined with other evidence.
Why non-coding DNA makes this hard
The phrase “non-coding DNA” can create the wrong impression. Non-coding does not mean useless, and it does not mean that every base has a known regulatory role. It means that the sequence does not directly encode a protein. Many non-coding regions help control gene activity, but regulatory function is often conditional. A sequence can matter in a liver cell but not a neuron, during development but not adulthood, or only after a particular signal activates a pathway.
The training data for models such as AlphaGenome therefore matters as much as the architecture. DeepMind says AlphaGenome was trained using measurements from public resources including ENCODE, GTEx, the 4D Nucleome project and FANTOM5. These projects measure different aspects of gene regulation across selected human and mouse cell types and tissues. They provide valuable coverage, but they are not a complete recording of every cell state in every person.
A model learns patterns from what has been measured. If a tissue, developmental stage, ancestry group or cellular condition is poorly represented in the training data, the model may have less reliable information for that setting. The output can still be useful as a hypothesis, but its uncertainty should not be hidden behind a precise-looking number.
There is a second challenge: biological distance. Gene regulation is not confined to a short neighbourhood around a gene. Enhancers and other regulatory elements can act over long distances, and DNA folds into three-dimensional structures that bring distant regions into contact. AlphaGenome can use a context of up to one million base pairs, which is substantially longer than many sequence models, but that is still a finite window. Google’s technical documentation lists a one-megabase context horizon as a limitation, meaning interactions beyond that range cannot be captured in a single prediction query.
There is also a difference between molecular output and organism-level outcome. The model may predict that a variant changes an RNA-splicing signal or lowers a gene-expression signal. That does not by itself establish that the change causes a disease, determines disease severity, or responds to a particular treatment. The path from DNA to health includes cell biology, development, physiology, environment and chance.
A useful map still needs ground truth
The most valuable next step after an Atlas prediction is usually not another prediction. It is an experiment or an evidence review designed around the specific claim. If the model predicts altered splicing, researchers may test RNA from a relevant tissue or use a validated cellular assay. If it predicts a change in enhancer activity, they may compare reference and alternate sequences in a reporter system, then check whether the result holds in a more realistic cellular model. If a variant is associated with a trait, researchers need statistical replication and a plausible biological mechanism.
Experimental design matters because computational predictions can be correlated with the data used to train or evaluate the model. A result that looks convincing in silico may fail when the sequence is tested in a different cell type or with a different assay. Conversely, a model can help laboratories choose experiments that would otherwise be too numerous or too expensive to attempt. Its strongest role is often to make validation more selective.
Clinical genomics has its own rules for combining evidence. ClinGen’s guidance for variant classification uses an ACMG/AMP framework that incorporates population data, computational evidence, functional studies, segregation, phenotype and other information. Computational predictions are one part of that process, not a replacement for expert curation. ClinGen explicitly warns that its guidance is not intended for direct diagnostic or medical decision-making without review by a genetics professional.
AlphaGenome’s own terms are similarly clear. Google DeepMind says AlphaGenome Atlas has not been validated or approved for clinical use. The model is intended for research, and predictions should not be used as medical advice, a diagnosis or a basis for changing treatment. That boundary is not a minor disclaimer. It is the difference between a tool for generating evidence and a regulated clinical test.
What researchers can do with it now
For academic users, the Atlas lowers several barriers at once. A scientist can query a candidate without running the full model locally. A team can inspect non-coding variants that would otherwise be difficult to rank. A student can explore how a one-letter change is predicted to affect several molecular processes at the same time. A computational group can use the API for smaller analyses and build its own comparisons around the returned scores.
The resource may also make negative results more informative. If a variant scores as having little predicted effect across several relevant molecular outputs, it may move down a laboratory’s priority list. That does not prove the variant is harmless, but it can help allocate scarce experimental capacity. The reverse is also true: a strong predicted effect can identify a candidate for deeper work without declaring it pathogenic.
There are practical limits. The public API is intended for non-commercial research, and the project’s documentation says it is better suited to smaller or medium-scale analyses than to requests involving more than one million predictions. Researchers must also keep track of the genome build, reference sequence, transcript choice and biological context used for an analysis. A coordinate without that metadata is not a reproducible result.
Precomputation creates another responsibility: versioning. A genome atlas is not a timeless answer. Models can be updated, reference assemblies can change, annotations can improve and experimental evidence can overturn an earlier interpretation. A serious analysis should preserve the model or Atlas version, the queried variant, the reference genome, the output fields and the date of access. Without that record, later users may not be able to understand why a result changed.
Why this is good technology news, carefully stated
The good news is not that AI has solved the genome. It has not. The good news is that a familiar research bottleneck has been attacked in a concrete way. Instead of asking scientists to choose from an almost unmanageable space of possible single-letter changes, AlphaGenome Atlas gives them a broad, searchable first pass across that space. It can bring non-coding variants into the same practical workflow as better-known coding variants and help connect different molecular signals around one candidate.
That matters because discovery often depends on deciding what to test next. A laboratory does not have unlimited time, samples or funding. Better prioritization can shorten the route from a genomic observation to a well-designed experiment. In rare disease, it may help a team notice a plausible regulatory or splicing mechanism. In population studies, it may help group variants into biologically meaningful sets. In basic research, it may reveal patterns that suggest which parts of the genome deserve closer measurement.
The limits are equally important. The Atlas predicts molecular effects, not complete clinical outcomes. It covers a reference genome, not every personal genome context. It cannot see biological processes that are absent or weakly represented in its data. Its finite sequence window leaves some long-range interactions unresolved. Its outputs need validation, careful uncertainty assessment and integration with established genetic evidence.
That combination is a healthy standard for judging computational biology. A research tool earns trust when it makes the next experiment clearer, records what it knows and leaves room for what it does not know. AlphaGenome Atlas appears useful on those terms: not as a diagnostic oracle, but as a map for navigating a problem that has become too large for manual inspection alone.
Sources and further reading
- Google DeepMind: AlphaGenome Atlas — announcement, Atlas scope, AVI score and early research examples.
- Google: AlphaGenome Atlas overview — dataset scale, access model and examples involving rare disease and UK Biobank data.
- Nature: DeepMind’s new genome atlas — independent scientific-news context.
- Google DeepMind: AlphaGenome model — model inputs, predicted molecular properties, benchmarks and research-use warning.
- AlphaGenome API documentation and repository — access, supported outputs and practical usage limits.
- Google Cloud AlphaGenome documentation — stated limitations including research-only use and the one-megabase context horizon.
- ClinGen variant-classification guidance — how computational, functional, clinical and population evidence fit into expert variant interpretation.
Comments
Sign in to comment.
No comments yet.