Where do nine billion DNA changes come from?

The human genome contains roughly three billion DNA positions.

At each position, one of four letters (A, C, G or T) is written. If we take one letter and replace it with any of the other three, we get three possible changes at every position.

Three billion positions multiplied by three possible replacements gives us roughly nine billion single-letter changes.

These are known as single-nucleotide variants, or SNVs.

AlphaGenome Atlas contains a prediction for almost every one of them.

However, it does not cover every possible type of genetic variation. DNA can also lose letters, gain new ones or undergo much larger rearrangements.

So, nine billion does not mean that the Atlas includes every genetic change that could ever happen.

It means that it covers almost every possible change from one DNA letter to another.

So, what does AlphaGenome actually do?

AlphaGenome compares two versions of the same DNA sequence: one with the original letter and one with the change.

It then predicts whether that change could affect how the cell uses the DNA.

For example, could it change how active a gene is? Could it interfere with the way the cell edits RNA before using it? Could it make an important region of DNA easier or harder for the cell to access?

It is important to understand that the Atlas does not simply label a change as “good” or “bad.”

What it does instead, is that it offers clues about what the change might be doing inside the cell and gives scientists a place to begin their investigation.

How does it make those predictions?

AlphaGenome learned from large collections of experiments showing what happens in different human and mouse cells.

These experiments measured things such as which genes were active, how RNA was processed and which parts of the DNA were open and available for the cell to use.

By studying the relationship between the DNA sequence and these measurements, the model learned to recognise patterns.

When researchers enter a DNA change, AlphaGenome compares its predictions for the original sequence with its predictions for the altered one.

If the two results look different, the model can suggest what the change may affect.

But it has not watched that change happen in a real person. It is using patterns from previous experiments to make an informed prediction and that prediction may still need to be tested in the laboratory.

Why is this especially useful for non-coding DNA?

Only a small part of our DNA contains instructions for making proteins.

Most of the genome does not make proteins directly. Instead, some of these regions help control when and where genes are switched on, how active they are and how their RNA is processed.

Changes in these areas are much harder to understand.

A variant may sit far away from a gene and still affect it. It may matter in the brain but not the liver, or during development but not later in life. Some changes have no meaningful effect at all.

AlphaGenome tries to connect these difficult-to-interpret variants with the processes they might disturb.

It does not completely decode non-coding DNA.

But it can give researchers a much better idea of where to start looking.

How does the Atlas decide which changes matter most?

Nine billion predictions are not very helpful if scientists have no way to sort through them.

That is why DeepMind also created the AlphaGenome Variant Impact score, or AVI score.

The score brings together predictions from AlphaGenome and AlphaMissense, another DeepMind model that focuses on changes affecting proteins. It gives each variant a single number based on how strongly it is predicted to affect the cell.

Researchers can then use that score to bring the most promising variants to the top of their list.

The Atlas also shows what may be behind the score. It might suggest that a variant affects RNA splicing, gene activity or the structure of a protein.

AVI score though is not a diagnosis.

A high score does not prove that the variant causes a disease or that it will have any noticeable effect on the person carrying it. It simply means that the variant looks worth investigating.

This is still a prediction

This is the part that can easily get lost in the excitement.

AlphaGenome has not tested all nine billion DNA changes in a laboratory. It has studied patterns in existing data and used them to make an educated prediction about what each change might do.

It might flag a variant because it appears likely to disrupt RNA splicing, for example. That is a useful lead, but scientists still need to test whether the disruption actually happens.

And biology does not always behave as neatly as a computer model expects. An effect may appear in one type of cell but not another. It may be too small to make a real difference, or the cell may find another way to keep things working.

Has it already helped scientists?

There are early signs that it could be useful.

In one rare-disease project, researchers used the AVI score to sort through variants that had previously been difficult to interpret.

One of the variants they identified affected DNM1, a gene already linked to a severe form of epilepsy. AlphaGenome predicted that the change created a new splice site, causing the cell to process the gene’s RNA incorrectly.

But it is important to mention that the researchers did not stop at the AI result. They tested the prediction experimentally and confirmed the effect on splicing. They also found nearby variants that caused similar problems. And that is exactly how a tool like this should be used:

The AI finds a promising clue. The experiment tests whether the clue is real.

How accurate is AlphaGenome?

In the study published in Nature, AlphaGenome performed as well as or better than the strongest comparison models in 25 of 26 tests involving variant effects.

That is impressive

But it does not mean AlphaGenome was correct in 25 out of every 26 predictions.

Each test measured a different task, using a different dataset and method of scoring. The results tell us that AlphaGenome often performed better than the other models included in those comparisons.

They did not confirm that every prediction is accurate.

The best available model can still make mistakes, miss important effects or produce a confident-looking answer where the biology is uncertain.

What can it still miss?

AlphaGenome is powerful, but it can only learn from the biological data available to it.

Some cell types, stages of development and biological conditions are represented better than others. This means the model may recognise an effect in a well-studied tissue while missing one that appears only in a very specific setting.

It also finds some long-distance genetic interactions difficult to predict. A regulatory region can influence a gene located far away, and those relationships are not always captured accurately.

The available data are also weighted towards protein-coding genes, while areas such as non-coding genes and microRNAs remain less well covered.

These are important gaps and do not form reasons to dismiss the Atlas, but reasons to interpret its predictions carefully.

So, what does AlphaGenome Atlas actually change?

Its biggest contribution may be speed.

Researchers often begin with thousands of genetic variants and very little idea of which ones deserve their attention.

The Atlas can help bring the most promising candidates to the top of the list and suggest what they might be doing inside the cell.

Scientists could use it to investigate rare diseases, study difficult non-coding variants, explore changes in RNA splicing and choose better targets for laboratory experiments.

It does not replace those experiments. It helps researchers decide which experiments to do first.

And when time, funding and biological samples are limited, that can make a real difference.

A map is not the territory

There is something genuinely exciting about being able to explore the predicted effects of billions of DNA changes in one place.

But the Atlas is still a starting point. Some of its predictions will lead researchers towards important discoveries. Others may turn out to be incomplete or simply wrong. That is not a failure of the technology. It is part of working with a model built to guide research, rather than replace it.

The real test will be what happens next: which predictions hold up in the laboratory, which ones help solve difficult cases and which ones teach us something new about how our genome works.

AlphaGenome can help scientists decide where to look.

Finding the answer is still a job for biology.

This article is for education and does not replace personalised medical advice.