Google DeepMind Maps 9 Billion Possible DNA Variants
DNA is often explained as a codebook or set of instructions for producing proteins, and ultimately, life. Some stretches of DNA, called genes, code for proteins, but the vast majority of DNA is considered “noncoding.” Some of it has no known function, while other segments are critical to regulating

DNA is often explained as a codebook or set of instructions for producing proteins, and ultimately, life. Some stretches of DNA, called genes, code for proteins, but the vast majority of DNA is considered “noncoding.” Some of it has no known function, while other segments are critical to regulating gene activity. These regulatory elements can interact in complicated ways, and their effects can vary across different cells and tissues. Some also influence genes located far away in the genome. Understanding how changes in DNA affect this regulation “is fundamental to understanding most disease,” says Carl de Boer, a genomicist at the University of British Columbia. That’s why researchers are working to understand what every imaginable small variation in human DNA across the entire genome might mean for gene regulation. A recent AI tool built for that purpose from Google DeepMind, AlphaGenome, was originally announced in 2025. In January, a paper published in Nature provided more details, and the model was released for public noncommercial use. The AI model can compare an original DNA sequence with an altered one and predict how the change might affect gene expression and other regulatory activity. But researchers had to select the variants they wanted to test, write code, and run the computationally demanding model themselves. Now DeepMind has done that work in advance for all 9 billion possible single-letter changes to a reference human genome. Today, on 8 September, DeepMind announced the creation and public release of the AlphaGenome Atlas, an online repository of precomputed predictions made using the AlphaGenome model. The Atlas offers a more approachable interface for scientists, without the need to write code or run the AlphaGenome model themselves. It also includes a much-requested new feature, a single-number impact score intended to show at a glance if a variant is likely to be meaningful. “Understanding our DNA is a grand challenge,” says Pushmeet Kohli, VP of science at Google DeepMind. “Understanding this language of life can unlock so many things.” The AlphaGenome predictions have some important limitations. For example, many diseases are associated with multiple genetic variants. And although AlphaGenome looks at a relatively large segment of DNA surrounding the variant in question—1 million base pairs—some DNA sequences, called enhancers, can regulate genes over very long distances, sometimes beyond the model’s field of view. Their effects are difficult to predict. But the Atlas could still help scientists filter possibilities and prioritize lab experiments that would validate its predictions. In that way, it could greatly accelerate work in fundamental biology, disease research, and treatment development, says Žiga Avsec, the genomics lead at DeepMind. “It seems like they made a useful resource for people,” says de Boer, who recently helped create a framework for better comparisons of computational models similar to AlphaGenome. He is not affiliated with DeepMind. Although de Boer considers AlphaGenome the “field’s leading model,” he notes that it’s also “very slow and computationally intensive.” The Atlas could benefit people without access to newer hardware, or simply reduce the number of people repeating the same simulations. The Atlas is freely available for noncommercial research, with the potential for commercial licensing. Computing 9 Billion Predictions The entire human genome contains roughly 3 billion base pairs. At each position there are three possible single-nucleotide substitutions, and therefore 9 billion variants in the Atlas. The complete dataset is around 1 petabyte. “When we started thinking about this project, it seemed impossible to do that computationally,” says Avsec. Early estimates told the team they would need to improve their calculation speed by a factor of 80 in order to compile the Atlas in a reasonable amount of time. To reach that target, the team gained advantages using a few different techniques, including model distillation, GPU kernel optimization, and the elimination of redundant calculations. “There was a lot of thought and engineering that we had to do in order to make this happen at this scale,” says Avsec. AlphaGenome and the Atlas build on years of related work at DeepMind. In 2020, AlphaFold predicted the three-dimensional structure of proteins from amino-acid sequences. In 2023, AlphaMissense predicted whether 71 million possible variants that alter proteins were likely benign or pathogenic. Similar to the new Atlas, prediction results from those projects were made available in a public database. The Atlas allows a scientist to look up a single variant and see more detailed information about the model’s prediction, including 11 different output types. But the top-line figure is a single-number impact score, which by its nature is a simplification of many aspects of those predictions. “It has a clear use, but it also is probably going to be easily misinterpreted,” says de Boer. “We’re talking about a very complex system, and there’s a lot of moving parts.”
Key Takeaways
- •DNA is often explained as a codebook or set of instructions for producing proteins, and ultimately, life
- •This story was reported by IEEE AI, covering developments in the research space.
- •AI advancements continue to reshape industries — read the full article on IEEE AI for complete coverage.
📖 Continue reading the full article:
Read Full Article on IEEE AI →
