Google DeepMind has unveiled a massive new database that aims to predict how every single tweak to the human genetic code might play out across the body. The tool, called AlphaGenome Atlas, is a freely available collection of data on 9 billion possible changes to the human genome. It estimates how those changes will affect different tissues and cellular processes, and it scores each prediction to make the results easy to compare.
The announcement comes one year after DeepMind launched AlphaGenome, an earlier tool that predicted how DNA changes alter proteins, cells and the human experience. Now the company has scaled that work to cover every possible base change in the human genome, storing the results in a single searchable database. The goal is to make powerful AI tools accessible to researchers who previously lacked the computing resources to run them.
The scale of AlphaGenome Atlas
DeepMind researchers computed every possible base change in the human genome, then packaged the results into a single database. The total data set runs to 1 petabyte, or 1 million gigabytes, of information. That is a vast amount of raw data, and DeepMind has made it freely available to anyone who wants to use it.
The database is wrapped in an easy-to-use web portal. According to Tuuli Lappalainen, a professor in genomics at KTH Royal Institute of Technology in Stockholm and a senior associate faculty member at the New York Genome Center, the interface is designed to be accessible even to researchers who are not deep bioinformatics experts.
“You really do not need to be an expert in these methods to be able to go there and look something up in a browser,” Lappalainen said.
What AlphaGenome Atlas actually predicts
AlphaGenome Atlas is not a replacement for traditional wet-lab research. It is a computational tool that makes predictions about how DNA changes might affect the body. The database is designed to help geneticists quickly assess the potential consequences of a particular genetic tweak.
The tool assigns a score to each prediction, which is meant to capture how much biological effect a change is likely to have. In an accompanying preprint paper, DeepMind showed that the score could separate disease-linked mutations from harmless ones in a clinical dataset. That suggests the score has some practical value for researchers working through lists of gene variants.
The limits of the predictions
The tool is not a perfect predictor. A Sept. 11 preprint study from a research team led by Katie Pollard, director of the Gladstone Institute of Data Science and Biotechnology and a professor at the University of California, San Francisco, found that while AlphaGenome was adept at finding causal mutations, it persistently underestimated their impact.
Pollard’s team found that the model could not always link changes in regulatory elements to the genes they controlled, especially when those elements were not close together on the genome. Lappalainen said the findings should give researchers pause before treating the predictions as settled answers.
“My perhaps naive hope would be that people take these predictions with an appropriate grain of salt,” Lappalainen said.
The “junk” DNA problem
Most of the human genome is noncoding DNA, which means cellular machinery does not directly read it to make the proteins that keep cells functioning. For years, scientists treated much of this material as “junk” DNA, assuming it had no function.
That view has changed. Geneticists now understand that this code manages how and when the genome’s coding genes are read. Some regulatory instructions are distributed widely throughout the code, rather than being clustered near the genes they control.
AlphaGenome allows researchers to explore how changes in a single pair of DNA letters can influence up to 1 million pairs of surrounding code. That captures much of each gene’s regulatory network, a far greater spread than previous models were capable of.
Why the database exists
Before AlphaGenome Atlas, researchers who wanted to use DeepMind’s models faced a significant barrier. The original AlphaGenome required users to have enough bioinformatics experience to access and use the models’ automated programming interface. Calculating the effects of each change also strained academics’ computing resources.
The new database solves both problems. By precomputing every possible change, DeepMind has removed the need for individual researchers to run their own simulations. The web portal handles the technical work, so scientists can query the database without needing to be experts in machine learning or bioinformatics.
What the experts say
Greg Findlay, a group leader at the Francis Crick Institute in London who is not involved with AlphaGenome Atlas, called the resource a useful addition to the field. His assessment was measured rather than enthusiastic.
“It looks like a great resource,” he told Live Science.
Tuuli Lappalainen, meanwhile, was more cautious. She praised the tool’s utility but warned that it is not a solution to every question in genetics.
“It’s not this holy grail,” she told Live Science.
Lappalainen’s lab has been using AlphaGenome to study the effects of variation in genomes over the past year. She noted that the technology is not entirely new, but she said DeepMind’s models outperform previous approaches in the resolution at which they can predict the impact of changes to the genome.
The path forward
Despite the limitations, Lappalainen sees the tool as part of a broader trend toward more collaborative, data-driven genomics. But she stressed that the field still needs to invest in the kind of experimental work that generates the data AI models rely on.
“We’re still very much data-limited in biology, and that data needs to be created,” Lappalainen said.
She warned that the field would need to give equal focus to “wet-lab” experiments, which work out the accuracy of variant predictions. These tests are tedious and consume much more time than a search on AlphaGenome Atlas’ interface, but they are essential for improving the underlying models.
Key facts at a glance
| Fact | Detail |
|---|---|
| Tool | AlphaGenome Atlas |
| Developer | Google DeepMind |
| Announced | Sept. 8 |
| Database size | 1 petabyte (1 million gigabytes) |
| Predictions | 9 billion possible changes |
| Preprint study | Led by Katie Pollard, Sept. 11 |
| Score metric | AlphaGenome Variant Impact score |
The announcement marks a milestone for computational genetics, but it also raises questions about how much faith researchers should place in predictive models. The database is a valuable resource, but it is not a substitute for careful, hands-on experimentation. As Lappalainen put it, the data needs to be created.
For now, the tool is freely available, and researchers can start querying it today. Whether it lives up to its promise will depend on how well the predictions hold up against the real world.
Source material: “Google DeepMind's latest AI tool promises a new era for genetics research, but experts warn to take its predictions with caution,” Live Science.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

