AlphaGenome Atlas, and the Gap It Is Built On
Published:
Today Google DeepMind released the AlphaGenome Atlas: precomputed predictions for 9 billion single-nucleotide variants — every single-letter change possible in the human genome — across hundreds of human and mouse cell types. About a petabyte, which they note is more than thirty times the size of the AlphaFold Database.
It is worth understanding both what this is, and what it is built on top of.
A decade of making the window bigger
Regulatory variant prediction has a clean intellectual lineage, and almost all of it is one idea repeated: show the model more sequence at once.
DeepSEA (2015) established the paradigm — a convolutional network taking DNA in and predicting chromatin tracks out, so a noncoding variant could be scored by how much it perturbed those tracks. Basset (2016) did accessibility. Basenji (2018) used dilated convolutions to reach ~131 kb of context. Enformer (2021) swapped in transformer blocks and reached ~196 kb, buying enough range to connect enhancers to promoters. Borzoi extended to 524 kb and modelled RNA-seq coverage directly.
AlphaGenome, announced 25 June 2025 and published in Nature in January 2026, takes 1 Mb of sequence at single-base resolution and predicts eleven modalities at once — RNA-seq, CAGE, PRO-cap, splicing, DNase, ATAC, histone marks, transcription factor binding, contact maps. It led on 22 of 24 benchmark evaluations.
Ten years, roughly a thousandfold more context. The Atlas is what you get when you run that model over every position in the genome and store the answer.
The result that should temper the enthusiasm
In 2023, Nature Genetics published a study that tested four of these models on paired personal genome and transcriptome data. The finding was uncomfortable: the models are good at predicting expression across genes, and weak at predicting expression across individuals. They frequently get the direction of a cis-regulatory variant’s effect wrong.
The sharpest part was the baseline. Per-gene regularised linear regression on nearby variant dosages — a model that learns no generalisable sequence features at all — explained substantially more cross-individual variation than the deep models did.
That distinction matters because across-individual is the clinically relevant axis. Nobody needs a model to tell them that a housekeeping gene is expressed in liver. They need to know what this patient’s variant does.
AlphaGenome does improve here. Against Borzoi on eQTLs it moves tissue-weighted Spearman ρ from 0.39 to 0.49 and sign auROC from 0.75 to 0.80. Real progress, and still a long way from settled.
The third time they have run this play
The Atlas is not a new kind of object. It is the third instance of a strategy DeepMind has now used repeatedly:
AlphaFold DB precomputed protein structures. AlphaMissense (Science, 2023) went further — it scored all 216 million possible amino-acid substitutions across 19,233 human proteins, yielding 71 million missense predictions, and classified 89% of them. Human expert curation had covered roughly 0.1%. The Atlas applies the same move to regulation.
The pattern each time: enumerate the entire variant space, precompute, publish as infrastructure. Running a model is a research act that requires a GPU and a willingness to install something. Looking up a table is not. That shift in who can participate is the actual product.
Regulation came last for a structural reason. A missense variant has essentially one answer per protein, so AlphaMissense fits in a downloadable file. A regulatory variant’s effect depends on cell type and tissue, so the answer is not 9 billion rows but 9 billion variants times thousands of tracks. Hence a petabyte, and hence a portal and an API rather than a download.
What a table changes
Two things happen when a prediction becomes a lookup.
The first is good and obvious: the marginal cost of asking goes to zero, and the set of people who can ask expands enormously.
The second is subtler. A stored prediction loses its uncertainty on the way into the table. When you run a model yourself, its confidence, its assumptions, and its failure modes are in front of you. When you query a row, you get a number. The 2023 direction-of-effect problem does not disappear because the output has been tabulated — it just becomes much harder to see.
Precomputation also freezes one model version as the reference. AlphaFold DB shaped what “the predicted structure” means for years. A regulatory atlas at this scale will exert the same gravity, and it will be far more expensive to regenerate.
DeepMind is explicit that AlphaGenome “has not been validated for, and is not approved for, any clinical use.” That line is correct and it will be under continuous pressure, because a fast, free, comprehensive table is exactly the kind of resource that gets used slightly beyond its warrant.
The conclusion
The Atlas is a genuine contribution and the most useful form this work could take. It is also a bet: that precomputing 9 billion answers is worthwhile even though the best-documented weakness of this model family is precisely the axis those answers will be read along.
Both things are true. The right response is not to distrust the resource but to keep asking the question the 2023 paper asked — does it get the direction right for this variant, in this tissue, in this person? — and to be suspicious of any workflow where a petabyte of convenience quietly removes the occasion to ask.
Infrastructure changes what is easy. It does not change what is known.

