In another HN thread about this AlphaGenome Atlas, someone has posted a link to:
https://www.science.org/content/blog-post/mutate-em-all-and-...
which comments the results of this study:
https://www.biorxiv.org/content/10.64898/2026.07.25.740675v1
That study has done in reality what the AlphaGenome Atlas does in fiction, but instead for a human they have done it for one of the simplest viruses.
So they have fuzzed the virus by mutating one by one each position of its DNA.
And various dedicated AI models all made poor predictions of the results of that experiment, which casts doubts about the value of the AlphaGenome predictive map.
A virus is much simpler than a human, but even for that simple virus the effects of most of the mutations could not be predicted. A half of the mutations had harmful effects, and for a half of those it is unknown for now why they were harmful.
For a human the uncertainty about the effects of a mutation will be far greater than for one of the simplest viruses.
Just because a virus genome is small doesn't make it simple, actually quite the opposite (think of it as an obfuscated, compressed package, that is hugely variable with no checksum). In the case of this virus there are even overlapping open reading frames, which is something you would never find in a eukaryotic genome, and almost the entire genome is protein coding, whereas only 2% of the human genome is protein coding.
The real value of AlphaGenome is not the effect of SNPs on protein coding regions (there are other tools for that, like AlphaFold), but rather identifying regulatory elements, such as promoters or alternative splicing patterns or microRNAs, within the ~98% of the genome that hasn't been well characterised yet.
That 98% is vastly underexplored, so a tool like this could help researchers interested in expression profiles or alternative splicing patterns of a protein, identify the source. Obviously not every "important" SNP will be consequential, but it helps narrow the search for that needle in a haystack.
Yep. Sequence-to-function models are still very limited. AlphaGenome Atlas, despite the flashy branding, is unlikely to provide significant benefit to researchers.
May be that's the reason for the alpha naming. We are waiting for a stable release (just kidding).
And it makes a lot of sense why they are limited. DNA is not an instruction set. It's more like a heavily encrypted dataset where the encryption key is the totality of physics and biology. The interactions with the physical world that result in the end product of life are enormously (it would seem hopelessly) complex.
For a machine intelligence to turn DNA sequences into organisim phenotype prediction requires modelling all that in latent space.
I imagine that is going to take a monumental amount of example data
It seems like we can skip much of the expensive modelling and use evolutionary conservation data to shortcut building a full latent space that captures all salient interactions. It's unclear to me whether we truly need to model the entire latent space. And given that biology develops in a generative way with feedback, it may be that attempting to model this using a static latent space is unproductive.
> And various dedicated AI models all made poor predictions of the results of that experiment, which casts doubts about the value of the AlphaGenome predictive map
I would not group AlphaGenome into the pile of failed predictions of other models. AlphaGenome deserves to get evaluated based off its own merits.
Why should we discount this model just because other models that are already thought to be worse weren’t very accurate? I’m not in this field at all, so curious if I’m missing something here.