People are upvoting this because it has the “Alpha______” prefix. Meanwhile, everyone in the field of genomics knows that AlphaGenome provides essentially zero improvements over the previous SOTA, Borzoi…
People are upvoting this because it has the “Alpha______” prefix. Meanwhile, everyone in the field of genomics knows that AlphaGenome provides essentially zero improvements over the previous SOTA, Borzoi…
> a database that predicts the effects of every possible single nucleotide variant in the human genome. We used the AlphaGenome AI model to pre-calculate the regulatory impact of all 9 billion single-letter genetic changes, resulting in a massive, 1-petabyte dataset.
This is for a database, no? While Borzoi is a model?
> Here, we introduce Borzoi, a model that learns to predict cell-type-specific and tissue-specific RNA-seq coverage from DNA sequence.
https://www.nature.com/articles/s41588-024-02053-6
This type of model, of which there are many, either directly releases their results as a precomputed database right away, or others release that database, or the method gets ignored.
Check out VIPdb for the broad category of methods/databases
https://genomeinterpretation.org/vipdb.html
These are predictors for "pathogenicity". Which is the vague concept of "does it cause genetic disease in humans" where "disease" itself is defined as the broad set of things that "brings patients into the doctor to figure out what's wrong."
The "regulatory impact" part of this is what differs from other predictors, in that it predicts the internal states as measured by several different assays, such as transcription regulating proteins are bound where in the genome, etc. Those predictions may or may not be usefel to people trying to reason about things going on in the cell, but my guess is that it's not going to get much use my molecular biologists, because the way the paper was described is pretty bad, and there are no experimental results I saw towards validating that.
But then, I'm not super interested in this paper. It's my field, but if there's something interesting I'm sure I'll hear about it from colleagues. Google is a fantastic advertising company that sometimes also does a bit of science, but this PR push is just an advertisement. A standard work-a-day paper gets covered as if it were ground breaking, and it will get enough eyes that I feel like I can safely ignore it until somebody in the field points out something interesting.
Any model can be expressed as a database.
Kolmogorov looks at this with a ‘duh’ face
> Meanwhile, everyone in the field of genomics knows that AlphaGenome provides essentially zero improvements over the previous SOTA, Borzoi…
Can you elaborate on this? I'm confused why Google would build something that provides zero improvements over SOTA, Borzoi... as you mention. I'm not familiar with this field, just curious.
Speaking as somebody who has worked within Google Research before: the researchers are under tremendous pressure to publish SOTA and sometimes they juice their results a bit to look competitive when they can't match. This is not uncommon in the field- it's remarkably easy to edit a paper to make yourself look good by omitting information.
One of the most egregious cases of this, in my opinion, is only publishing metrics that cover part of the confusion matrix. “The false-negative rate? That could not possibly matter for a variant effect prediction model; why would we include that in the paper?” Example: AlphaMissense.
The use of an exact quote in an ungrammatical fashion is a bit of a language model smell. I can’t help but be reminded of the purely nonsensical AI interview answers. “It’s a pleasure to meet you, Chick Bongo”
Creating an account 12 minutes ago (from the time of this posting) to comment on how another comment seems to you like a "language model smell" is itself, a language model smell, or a scammer.
Please stop accusing or hinting at others being a language model or bot. Not only is it a dumb waste of time, it's wrong in this instance and you are not only going to continue to be wrong but you have no way to prove or demonstrate that any single post comes from a bot nor the ability to do anything about it if you did in fact believe some comment to be attributed to a bot.
They did benchmark the model and beat the SOTA on every metric though
Is it so bad to have another entrant, especially with the resources Google could bring to bear?
Imagine if the Apple EV had actually happened, you think the EV enthusiasts would roll their eyes like you are?
Are you really taking a holier than thou approach on a google article?