I think their move makes a lot of sense. Frontier models can probably not advance much further with the datasets we have currently available. Most of the text that goes in is online chatter, images and some scientific texts. That's great for chatbots, knowledge retrieval and programming. But with that database genuine discovery is hard to do. I think for the next step in intelligence the models need to have access to much more data: Data from physics, chemistry, biology experiments - and so on. And they need the data in much higher fidelity than you can currently access. If we just feed AI with all the knowledge we've acquired it's much harder for it to become smarter than us - it basically needs it's own eyes, ears, nose and so on.

> Data from physics, chemistry, biology experiments

There are PetaBytes of important scientific data locked in archival file formats. The first step is to make this efficiently readable.

https://www.earthmover.io/blog/virtual-zarr

https://news.ycombinator.com/item?id=46659254

> Frontier models can probably not advance much further with the datasets we have currently available

I’ve been reading this sentiment on HN since GPT4o, yet models got better and better

There’s a difference between things getting incrementally better and a step function.

LLMs will obviously keep getting incrementally better, but in order to get the kind on jump that LLMs themselves were, the sentiment is that we need something more.

Fable is a clear step function for me

That’s because a significant portion of the work and spending going on at frontier labs is generating new, more curated data, in select domains like software engineering and now biology[1].

[1] https://www.synbiobeta.com/read/anthropic-is-hiring-biologis...

I think that was his argument as well. It looks like there is no more data to find, but then it becomes important to find more data - lo and behold we can make more.

Sounds like the oil scare from 90' - we thought we were gonna run out of oil. But as oil gets more expensive it pays to dig further down to find the stuff that didn't make sense to dig up before.

Yes. The web is enormous, but just think of home much information within a company doesn't even make it into the internal knowledgebase, let alone anywhere public.