Agents require training data for RL. This data is rare in comparison to what we feed foundation models, and Google is sitting on a dragon's hoard of sensor and tracking data. I'm not counting them out yet.

Okay but would it kill everybody to start with commoncrawl?

We don't need Google's data. Once we instrument the world, we'll get a Google's worth of data in short order.

It's easier than ever to ingest and label data.

This is happening with or without Google. The recursive improvement does not require them at all.