Google seems to have anorexia when it comes to model intelligence. They have an internal hard constraint on price per token it seems, and they are trying to squeeze out intelligence with limited compute.

I wonder if there is something with their TPU cycles that makes them want to postpone training a new model. My guess is that they have been on the same base model for 6 months and they may have waited for the next gen TPUs to train Gemini 4, which greatly limits how much intelligence they can increase and forces them to do cost efficiency increases.

It's probably a mix of things but I do think they are viewing "edge AI" as their strategic play: on-device, small efficient models (Android / iOS) and instant AI summaries in google search etc. So all of their focus is on delivering strong performance in a compute constrained environment.

I do think it's still also simultaneously true that they have an actual problem with competing with current frontier progress. It's just that has gone from an existential threat to something they are willing to defer addressing because they see the long game for them sitting at the smaller end.

Could it be that they have to serve their models to billions of users?

> Could it be that they have to serve their models to billions of users?

And how is that different from their competitors exactly?

[dead]

I'd guess they did model-hardware codesign but the design ended up limiting the scaling capability of the model (i.e. they overoptimized too soon).

Google Cloud is probably Google Deepminds biggest competitor. Big company kinda bullshit.

How so?

Google cloud sells compute out from under Deepmind to other labs. So they basically are in competition with Google cloud for compute.