Centralized inference can easily increase batch size, leading to huge efficiency gains in the usual scenario where most users have just one or very few session. Using local resources efficiently requires some way to increase the batch size. I'm not sure if we are there yet.

I think the point is that if people are able to run inference on their laptops batch size efficiency won’t matter.

And before that, businesses will be able to get decent results with dedicated inference hardware.

Jevons paradox: large purpose-fit data centers increase efficiency such that you can use AI in more places, and use more tokens for those tasks.

The future is not a single chat bot session of bs=1. The future is many agents performing many tasks in parallel for a single user. Large GPU clusters will always have the edge in efficiency.

Centralization without proper controls against monopolization becomes sloth and gluttony. If american labs were constrained like china, their models would benefit.

The abstract benefits are quickly outstripped. The same way adding more highway lanes never improves gridlock. Its inducement.