I think the point is that if people are able to run inference on their laptops batch size efficiency won’t matter.
And before that, businesses will be able to get decent results with dedicated inference hardware.
I think the point is that if people are able to run inference on their laptops batch size efficiency won’t matter.
And before that, businesses will be able to get decent results with dedicated inference hardware.
Jevons paradox: large purpose-fit data centers increase efficiency such that you can use AI in more places, and use more tokens for those tasks.
The future is not a single chat bot session of bs=1. The future is many agents performing many tasks in parallel for a single user. Large GPU clusters will always have the edge in efficiency.