Wouldn't improving LLM efficiency make them even more useful across the board, then they can enjoy the nice economies of scale?
The plan is to have LLM working completely autonomously, in that case, the more resources you have, the better. Perhaps people will use local LLM to ask questions, or coders use them for their personal projects, but that's not where the real money is.
The problem is that if the AI companies pass through the actual costs they're incurring, then charges to those companies will >10x.
If the companies don't see that kind of value (so LLMs don't become dramatically better in some kind of quantum leap from where they are now), they won't want to pay those costs. Already, most AI projects in corporations tend to fail.
If the efficiency of LLMs gets 10x better, then either corporations will "private cloud" their own AI or start using competitors that aren't carrying those kinds of debt loads from the "gold rush" phase.
If it's efficient enough you just run it all locally & screw all the rent seekers who want to tell you how you can't use their model & who will sell and misuse all your data they capture.
What happens when the AI god doesn't appear, and these models plateau in regimes supportable with high-end laptops?
If apple puts an inference SOC in their phone, the datacenters are all dead.
Truly, people have been saying this since 2017 and the Apple Neural Engine has proven them wrong time immemorial.
Apple's own desktops, with the fastest Apple Silicon GPUs and TDPs 20x higher than an iPhone still can't compete for real-world datacenter use even with RDNA clustering. Apple's GPGPU architecture is behind AMD at this point, there's a reason why Apple Intelligence is critically reliant on Nvidia and Google to provide inference backends.