Yeah, local LLMs are probably an order of magnitude less efficient at a fixed level of "intelligence" if not more.

Inefficient per watt, yes, but local inference capacity is greatly underutilized in aggregate. If a model can run on a machine that already exists, that's a bunch of additional chips that don't need to be built.