Given how heavily subsidized it is at the moment, the efficiency isn’t as important. Typically efficiency would give you more at lower cost, but with token prices so removed from actual cost that plays less of a role here.

If the inference gets an order of magnitude cheaper, labs can afford to subsidise an order of magnitude more usage for the same marketing cost. So that part of usage will, if not accelerate with efficiency, at least still grow linearly with it. And there is a substantial amount of usage at or above true costs - everyone using a 3P harness, everyone on enterprise contracts, and everyone self-hosting an open weights model in a 3P cloud.