Is the hardware really getting better? It feels performance per watt is not getting better at all which is the metric that will matter eventually when supply-demand stabilizes.
As it is, it seems the improvements are about making the hardware cheaper (as in capex, not opex).
This is just feels from me from what I hear on the news and see on the products though.
The DGX Spark apparently consumes up to ~150W while being able to run many models at decent speeds.
I think really good efficiency is possible right now, but the GPU makers don't want to make their consumer GPUs too good for AI - if the cards were more efficient, it'd be much easier to run multiple - while data center ones have a bunch of additional power overheads.
Solar power and batteries are getting cheaper and cheaper at the moment. So Watts should become cheaper in the long run.
Especially when chips are becoming cheaper (in the capex sense), then you can afford to only run them when power is cheap.
Btw, from where do you take the notion that performance per Watt ain't increasing? We are also still using what's more or less general purpose GPU hardware; we could get a lot further if we were willing to specialise more. Which would be the natural avenue to explore, if progress in general purpose hardware slows down. Google is already looking.
Like I said, just feels I have from the consumer-hardware space. For several generations of GPU now most improvements come from packing more transistors into a larger die than packing more transistors closer to each other.
GPUs have been getting physically bigger with huge heatsinks and fans to support those bigger dies power consumption. Just compare the TDPs:
2020 RTX 3090: 350W
2022 RTX 4090: 450W
2025 RTX 5090: 575W
Bigger dies means lower capex of course, but the similar opex (maybe slightly lower as there is less physical hardware to maintain).
I seen some specialized hardware like google's TPUs. Not sure how they compare on performance per watt with GPUs though. Regardless the manufacturing processes are still the same (EUV) which is the thing that hasn't been improving. A fully optimized specialized hardware can at most deliver a single-time linear improvement (that could be very significant, for example 30% is still huge of course) and then little compared to normal GPUs.
I don't think renewable power generation is going to massively reduce costs for data centers, especially considering power transmission hasn't meaningfully reduced in cost. If anything the only thing that I think will have significant impact for data centers would be dedicated nuclear power plants physically located right next to the data center.
In fact I expect power generation to get more expensive as demand can increase faster than supply can be established. I imagine setting up new solar farms and transmission lines to be significantly harder (as in, takes longer time due to approvals and so on) than new data centers (which requires a single large location and I assume less approvals).
If you have somewhere to put them, you can get 2 440W panels for ~$350 if you're okay with intermittent or mostly daytime use. Or for ~$130/kWh you can extend that with batteries. Then you need an inverter or some kind of regulator, but all in your fully capitalized power is still less than an AMD or Intel GPU and a lot less than an nvidia GPU for home use (the context of this thread is how intelligence is not limited to datacenter deployments and can be done at home). Most home users probably aren't going to leave it running overnight all the time, so you really only need storage for morning/evening, maybe. If you had a 27B model on a cerebras-like chip, it'd probably be much faster than any human could interact with it, so it'd race to idle just like CPUs.