Capitalism eats itself this way. Second and third order effects will collapse the demand.
You need to keep the market healthy, not some insane Bitcoin style HODL pump - that's how you get wrecked.
I mean I'm not a neoclassicalist but I've read all of them. I'm in consensus with them here. There's a bunch of theories on what a healthy market is but what we're currently seeing matches none of them.
It's short term profitable but long term disastrous, especially in a world where new mathematics and techniques could literally collapse the demand overnight.
Imagine if some paper hits arxiv and the 256 GB requirement for some model now becomes 64. Woops!
Some clever trick about how attention heads and context Windows work could potentially slash a bunch of requirements by giant margins and all they're doing is firing the starting gun at that global race with every obscenely priced unit they sell.
But if prices were reasonable, this wouldn't be an apocalypse. It'd be fine. Consumers wouldn't rush to 64GB, they'd say " Cool I can multitask now at 256" or " great I can do horizontal scalability' or something else.
But no they created the market conditions so now what would happen is the consumer will immediately flip the 192GB they don't need on eBay, hoping to snatch a profit before the prices tank and the second hand market will be flooded the rug will be pulled out from the luxury pricing and everyone will get screwed.
This has happened in electronics markets before. Many times.
When Engels talked about the grave diggers of capitalism they were looking at it through a 19th century labor/manufacturing lens but arguably this same dynamic is at play here.
> Imagine if some paper hits arxiv and the 256 GB requirement for some model now becomes 64. Woops!
If you were a DRAM manufacturer, isn't this exactly the kind of thing that would make you think twice about investing years and $billions in new fab construction?
What "second- and third-order effects" do you suppose will collapse the demand for RAM? The people complaining most loudly about RAM costs are the people who want to run local models; if that becomes popular it will supercharge RAM demand, because locally-hosted models can't parallelize runs from many users the way cloud-hosted ones can. I don't see any slackening in RAM demand at any point in the foreseeable future, even if the big AI companies all go bust.
This is all hypothetical and debating hypotheticals isn't productive so let's roll back to markets.
Let's say ram used to cost $100 and now that same unit costs $1000. You paid say $500x1,000 for that unit during the price increase or some price where you can currently flip for profit.
You have a very expensive data center and you're in debt financed on the premise that you have these special computers.
Now a new technique comes out and it turns out you only need 1 memory unit for something that used to require 8 or 4 or some meaningful multiplier.
This stuff happens all the time. It's why we don't use BMP files on websites or serve giant MOV files on YouTube. It's why postgres queries are faster now than they were 10 and 20 years ago.
You rent out your machines. You need to service your debt.. Demand may 8x overnight to accommodate but you have a monthly bill to pay and that's unlikely. It's likely going to drop.
Think about it. Your customers are paying maybe $10,000 a month and serving their customers. Now they can drop that to $1,250.
On market if you were to sell some of that ram you have 100% profit right now but not for long.
Jevons paradox assumes unlimited capitalization, zero debt servicing, infinite time horizons...
We live in the real world so what do you do?
Historically the answer has been "sell that shit"
There's an aphorism for this "stairs on the way up elevator on the way down"
If we had a healthy market with sane prices where you can't flip the thing you bought for 100% profit the answer would be "create more value."
>Your customers are paying maybe $10,000 a month and serving their customers. Now they can drop that to $1,250.
Or they could stay at $10,000 per month since they are willing to pay that much already.m, so they just use AI more and in more places.
> locally-hosted models can't parallelize runs from many users the way cloud-hosted ones can
Why not? Unlike many other workloads, LLM inference actually seems pretty suitable for decentralization (effectively stateless means no availability concerns; bandwidth and latency are relatively forgiving too).
I think locally-hosted models at the org level will definitely be somewhat popular, but you seem to be talking about decentralizing for people's personal, non-business use, and I just don't think that's going to happen to any real degree.
People who say they want local runs really mean it: they want local runs on hardware in their room, not on some decentralized system which, if it existed, would almost certainly just be a worse, less-reliable version of cloud hosting. I'm not saying nobody would use it, but it sounds a lot like things like IPFS, which have also completely failed to displace either cloud storage or buying a bunch of disks for your own private use.