I found the OP insightful and worth a read. Thank you for sharing it on HN.

The only aspect that is poorly analyzed by the OP is business model viability. All players are investing insane amounts of money in infrastructure with the expectation that their future profits will justify all that investment. The winner or winners in the AGI race, they believe, will find the proverbial "pot of gold at the end of the rainbow."

The OP glosses over questions of business model viability with a brief qualitative discussion and very little hard data. For example, to earn an annual return > 10% on every trillion dollars of capital sunk into infrastructure, the owners of that infrastructure must earn free cash flow (operating profit less investment) in excess of $100 billion per year in perpetuity. Is that feasible? Why? How?

The OP does not really consider such questions.

>The winner or winners in the AGI race, they believe, will find the proverbial "pot of gold at the end of the rainbow."

The hubris of this is really astounding too. There is no technology out there that some company develops and has not been reverse engineered and copied and manufactured at scale by competitors before long. You can't stop this from happening. People will leave the company or be poached and proliferate what they have built in the past. Every country that wanted a nuke has a nuke, after all.

The only way to keep the secret fire from leaking out would be to have AGI's first move be to lock the doors and prevent anyone from ever leaving company property again.

I think the article's analysis is basically right in a vacuum. That is, I think it's clear that inference is a viable business model. But what isn't clear is whether it will be such a profitable business model for any given company that it will justify the investment that company has taken. I kind of think the winners might be a follow-on generation of companies that focus on this commodity inference business model instead of the invent-machine-god-first "business model" and thus are wiser about their level of investment and capital costs.

> I think it's clear that inference is a viable business model.

You may be right. I'm not so sure. Inference looks like a viable business model for those operators that have SOTA infrastructure in place, but the investment required to have it is enormous, and appears to be never-ending, because if an operator stops investing aggressively, its infrastructure quickly becomes non-competitive, and customers will quickly leave for alternatives. SOTA infrastructure is a moving target.

Renting out compute is a viable business model. That's part of what the big and the small cloud providers do. It's just (on its own) not a "get rich quick" business model anymore.

Why would it be any different for inference? If we believe OP, it'll just become part of regular compute infra, and thus part of the renting-out-compute business model.

I think it's an open question if the current generation of inference investment will pan out, but in the long term, there'll be a balance between investment cost and margin, just as in every other industry.

What about inference on a distilled version of someone else's model? What about inference on an open weights model? I don't think all LLM business models are going to revolve around developing state of the art frontier models.

Right. This is what I was implying in my comment, but it's good to be explicit about it. I think there will be successful companies that just sell inference against the best models they can get without needing to invest anything (or very little) in training anything new. This could even be a spinoff of one of the big frontier labs. I would personally rather invest in an IPO for a company that only owns the gpt-6-astra implementation and infrastructure than in openai itself. There is certainly less upside, but IMO also way less downside risk. (I'm not saying there is any chance this kind of a spin-off would happen, it's just a thought experiment.) And I think this same calculation applies to open weights inference providers.

Edit to add: Or or might just be AWS / GCP / Azure that benefits from this business model. They're already pretty good at selling commodity infrastructure.

Maybe... I'm not enough of an expert on the financials to say, but it seems to me that inference should be able to recoup the cost of SOTA infrastructure, unless you then also use a large portion of that infrastructure to train new models. And I also think the race to remain SOTA itself is also largely a function of training, because my understanding is that training benefits more from the leading edge of hardware.

But yeah, I definitely don't have high confidence in any of this!

Yes, that makes sense. I don't have "the" answers either :-)

> I think it's clear that inference is a viable business model.

Only if you also have the model thats better than anyone else's.

As soon as models are free, or there are no newer models (assuming thats going to happen, and thats not a given) then the only thing you can compete on is price.

This means that the only thing you have to differentiate is either price, speed or ease of use. (or regulatory capture...)

We are at pets.com level of spend currently. Unless model development becomes cheaper, then we are going to run out of novel debt but not really debt mechanisms.

I mostly agree with this!

But I also think there are multiple ways to differentiate. There is at the very least: "intelligence", price, latency, throughput, reliability. It's not clear to me yet what this looks like, but maybe there is also a services and integration level of differentiation. And then there is the universal stuff: sales, marketing, branding. And then on the other side of the ledger there is operational efficiency, management capability, cost of capital, that kind of stuff.

I mean, there is no kind of "model quality" difference between AWS and GCP or between Delta and Southwest or between Wal-Mart and Costco, etc. but all of these businesses remain viable in very competitive markets.

I totally agree that the level of investment / capex is not sustainable though! But I think what's going to happen is that it is not going to be sustained, while AI continues past that point as a viable business (but maybe with different specific companies leading that industry).

A phrase comes to mind: "Your margin is my opportunity."

The problem source is that the "cost" of tokens are taken at face value from business that are losing money at record speeds. E.g. https://artificialanalysis.ai says "doing task A costed us $10 using OpenAI", and that is the "cost" the OP used as basis for "tokens are cheap". Meanwhile OpenAI is losing $19 for each $1 in revenue... So right now OpenAI should be charging around $200 to do task A just to break even, but that would mean their use base would collapse.

There are two different things occurring here.

One is how much does it cost OpenAI to train the model.

The other is, if I stole OpenAI's model how much would it cost for me to run it?

R&D costs versus operational costs. Operational costs are very likely profitable. R&D is catastrophically expensive currently.

There is also the cost of the inference hardware that gets ignored because they already have it from training.

The main problem with the scenario of just doing inference is it relies on nobody else training models better than yours. As long as people are training private models that are better than yours, just inference isn't a viable business model.

They are already turning profits and inference has shown to be a cash cow. And they've already secured compute for the next several years.

Some frontier labs are reporting positive "adjusted EBITDA" (earnings before interest, taxes, depreciation, and amortization, with extra adjustments to make the figure positive).

Free cash flow (operating profit less investment), actual cash coming in, is deeply in the red.

EBITDA can be a sensible measure of profitability when there isn't much need for additional investment. That doesn't seem to be the case with these operators. They need to invest aggressively to avoid losing customers to competitors. All of these operators have made multi-year commitments to invest more in infrastructure. In addition, they have guaranteed quite a bit of debt to fund it.

Maybe it all will work out fine (and I sure hope it does!), but I didn't see any hard data from the OP, or from you, supporting that view.

EBITDA might make sense for the resellers who package up open weight models and sell inference. It is not appropriate for the labs who have billions in debt for RAM, new data centers, gobbling up competitors, etc.

Those real debt obligations are going to want to be paid back.

Definitely. The question is: Is it enough to recoup the enormous capital costs and justify the level of investment they've received. I think there's a decent chance that it will be. But maybe not. And the longer they keep focusing on training new models more so than on inference, the more uncertain I become that it's all going to work out.

Honestly at this point with training costs I don't see how it could ever pay itself back unless you get RSI, in which the talk of money really isn't the main problem any longer.

We're in a situation where AI isn't going to go away, but whatever financial mode we're in right not is not going to work.

Labs are playing money games with EBITDA, which is not uncommon, but also hides the extent to which they are in the red (deeply, deeply, in the red, and projected by them to get worse).

Who is the "they" that are turning profits?

People are saying. If you know you know.

Think of it like spending $1 trillion to become the next Google. That hardware itself may never turn a profit. But if 10 years for now you're the software provider that owns the ecosystem around too cheap to meter tokens you've got a money printing machine.

Exactly. "Cost-to-distill" is a critical parameter. Right now usage of frontier models for all tasks is both subsidized and irrationally popular even at the subsidized price. Deepseek would solve most tasks faster and 10x cheaper. I agree with the author that just as Deloitte exists, frontier labs will exist. But not because their products are proprietary technical marvels or gods, but rather because of branding.

DS wouldn't be 10x cheaper than the subsidized subscription plans from openai/anthropic. Although it is of course much cheaper than the enterprier/API pricing- I think if you're on the subscription plans, you can't beat that on performance per price.