This is a good counter argument. But you have to note that this is after OpenAI cut Luna costs by 80%. If you compare launch pricing, Qwen probably comes out ahead on a cost-performance basis.

The luna cost cuts were real though, not a one time promotion or something, due to some optimization (probably distillation?) that openai did.

You assume that openai's inference is profitable and that they aren't just trying to bolster revenue before their IPO.

The only indication that openai is profitable comes from openai (whom I wouldn't trust with any statement, especially when it comes to profitability).

In fact there is evidence that inference is not profitable simply because the rate of losses doesn't seem to reduce as revenue increases: if inference had great margins, we would expect that as revenues increase, the amount of spend on training reduces as a fraction of total expenses. Since the loss-making fixed costs shrink as a fraction compared to the profitable inference, we should expect profitability to rise with total revenue.

However, all leaks of openai's numbers seem to suggest the opposite: as revenues increase so do the losses.

The indication that OpenAI's inference is profitable is that 3rd party providers host large models for cheaper.

Given that OpenAI is ahead in intelligence, it's also reasonably likely that they are at the frontier of efficiency too.

Your "evidence" for OpenAI's inference not being profitable is apparently based on leaked financials supposedly showing growing losses for reasons entirely unknown.

With their research, training, data centers, chip development, and hardware product development, there seem to be a number of reasons that might explain growing losses.

> Given that OpenAI is ahead in intelligence, it's also reasonably likely that they are at the frontier of efficiency too.

Frontier labs have no incentive to be at the frontier of efficiency.

Claude still leads the pack in general intelligence yet has the worst efficiency by far.

I don’t pay OpenAI’s bills - I pay what they charge me. Their cost accounting isn’t relevant to a user.

Argument was that open ai cannot be profitable with this. But sure, use it while you can.

You can make the other argument that China subsidizes the price and that they can't be profitable at this pricing level. From an industrial strategy standpoint, they already do this for many other industries with huge subsidized state loans.

So we can go round and round on this, each with our made-up objections about how it's temporary or unrealistic or impossible or whatever, or we can just accept the prices as listed and use that to guide our economic decisions.

Private companies cannot play that game too long. Profit from current state of AI is a mirage and sooner or later stuff will hit the fan.

Just look at the prices that inference providers charge for small models. The argument that these unit economics are negative is trivial to disprove.

DeepInfra sells DS v4-flash at 0.08 in, $0.18 out. Gemma4 they sell for $0.07 in, $0.34 out. OpenAI's price for luna is $0.20 in, $1.20 out.

Why would you assume OpenAI is somehow uniquely incompetent at making small, fast models? And that they're worse at serving it than DeepInfra? Any observer can see they are making money here.

I never understand why people who are convinced there is a big con just don't check market prices and see if there's money to be made.

That doesn't mean their business is great -- they're losing tons of money, but it's because they spend too much on fixed costs, and they can't stop spending money on training next generation models with no end in sight, not because the inference is margin negative, which is a flimsy idea that just clouds the actual business issue.

Because Openai is just another player. Nothing really special for now. Their valuation is ridiculously overblown.

Didn't OpenCode CTO state they could replicate deepseek pricing on rented hardware?

There's a difference between the Deepseek.com provider lunch pricing and the pricing every other provider is doing now.

Right now DS4-Pro-0813 is available from multiple providers for $1.32/million input tokens[1].

It's pretty easy to work backwards from B200 and electricity prices and see this is profitable even without the heavy serving optimization these providers are doing[1.5].

The OpenCode CEO said: "inference is very profitable and probably a good opportunity to understand some basic business math"[2] and "the inference we do is already profitable and that's with some middlemen involved"[3]

If at this point people don't believe inference can be profitable, and providers can turn the prices up and down to choose exactly how profitable they make it I don't know what to say.

[1] https://openrouter.ai/deepseek/deepseek-v4-pro-0813#provider...

[1.5] https://www.seangoedecke.com/ai-inference-is-obviously-profi...

[2] https://x.com/thdxr/status/2042277156940587469?lang=en

[3] https://x.com/thdxr/status/2042614323344818520

what if it was because of quantization and they haven't released the new benchmarks for it?

Anything which changes the model needs new benchmarks I guess to compare with other models, otherwise you can benchmark Fable, and distill it to student model and keep claiming this is the Fable model

ARC Prize has retested Luna after the discount and validated identical performance.

(Also, quantization isn't inherently bad or damaging when done properly, e.g. QAT).

These APIs are used heavily by enterprises at scale; with lots of performance telemetry, live evals, etc. You can't really silently nerf API models at scale without people noticing.

Of course, what I said doesn't apply to non-API consumer sub models; there's many documented and officially confirmed instances of under-the-hood "juice/effort" adjustments. (Juice = a number your effort tier maps to underneath the hood; much like Inkling's effort=0.00 to 0.99).

Was it?

Given the timing, I think they A. shat their pants since Deepseek flash just came out with insane pricing before the price hikes, and B. Anthropic is really struggling in model tiers below opus.

It was smart for them to cut prices regardless of whether they had 80% efficiency gains or not

Why would anyone car what the launch price is? Comparing launch pricing is just an odd thing to do.

Because labs can learn to optimize inference post launch, plus can move to use bigger/better clusters depending on demand. It is not impossible to imagine Qwen cuts prices further with QAT/MTP-like improvements.

>If you compare launch pricing

Why?