Price is confounded by VC subsidies, economies of scale, and inference optimizations. I think a more interesting chart would be ARC AGI vs forwards pass flops or ARC AGI vs training tokens. Of course we don't have those numbers for the closed source models or even some of the open weight ones.

DeepSeek V4 Flash 0731 is an open-weights model which means price is determined by competition/invisible hand of the marketplace: https://openrouter.ai/deepseek/deepseek-v4-flash-0731

With the exception of cache costs, all providers have similar input/output costs.

Not counting the cost of making the model, which is subsidized by… someone? The chinese gov i think?

DS comes out (one of, or) the most successful quant fund in China.

They don't strictly need any kind of subsidies.

FWIW they have a funding round planned (kerfuffle about leaks from CEO presentation few weeks back) -- presumably because infrastructure needs have ballooned.

Naturally there will be some PRC government interest in one of their flagship AI companies. From what is visible seems to be more along the lines of ensuring that DS gets its fair share of resources -- e.g. Xi Jinping meeting founder and positive comments about success of DS means that (hypothetically) Alibaba can't screw DS too much on infra charges to kill off a 'competitor'. Also would imagine that DS's top guys have been clearly identified and will have been 'discouraged' from going to work for one of the SV polycules. But even here as much carrot as stick -- none of the DS top guys will ever need to work again except for love of the job.

> DS comes out (one of, or) the most successful quant fund in China

> They don't strictly need any kind of subsidies.

You understand how these two sentences directly contradict each-other, yeah? The money-losing operating of training a model is paid for by momey earned from prior investments. So… the work is “subsidized” by its parent company’s investments in it.

[deleted]

Subsidized by inference profits and volume.

Is deepseek actually turning enough of a profit off inference to fully pay for training the next model? And do those profits depend on releasing model weights somehow?

weak argument. deepseek v4 flash is open weight, you can easily find other providers with competitive price with Deepseek (except for input caching), some even half as cheap.