If what you're saying is true and accurate, then US-based AI labs are in big trouble. The only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models across US federal and state governments and large enterprises (especially ones with federal government contracts), but even still, US AI labs will probably lose out massively on international market if a smaller model can match SOTA of just a few months ago.

There's no way large companies outside the US will pay the "US AI lab" premium if they can get the same workloads done at a fraction of the cost using open-weight models that they can self-host and optimize/fine-tune on.

I think you're overlooking the fact that for long-horizon tasks, even small errors compound over time and can lead to catastrophic outcomes.

For simple queries, we have reached the threshold since the beginning of the year, and models are good enough from every provider to make a meaningful difference between one another. (ChatGPT, Claude, Gemini, Grok, MuseSpark, Kimi, DeepSeek, GLM...)

The real unlock will be, and you can already see it with GPT-5.6 and Fable-5, to delegate complex enough tasks that will take more than 24 hours to get done and they will not lose track. I'm not talking about a loop, but the actual intelligence to recover from these compounding errors that accumulate in dumber models.

We're still a long way from the intelligence needed to let one of these agents go ahead and supervise multiple layers of sub-agents underneath to do complex orchestration. The future looks very promising and exciting. Imagine having the possibility of a Frontier model orchestrating as many sub-agents as needed that are running on cheaper models like DeepSeek.

That doesn’t make sense. It’s not like SOTA models are error free, yet we still use them.

You use Fable 5 right? If that’s good enough for you now, why wouldn’t a Chinese model that’s as good as Fable 5 but at 10% the cost be good enough in 6 months?

I think we put up with Fable's occasional hiccups because there's nothing better at the moment.

I use Claude Code semi-heavily for my small business, and the $100/mo I pay for that is a rounding error compared to the value it provides.

If I can avoid spending an hour or two "massaging" the output from a lower-end model once, or it avoids introducing one load-bearing (sorry, couldn't resist) bug, then that's the entire $100 right there.

Hell, you could argue that the best "coding model" that we have at the moment is the human brain, and people will gladly pay $10,000/mo for one of them.

Arguing over $20 vs $100 for something that actually puts in work just seems insane to me.

The question low cost models will create: Why would you massage output?

Fable 5 is still going to mess things up at any sufficient complexity. The advantage of low cost models with "good enough" intelligence is they can recursively correct. Why? Because it is cheap. Proper requirements and tests and subagents take away increasing amounts of work, at a cost that is not prohibitive.

If you are reviewing code manually you might consider Fable 5 a worse option. As it articulates itself with higher confidence and you already know it is capable, you are may be more likely to miss a mistake. You know to be on guard with a junior engineer. Reviewing a senior who suddenly makes some weird stochastic mistake can be a lot harder. It would be like if the smartest human engineer you knew was capable of some random brainfart in the middle of their massive diff. Imo, much harder to deal with.

Of course, we should keep in mind Fable 5 is only expensive today. It will be cheaper in the future. Autonomous, recursive prompting and improvement is the clear end state. Especially for entities that will always have the budget for that at the SOTA frontier.

> I think we put up with Fable's occasional hiccups because there's nothing better at the moment.

Which was an argument for using every less powerful model since the moment they got useful, right?

When was that? Opus 4.5 maybe? Let's say Opus 4.5 for the sake of the argument. So back then we were like "DeepSeek is not good enough, I need Opus 4.5". Now DeepSeek is better than Opus 4.5. So if Opus 4.5 was good enough back then, DeepSeek is better than that now.

Sure, it's always nicer to have a slightly better model. But the price difference starts mattering a lot more when all the models are already sufficiently good.

To put actual numbers on it, since using AI to start solving all kinds of bottlenecks/inefficiencies in our small business, we've seen monthly net profit go up by around $4,000 USD. These are semi-permanent fixes, and the tech is only partially deployed. I am the only one using it, and I only use it part time.

We've just spun up our first Hermes agent, with direct API access to our main inventory system and that's expected to find another few grand per month in misallocation/inefficiency.

I wouldn't be surprised if we were doing more like $10k/mo higher in 6-9 months' time.

When you're talking about numbers like this, the fact that one AI is $100/mo and another is $10/mo or $40/mo doesn't matter. They could make GLM-5.2, or any other Opus 4.5-class model free and it still wouldn't make sense to deploy in a commercial context.

The other angle I'd approach things from is that Opus 4.5 (and I'd agree with you that that model was the saddle point) was "good enough" for the types of things we were asking it to do back then, but as the models have become more capable the tasks we're asking them to do have also expanded with it.

I know I've personally gone from "hey can fix this race condition with a Redis mutex" 6 months ago to "Independently redesign this full embedded USB stack and QA it end-to-end, working around a specific Kernel bug in macOS Tahoe that requires decompilation to find the source of, while keeping in mind the constraints of our 8-bit AVR chip from 2011" now.

But that said, yes, maybe in 5 years' time we will reach an "intelligence saturation" where the average person won't be able to even conceive of how to use the new SOTA.

I think we're even starting to reach that saturation point now for a lot of people. In my industry (law) plenty of people have tried CoPilot once or twice, or tried ChatGPT a year ago, and as a result have basically dismissed AI as being useless. The setup required to be able to get it to do end to end tasks to your liking is also substantially more work than most people are willing to put in.

"оur first Hermes agent, with direct API access to our main inventory system" – let me assure you that absolutely nothing can go wrong here, mate. /s

[delayed]

Except Fable won’t be costing $100 for enterprises that will be considering the Chinese models.

If $100 Claud Max subscription works for you, then great.

But you have to remember your pricing is subsidized by enterprises that pay hundreds of thousands of dollars each month, if not more, to Anthropic.

For those companies, a Chinese model that can cut their AI spend from $1M/month to $200k suddenly seems attractive.

And unfortunately for the American tech industry, the valuation is based off those enterprise deals, not your $100/month Claude Max subscription.

Didn’t deepseek recently announce prices will go up significantly?

Right now the US dominates everyone else in actual chips in data centers. So even if deepseek etc tries to undercut, they’re very capacity limited.

It's far from significant, it's partially doubled during peak hours. They could 16x it and it would still be two orders of magnitude better value than OAI's $200/mo plan.

It's that good. They are far from capacity limited, and even if they were, you can rent a single MI300X from somewhere like Hot Aisle and get more tk/s than you'll be able to use.

Deepseek is open weight/source ie it will be running on US servers in the US maybe by on each companies own servers.

> not your $100/month..

This is made brutally obvious by anthropics customer support for people with such accounts.

Yeah, fair. If we were talking $2,000/mo vs $200 then the maths starts looking very different.

There are a lot of tasks that are hard for organisations to run consistently but require some intelligence - monitoring logs and metrics for anomalies and security events, backup audits, audit processes in general, ensuring document quality and consistency, database advice and tuning, customer experience management, process optimisation - that are not "long horizon" in the classical sense of each step depending on the last, but are the result of consistency and attention over a long period of time and a large amount of data.

For this genre of task execution can run with limited horizon and is independent but would be too expensive to do with "us frontier tokens", I think for these, there is value in availability of cheaper tokens.

These are not 24 hours of inference with floating point errors accumulating; largely the system guards against errors compounding. Tool failures, compile failures, test failures, etc, push back against the model taking a wrong turn and force it to correct.

Yes it's much easier to have a smarter model that goes straight to the correct answer first, but it may not be necessary or economical. There's a minimum bar for the model where it understands problems and knows the right step to correct them, and above that newer models give diminishing returns.

> it's much easier to have a smarter model that goes straight to the correct answer first

That's basically ASI not AGI, if you agree humans are NGI (natural general intelligence) and make mistakes and wrong decisions in solutions all the time. Right steps with some wrong ones is acceptable though for AGI.

Majority of white collar work absolutely does not require sota models

I have a silly (but honest) question. What's an example or two of a > 24hr task that people are actually asking something to do? Like real life ones.

Decompiling / disassembling and annotating old software, making sure it can build cleanly back to the original binary, and then look for bugs or subtle issues.

Another one I did was a printer data stream translator from an obscure format to PostScript/PDF (or just PNGs), complete with cups support, etc so these old apps can easily be hooked up.

Flash is capable now of running long range defined-goal tasks like this.

> these compounding errors that accumulate in dumber models

While SOTAs handle these errors better, they compound in all models and there's a term for that. It starts with cluster and ends with an expletive.

I wish I could, but I don't see the need for human steering going away soon if the task involves anything novel (see Terry Tao's chat).

> even small errors compound over time and can lead to catastrophic outcomes

So, death sentence even to frontier models?

It is true. I don't care about having infinite frontier-level intelligence, and I don't care if Fable can one-shot frobnicate a klaxelzorp with a benchmark performance of 97%. I doubt most people do, in fact. I just want something that meets the baseline level of intelligence needed to be a really, really good pair programming agent. It shouldn't have any silly dealbreaker issues involving laziness or hallucinations, it should be smart enough to bounce ideas off of, and it should automate doing tedious boilerplate. And - most of all - I want to be able to afford using it as much as I want. That's what has happened here.

I wonder when we crossed the "99 percentile of intelligence for 99% of the usecases" threshold. At this point, the gains seem to be right at the very edge of bleeding edge for narrow and specialized use cases, and wonder if it'll be a sort of diminishing return from here on.

In April

Probably the best counterexample is the games they are able to design. It's still mostly AI slop, few would want to play.

It is amazing how fast it happened. Right now one of my main projects is fully running on DeepSeek flash. My reason was that I was blocked by both of the main US AI labs from working on it because it involves viruses. DeepSeek flash has been killing it since I switched it on, completing the first phase of the project and setting up an iteration in another application space. It isn't the most brilliant model, but it is reliable and I don't have to manage my weekly token allowance. I just spend freely and end up spending only a few dollars a day. Intelligence is going to become a basic commodity. Only special stuff is going to drive us to use special models. And maybe not even that.

> If what you're saying is true and accurate, then US-based AI labs are in big trouble.

Well... I would think that the whole AI industry in the US are working towards public bailouts... Which I guess they'll get under the current administration... So they'll be fine... Nobody there really seems interested in actually creating a profitable business anyway...

> only saving grace might be some sort of a 'national security' proclamation banning the use of state-of-the-art Chinese (and non-US) models

In what kind of sad and failed dystopia is this a "saving grace"? For whom?

> If what you're saying is true and accurate, then US-based AI labs are in big trouble.

I've been working with DeepSeek V4 Flash 0731. I'd say that it's maybe not quite as smart as Opus 4.5, but it's willing to think things through carefully and keep going until it gets a good answer. So it's a decent Opus 4.5 replacement. Just let it cook.

It isn't Opus 5 or Fable 5. But it's nearly free on Open Router, and it's self hostable on a Mac Studio with plenty of RAM, or using an RTX Pro 6000 Blackwell or two. Which is chump change for any company that employs programmers.

It would absolutely have been a frontier model last December.

If programming in the US to become unconditionally 10x more expensive, then the exodus from the US is about to begin.

Could we look at manufacturing, for example the car industry, to predict what will happen?

Only to to find out they end up in a much worse place

I couldn't agree more, and think of all the wasted inference across accounts overpaying for their subscriptions.. Need a secondary marketplace for this stuff.

The bet isn't that people will be able to automatically reply on bugs and rack up API charges.

The bet is on using AI to gain competitive advantage. You don't win the stock market or make the deadliest drone by switching to the cheap model

> You don't win the stock market or make the deadliest drone by switching to the cheap model

Really? How many times a small team has outperformed a much bigger one just because they were "doing it right"?

I have been in software companies where most software produced was bad. Not just the code, the overall design everywhere. So... bad engineers with the most expensive model, or great engineers with cheaper models?

There are more than two options. What about great engineers with great models?

Again, I was answering to:

> You don't win the stock market or make the deadliest drone by switching to the cheap model

The question is not "can you win with the best model?", it is "can you not win without the best model?".

Yeah this is what I’m curious about. How good are they after the benchmarks. I’ve been told yeah they’re good but they’re just building to show off for benchmarks.

The ByteDance folks are apparently training a mythos level model 10T params apparently. If they do would it still be subsidized at these cheap rates?

It's been true for almost every business. "Cheap and good enough" usually trumps "excellent but expensive". Ikea, McDonald's, Ryanair, AliExpress, Aldi - these brands prove that catering to poor people is more profitable than catering to rich people simply because there are so many poor people that their collective spending power outweights the one of rich people.

Well, not universally. It’s a tradeoff. If what you said was universally true Apple wouldn’t exist; Spirit Airlines wouldn’t be bankrupt, etc.

Apple sells to half the American population. And by definition many of them are poor.

Spirit was broken by oil prices which everyone pays the same for. (There is no cheaper jet fuel alternative).

Not a good comparison to the point of wrong conclusions.

>Aldi - these brands prove that catering to poor people is more profitable than catering to rich people

At least here in Germany Aldi isn't even really limited to the poor, it's famously a place where you can run into anyone. Where I used to live in Berlin close to the government district I literally on occasion ran into the chancellor (and her bodyguards). Aspirational shopping where you buy premium goods to pretend to have higher social status honestly seems a bit on its way out. Even middle class people seem to consciously shop more utilitarian now.

[deleted]