Because DeepSeek is not "a month or two" behind as claimed in the article.

These open models still did not beat February's Mythos / Fable 5.

DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.

Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.

It's plausible that open models are 6 - 12 months behind, and there is no "good enough". As long as progress doesn't slow down, leading labs have nothing to fear.

Perhaps on certain benchmarks and for certain work, but anecdotally I've not been able to see a difference between it and Opus on a lot of dev work (web, Go, iOS/AppleTV native, scripting, general tasks)

I was thinking about this earlier today and I came to the following question:

If you had a model 10x as capable as the best model out today, but it cost 100x more, would there be a market, and, if so, how big?

I think there would be a market and I think it would be large.

So, I agree.

Many people would, and you'll find that they're building crappy webapps where you dont need SoTA. Like seriously who needs these frontier models?

Unless you're doing some extermely difficult post-grad lvl research, you do not need a 100x PhD research assistant, especially not for whatever silly SaaS product most people are building.

There's people at my job that get so much more done than everyone else using Fable/Opus/Astra. and all they use is the fastest cheapest models. I'd say the people who are using sota models for everything are doing it just because they prefer to be lazy.

You simply do not need these frontier models, they outgrew most people's needs 6 months ago, but for some reason people still want to run a 700k rack of gpus full throttle to center a div for them.

Doing what? How many jobs involve solving Millennium Prize math challenges?

99% of everything is CRUD LoB apps.

I am coding CRUD apps with a mix of astra, sol 6.1, fable and opus 5.5. A more capable model would still benefit me imo. Being able to follow high level guidance better, and being able to harness other models for each task would be a big improvement.

Do you know how what you're doing, or do you find yourself working on things you dont understand and need the best model because it's the only way to push your own capabilities (because you're avoiding learning how to do the thing yourself)?

Not asking to be mean, I just genuinely dont know why you'd need the frontier for basic applications.

It's a matter of bandwidth. The more I can offload onto the model, the more I can accomplish. For example, I had to do a lot of security work over the last 2 weeks to get ready for an event. This requires handholding current models on many fronts, like: 1) Do they actually implement the security fixes correctly. 2) Do their fixes create any new edge cases. 3) Do their fixes compromise existing interfaces or API surfaces.

I cannot trust current models to find all the necessary context, or to make what I consider to be good trade offs. A much more capable model would be able to see my existing patterns (or at least not have context rot make them blind to my convention docs) and make trade offs I agree with much more consistently, and I'd be able to do more with my time.

I've actually found models to be pretty poor at driving things I don't know well, so I generally don't do that unless its general design/product exploration and the end product code is throw-away.

Agree, will see after the dilution and CoT hack fixed, will they keep the pace now. MiMo had some good numbers recently because it's discovered that the post evaluation RL directly exposes answers to models, so RL and evaluation is runied.

that just sounds like openai/anthropic cope/propaganda, based on absolutely nothing objective lol

even their harnesses are far surpassed by pi and opencode at this point

also sick 'rumors' lmao, apparently marketing through rumors is in vogue these days

>> that just sounds like openai/anthropic cope/propaganda, based on absolutely nothing objective lol

Nah. There are benchmarks. They are free to look at. And they paint a very clear picture.

Yeah the picture they paint is that they're mostly bullshit

There absolutely is "good enough" and I agree with this author: DeepSeek 4.1 Flash is plenty good enough for all the things I would trust an AI to do at my job.