> I find it weird there's not more discussion here on HN on how the most used model now has clearly degraded in quality and it seems we've hit a peak and are on a downslope

It's not weird, because it's an anecdote, not an accepted fact.

Personally I've not been too happy with Opus 5, but I've had similar experiences with other models previously, feeling like they didn't quite fit with my working style.

So nothing indicates we've hit a peak.

> nothing indicates we've hit a peak

Opus 5 is arguably a regression but GPT 5.6 is pretty strong evidence that we haven't hit a peak. I think I actually prefer Sol to Fable at this point.

I'll say everything indicates we've hit or are near peak for the masses at least (unless you start paying 50x more) but to each his own.

4.6 was best for us and right now yeah OpenAI and others are edging forward, but slower while prices are increasing industry wide as much as 20x, time to completion is increasing wildly and i'm sure they'll do the same over at OpenAI as their compute constraints also start to take a toll ie degrade performance.

In my view 4.6 era was way faster and with less weirdness so we've gone downwards at least in my company, 4.7 was ridiculous, then 4.8 was almost 4.6 level, 5 is even worse than 4.7 - so it's not a bit up and down its down then a little up then further down.

And all of this is against a backdrop of zero ROI in this sector - so it makes sense we've hit a peak and we're now seeing the subsidisation phase begin to falter, will there be better models certainly but only for short amounts before they get quantised (or whatever is happening behind the scenes), and with diminishing returns over huge prices increases and slower responses.

Literally nothing indicates we've hit a peak, but I guess I'm discussing this with someone who thinks every iteration since Opus 4.6 had zero ROI so there's probably not much common ground here.

Four years ago LLMs were sometimes amazing, sometimes wrong, sometimes a huge time sink when it's almost there and you try to herd the tokens but it's like herding cats.

Just today I had the exact same experience. Every single testimonial is the same as I described above, just emphasizing a different bit to defend or attack LLMs or to make a case for nuance.

The two differences have been: (1) the 1.5 trillion dollar data center build out (2) everyone and their cats now has an opinion on "AI" and data centers. Software is not super amazing, nor are new useful features coming out super fast - It's about the same as 4 years ago plus 4 years of average long term progress as we've seen since 1990s,

They will always be wrong sometimes, but it’s becoming less wrong and wrong less often. Everyone will have their own opinion on “good enough”, but if you expect perfection you are bound for disappointment. No need for that when random chance and Murphy’s law will bring enough anyway.

May be getting harder to catch the mistakes but that makes them worse in my view. I'd much rather they make easy to spot mistakes because I don't expect factuality from them anyway just speedy transformation of information I already have available.

In fact it's the lossiest transformation tool I've ever used and it's still useful despite that. If it reaches one nine of reliability that would be huge but given the pace of growth in investment a first nine would cost an absurd amount of money, and the second and third nine would cost about the Earth's GDP

Or, the hard to catch mistakes were always there, and now we focus on them instead of the obvious ones that have been eliminated.

I still see obvious mistakes, so not eliminated.

The number of mistakes are also about the same.

But more mistakes are harder to catch. The output is more polished and convoluted and that make errors harder to catch. I don't want that.

I think you misunderstand what i'm saying: the companies are not profitable yet, it's the ROI on the investments, not that these tools are useless, very much the opposite, but the business model is not viable, hence the price increase and degrading quality, slower responses etc. I agree theres still progress but its slowing and we're probably near peak.