I've gone back to 4.8.
5 would constantly veer of in random directions if not working from 100% strict and narrow instructions.
I find it weird there's not more discussion here on HN on how the most used model now has clearly degraded in quality and it seems we've hit a peak and are on a downslope - because the model is clearly smaller or more economical for Anthropic no doubt about it, and the benchmaxxing they do is pure marketing bs - Fable in my view has also been not much better than 4.6 or 4.8 after a few days, disregarding the insane amounts of astroturfing and marketing everywhere.
Theres thousands of threads of twitter, reddit and the internet at large but silence here. Weird but not weird as crypto bs was also insufferably rampant here for a while.
Personally i think we've hit the top of the subsidisation phase and prices will probably 10-15x soon as foreshadowed with both API price policy changes from all the big providers, and now the 1100% deepseek API price changes from yesterday, this could domino into a market implosion and an AI winter, because expecting growth from the bizarre bubble carousel investments with little ROI atm is just not viable.
A bit worried about this as i've already grown quite accustomed to these tools.
> I find it weird there's not more discussion here on HN on how the most used model now has clearly degraded in quality and it seems we've hit a peak and are on a downslope
It's not weird, because it's an anecdote, not an accepted fact.
Personally I've not been too happy with Opus 5, but I've had similar experiences with other models previously, feeling like they didn't quite fit with my working style.
So nothing indicates we've hit a peak.
> nothing indicates we've hit a peak
Opus 5 is arguably a regression but GPT 5.6 is pretty strong evidence that we haven't hit a peak. I think I actually prefer Sol to Fable at this point.
I'll say everything indicates we've hit or are near peak for the masses at least (unless you start paying 50x more) but to each his own.
4.6 was best for us and right now yeah OpenAI and others are edging forward, but slower while prices are increasing industry wide as much as 20x, time to completion is increasing wildly and i'm sure they'll do the same over at OpenAI as their compute constraints also start to take a toll ie degrade performance.
In my view 4.6 era was way faster and with less weirdness so we've gone downwards at least in my company, 4.7 was ridiculous, then 4.8 was almost 4.6 level, 5 is even worse than 4.7 - so it's not a bit up and down its down then a little up then further down.
And all of this is against a backdrop of zero ROI in this sector - so it makes sense we've hit a peak and we're now seeing the subsidisation phase begin to falter, will there be better models certainly but only for short amounts before they get quantised (or whatever is happening behind the scenes), and with diminishing returns over huge prices increases and slower responses.
Literally nothing indicates we've hit a peak, but I guess I'm discussing this with someone who thinks every iteration since Opus 4.6 had zero ROI so there's probably not much common ground here.
Four years ago LLMs were sometimes amazing, sometimes wrong, sometimes a huge time sink when it's almost there and you try to herd the tokens but it's like herding cats.
Just today I had the exact same experience. Every single testimonial is the same as I described above, just emphasizing a different bit to defend or attack LLMs or to make a case for nuance.
The two differences have been: (1) the 1.5 trillion dollar data center build out (2) everyone and their cats now has an opinion on "AI" and data centers. Software is not super amazing, nor are new useful features coming out super fast - It's about the same as 4 years ago plus 4 years of average long term progress as we've seen since 1990s,
They will always be wrong sometimes, but it’s becoming less wrong and wrong less often. Everyone will have their own opinion on “good enough”, but if you expect perfection you are bound for disappointment. No need for that when random chance and Murphy’s law will bring enough anyway.
May be getting harder to catch the mistakes but that makes them worse in my view. I'd much rather they make easy to spot mistakes because I don't expect factuality from them anyway just speedy transformation of information I already have available.
In fact it's the lossiest transformation tool I've ever used and it's still useful despite that. If it reaches one nine of reliability that would be huge but given the pace of growth in investment a first nine would cost an absurd amount of money, and the second and third nine would cost about the Earth's GDP
Or, the hard to catch mistakes were always there, and now we focus on them instead of the obvious ones that have been eliminated.
I still see obvious mistakes, so not eliminated.
The number of mistakes are also about the same.
But more mistakes are harder to catch. The output is more polished and convoluted and that make errors harder to catch. I don't want that.
I think you misunderstand what i'm saying: the companies are not profitable yet, it's the ROI on the investments, not that these tools are useless, very much the opposite, but the business model is not viable, hence the price increase and degrading quality, slower responses etc. I agree theres still progress but its slowing and we're probably near peak.
I have a chunky bit of functionality in my hobby app using babylon.js to render 3D worlds using things like portals and LoD rendering to manage the visual load. I built it out with a combo of Fable and Opus 5.
I too got fed up with the prose of Opus in particular, and tried going back. Unfortunately, the previous models were less able to hack it. The prose was better but progress was worse.
It wasn't just conversation and comments. Some of the function names were wild. Like it instead of something like "isSolidWall(x)" it would write something like "weightyNotEphemeral(x)" or something - that's not quite it, but it really did embed overwrought antithesis into the identifier instead of a straightforward positive predicate.
> "isSolidWall(x)" it would write something like "weightyNotEphemeral(x)"
Is this a literal example? That is wild.
It is not the literal example because I told it to rename the function, but the function had the form thisNotThat for a boolean predicate that had a far more conventional name.
I haven't had time to complain on here because I spend all my time trying to work out wtf Opus 5 is talking about and googling words I've never seen or heard in 40 years as a native English speaker.
I’m either out of touch with contemporary terminology but Opus 5 dropped “pre-mortem” on me today. Figured I’d just figure it out with more context.
It's project management jargon: https://en.wikipedia.org/wiki/Pre-mortem
> prices will probably 10-15x soon as foreshadowed with both API price policy changes from all the big providers
I'll take a sportsperson's bet with you that prices per unit of inference will be far far cheaper in one year from now then they are today. I think the trend of cheaper for better/equal inference quality will continue hard.
I'm sorry, but: why?
Hardware prices have spiked.
I know software is optimized over time but it's a very slow process.
We can use OpenRouter pricing to get an idea about what competitive inference pricing is like without R&D or other costs, and indeed we'd be screwed if we had to pay those rates. We'd go from 100-200 USD to 2000-4000 USD/m.
Not sure what you're talking about here. Kimi K3 is frontier scale and sells competitively at $2.80 input, $14 output per 1M tokens.
edit: oh you mean month? Sure, but then it fully depends on your usecase. I agree that subscriptions are heavily subsidized though.
no if prices were that high, millions of developers would switch to hand writing, open source LLM help and outsourcing at those levels, AI companies know this, you should too.
If OpenRouter prices aren't competitive, the providers would need a cartel to sustain that. It's a commodity market.
Yes they will buy openrouter, pass AI laws and regulations so you all pay extra $$$.This always happen with new real industry. Remember how they built cities for cars. This new laws will day all innovation and slow it for few more decades too but this time they will let China win instead because in China they have all of it aligned, no need to have cartel there. Openrouter and it's new owner will restrict othe models, because they will be part of cartel too.
It is bizarre. There has been no statement, no mention of even hearing concern about Opus 5.
I presume something is forthcoming, but it may be they don’t want to come empty handed—-5.1 is intended to “fix the glitch.”
Yeah. I set my default back to 4.6 and only use 5 for code review etc. Also save s alot of tokens...
yeah I feel the same way with Opus 5 too. If I ask it to do something, it would go ahead and rewrite unrelated things and then in a less performant version of it.
I’ve instead moved to GLM, at least it has the courtesy to ask some steps of the way what I wanted exactly and only work on what I asked.
>it seems we've hit a peak and are on a downslope
Sol and Fable are great; we haven't hit a peak, Anthropic just tried to pull a fast one on its customers with Opus 5.0.
Opus 5 has a habit of taking what I asked for, doing something tangentially related to it, and then lying to me and saying it did exactly what I asked.
It even adds comments to say you asked for [thing] then leaves snarky comments about when you correct it.
At that point I decided it's just not worth the babysitting that's required, and you are better off working entirely with other models.
If the harness itself was open source then maybe we'd be able to wrap it up in a reasonable layer of sanity.
[flagged]
> A bit worried about this as i've already grown quite accustomed to these tools.
Ka-ching