Anthropic, if you're listening - by the time this crops up on Reddit, the front page of HN, etc.... you should be expecting calls from CEOs of major corporations next threatening to abandon ship...

We've seen this pattern before several times.. I hope they are listening and address this publicly.

I'm not sure what is going on, some users report it works fine or great, others report the degradation. I've experienced both at times, and it's been such a different experience it has made me wonder if there isn't some sort of hidden A/B test or model router in the background silently downgrading requests at times.

Also, regarding the subsidized access to models, in my opinion, the frontier companies owe it to society to continue it. After mining the public content of all of humanity, I personally feel it is a service they owe the public in return.. not that my feeling of this counts for anything though.

You are expecting consistent QoS from a randomly sampled mathematical function.

This is a fair comment, although I would add that is not my expectation personally.

I think nondeterminism does not have to be the same as non-coherency - i.e. just because something is randomly sampled does not mean the result has to be incoherent or inconsistent.

Also, if we speak purely about LLM based on how they are implemented now, I feel that is different than speaking about artificial intelligence. The field of AI is much more than just an LLM by itself, and the promise of these companies is not just LLM, whether the underlying models are limited to that technology or not.

FWIW, I have built rule based expert systems, used logic based reasoning systems like NASA CLIPS or rete-algorithm based systems, mathematical/symbolic solvers, written plenty of terrible case/conditional logic in programming languages, worked with ML in its infancy and now worked in AI/LLMs - I give this context only to clarify that I understand what an LLM is and isn't.

With all that said, LLMs have allowed humanity to make advances, at great cost to society (IMHO), and I'd hate to see the opportunity be wasted.

There is plenty of room past "attention is all you need" still to do incredible work, especially at the crossroads between deterministic and nondeterministic behaviors.

True, but this function wasn't handed down to us from the gods, it can be shaped by training and RLHF processes. They still have a little bit of control over its output.

Strong words coming from a blob of oxygen, carbon, and nitrogen.

No, they're expecting to see a failure rate consistent with previous failure rates, not periods of low failure rates and other periods of high failure rates, with the same model.

And you can expect more consistency from SOTA models than you can from an old model like GPT3--you agree, right?

GP expects the same level of consistency throughout their time using the same model. Not request-to-request, more like day-to-day.

I get consistency out of the ridiculous pile of quantum noise that is my CPU. Plenty of random processes produce consistency when handled properly. An LLM won't usually give identical outputs for identical inputs, but it's entirely reasonable to expect similar output for similar input when considered on broad metrics like "intelligence" or flowery language or staying on task.

[dead]

I swear I've noticed a sustained and consistent degradation in intelligence of Opus 5 since its launch. I seem to experience this every time a new model is launched. This time I have a baseline: Sol. I switched from Sol because performance was better with Opus 5. I just switched back because performance is much better with Sol. Either Sol became much better over time, or Opus 5 became much worse. I'm certain it's the latter.

> owe it to society

In America? lol if only, only a law would get them to act for that reason, maybe not even that these days..

Or, maybe competition.. but your point is taken.

I like to hope that those in positions of power do have a sense of morality though too.. but their worldview is quite different than an ordinary citizen.

ROTFL

They have usage metrics that are 100x more unbiased and rich than random internet complaints.

They can see which models people are using, how irritated they are during conversations, and how often people drop or shift to a different model. There is just no world where listening to random complaints on the internet gives them information they don't get from actual conversation logs.

please god I want anthropic to fail, so that they could learn that their current approach is wrong