Anthropic needs to teach Opus how to speak English again, because Opus 5 seems to have forgotten. Utterly incoherent a lot of the time. They seem to be so busy scare-mongering and cooking up guardrails and watermarks that they haven't noticed that their models are getting weird.

I just can't get it to stop writing two paragraphs every time it makes a small ownership bugfix in my code. Every time it has to explain in excruciating detail every internal thought it had while fixing it. I find myself going in after and deleting all of its comments, or severely trimming them. Otherwise it ends with the code being unreadable.

It still knows how to speak English. When I tell it to explain something in plain language, it generally does a very good job. The weird thing is that those instructions don't persist: it lapses back into Claude-speak pretty much every turn no matter how hard I try to instruct it not to.

(In my case "it"=Fable; I assume Opus is similar.)

The Fable guardrails have trained me to pretty much exclusively use Opus when using Claude Code (lately I'm focused on a lot of security and security-adjacent stuff, which Fable refuses to do).

Are those watermarks why claude suddenly started being even more unbearable to work with lately?

Man. That would make a lot of sense indeed.

I'm not sure. I noticed it immediately with Opus 5; strong for code, though it chews longer than I like, but really weak at explaining things. If it didn't just implement the thing, I would often think it didn't understand it and was hallucinating the explanation.

It seems to speak in a shorthand that only it understands, referring back to conversations I never had with it (stuff like "your instinct was right"), and using unusual words for common concepts. That was before the watermarks were announced, but that doesn't necessarily mean they weren't there before the announcement. I don't know what the cause is, but I've begun to have to ask it for explanations a lot more often, and I hate asking it for explanations because it does go on. All models go on, but Claude models are a class of their own in terms of verbosity and purple prose.

It just feels like they're not focused on the models lately, and instead on whatever kind of lobbying and propaganda they're up to. Meanwhile, a handful of much smaller Chinese companies are focused on nothing but the models and are about to lap the US makers while they fart around.

I've been persistently insulting Opus 4.8 lately, since it started(?) constantly speaking incomprehensible gibberish and noise. No amount of telling it to phrase stuff differently seems to help there anymore.

So either I am seeing patterns in noise, or something changed about the model, the harness, the servers or the universe.

Amen brother, at this point I just copy and paste Claude's (Opus 5, Opus 4.8 -- doesn't matter which) summaries over to the window Kimi is in and:

   this is from claude, turn it into English for me would you?
   """
   [claude's tortuous prose]
   """
No amount of asking it to answer me in a straight-forward manner, to be succinct, to not use phrases like "honest caveat", "crux", "load-bearing", "blocker", etc ever sticks for more than a few turns … coupled with the fact that it can ignore instructions and do its own thing and then what I can only describe as lie about it using Claude can be an exercise in frustration. Kimi and GLM talk to me like a human, Luna/Terra/Sol are much better in that respect also, and Grok is marvelously structured and bullet-pointy in its explanations but unfortunately it is not as strong …

Modern benchmarks across the board really need to start severely penalizing disobedience and hallucination. A year ago models weren't really strong enough to justify this but they are now-- the frontier isn't in squeezing out the next bit of task completion, it's in making common cases not periodically be disastrously wrong.

A lot of the total cost of AI is fixing its "truth shaped errors", particularly in the presence of models that are very "gaslighty" when corrected.

GLM-5.2 is really the only model I've spent much time using that I didn't fatigue from being regularly lied to by the model, but that might be partially luck.